FlashMLA is DeepSeek's MIT-licensed CUDA kernel library for optimized attention in DeepSeek-V3 and DeepSeek-V3.2-Exp style inference. It includes dense MLA decoding plus sparse attention kernels for DeepSeek Sparse Attention, with README-reported H800/CUDA metrics up to 3000 GB/s, 660 TFLOPS, and sparse 640/410 TFlops paths. It has 12.7K+ GitHub stars.
Best DeepEP Alternatives
2 editor-verified alternatives · DeepEP overview →
source: tools.alternatives · stored order · active records only; review scores are annotations and never change membership or order
DeepGEMM is DeepSeek's open-source library of FP8 matrix multiplication CUDA kernels optimized for LLM inference and training on modern NVIDIA GPUs. It provides efficient GEMM operations using 8-bit floating point precision that reduce memory bandwidth requirements while maintaining model accuracy. Designed for integration into inference engines and training frameworks. Over 6,300 GitHub stars.
Open-source DeepEP alternatives
FlashMLA, DeepGEMM — see all open-source developer tools.
FAQ
What is the best DeepEP alternative?
FlashMLA tops our editor-verified list of 2 DeepEP alternatives, scoring 80/100 in our hands-on review.
Are there open-source DeepEP alternatives?
Yes — FlashMLA, DeepGEMM are open source.