FlashMLA is DeepSeek's MIT-licensed CUDA kernel library for optimized attention in DeepSeek-V3 and DeepSeek-V3.2-Exp style inference. It includes dense MLA decoding plus sparse attention kernels for DeepSeek Sparse Attention, with README-reported H800/CUDA metrics up to 3000 GB/s, 660 TFLOPS, and sparse 640/410 TFlops paths. It has 12.7K+ GitHub stars.
Best DeepGEMM Alternatives
2 editor-verified alternatives · DeepGEMM overview →
source: tools.alternatives · stored order · active records only; review scores are annotations and never change membership or order
DeepEP is DeepSeek's open-source communication library optimized for expert-parallel training of Mixture-of-Experts models. It provides efficient GPU-to-GPU data routing for distributing tokens to expert networks across multiple devices during MoE model training and inference. Enables the distributed expert parallelism that powers DeepSeek's competitive model efficiency. Over 9,100 GitHub stars.
Open-source DeepGEMM alternatives
FlashMLA, DeepEP — see all open-source developer tools.
FAQ
What is the best DeepGEMM alternative?
FlashMLA tops our editor-verified list of 2 DeepGEMM alternatives, scoring 80/100 in our hands-on review.
Are there open-source DeepGEMM alternatives?
Yes — FlashMLA, DeepEP are open source.