Skip to content
aicoolies logo
DeepSeek logo

FlashMLA

DeepSeek's optimized attention kernel for Multi-Head Latent Attention

FlashMLA is DeepSeek's MIT-licensed CUDA kernel library for optimized attention in DeepSeek-V3 and DeepSeek-V3.2-Exp style inference. It includes dense MLA decoding plus sparse attention kernels for DeepSeek Sparse Attention, with README-reported H800/CUDA metrics up to 3000 GB/s, 660 TFLOPS, and sparse 640/410 TFlops paths. It has 12.7K+ GitHub stars.

About FlashMLA

FlashMLA is a production CUDA kernel optimizing Multi-head Latent Attention (MLA) inference on NVIDIA Hopper GPUs (H100/H800). Developed by DeepSeek-AI, the kernel addresses a critical bottleneck in modern LLM inference: attention operations are memory-bound, not compute-bound, so traditional kernel designs waste GPU compute while waiting for memory. FlashMLA achieves 3000 GB/s memory bandwidth utilization in dense inference and 660 TFLOPS in compute-bound configurations, reaching near-theoretical peak performance through kernel-level scheduling that overlaps CUDA Core operations, Tensor Core operations, and memory transfers.

The technical implementation merges several optimization strategies. FlashMLA uses programmatic dependent launch to overlap the splitkv_mla and combine kernels, reducing synchronization overhead. A tile scheduler allocates jobs to streaming multiprocessors for load balancing. The kernel supports BF16 precision natively and implements paged KV cache with 64-byte blocks, dramatically reducing memory pressure compared to contiguous allocations. For sparse workloads using FP8 KV cache, throughput reaches 410 TFLOPS. Variable-length sequence handling (padding-free) further improves efficiency for batched inference.

DeepSeek released FlashMLA as part of their open-source week initiative, targeting inference infrastructure teams operating large model deployments. The kernel integrates with vLLM and SGLang inference engines, allowing drop-in speedups for production LLM APIs. Infrastructure providers hosting Qwen, DeepSeek, or other MLA-based models benefit from 2-3x throughput improvements. For research teams fine-tuning MLA architectures, FlashMLA provides reference implementations demonstrating memory-optimal kernel design applicable beyond MLA to general attention optimization.

Pricing & Platform Specs

Pricing Summary

Free and 100% open source under the MIT license. FlashMLA has no software licensing costs or commercial restrictions; deployment costs are determined entirely by underlying NVIDIA Hopper GPU compute infrastructure.

full pricing breakdown →

Supported Platforms

CUDA, NVIDIA GPUs, Python/C++ integration

Explore categories, tags & use cases

DeepSeek's expert-parallel communication library for MoE model training

DeepEP is DeepSeek's open-source communication library optimized for expert-parallel training of Mixture-of-Experts models. It provides efficient GPU-to-GPU data routing for distributing tokens to expert networks across multiple devices during MoE model training and inference. Enables the distributed expert parallelism that powers DeepSeek's competitive model efficiency. Over 9,100 GitHub stars.

Open Source

DeepSeek's FP8 general matrix multiplication kernels for efficient inference

DeepGEMM is DeepSeek's open-source library of FP8 matrix multiplication CUDA kernels optimized for LLM inference and training on modern NVIDIA GPUs. It provides efficient GEMM operations using 8-bit floating point precision that reduce memory bandwidth requirements while maintaining model accuracy. Designed for integration into inference engines and training frameworks. Over 6,300 GitHub stars.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is FlashMLA?

FlashMLA is DeepSeek's MIT-licensed CUDA kernel library for optimized attention in DeepSeek-V3 and DeepSeek-V3.2-Exp style inference. It includes dense MLA decoding plus sparse attention kernels for DeepSeek Sparse Attention, with README-reported H800/CUDA metrics up to 3000 GB/s, 660 TFLOPS, and sparse 640/410 TFlops paths. It has 12.7K+ GitHub stars.

Is FlashMLA free?

Yes — FlashMLA is open source and free to use. Free and 100% open source under the MIT license. FlashMLA has no software licensing costs or commercial restrictions; deployment costs are determined entirely by underlying NVIDIA Hopper GPU compute infrastructure.

Is FlashMLA open source?

Yes — FlashMLA is open source.

Is FlashMLA still maintained?

Yes — FlashMLA is active. Its listing was last verified on September 6, 2026.

What are the best FlashMLA alternatives?

The first editor-selected FlashMLA alternatives are DeepEP, DeepGEMM.

How does FlashMLA score in our review?

The published editorial review lists FlashMLA at 80/100 overall across speed, privacy, and developer experience. Check the review's evidence status and test metadata for its verification level.