Skip to content
aicoolies logo
DeepSeek logo

DeepGEMM

DeepSeek's FP8 general matrix multiplication kernels for efficient inference

DeepGEMM is DeepSeek's open-source library of FP8 matrix multiplication CUDA kernels optimized for LLM inference and training on modern NVIDIA GPUs. It provides efficient GEMM operations using 8-bit floating point precision that reduce memory bandwidth requirements while maintaining model accuracy. Designed for integration into inference engines and training frameworks. Over 6,300 GitHub stars.

About DeepGEMM

DeepGEMM provides optimized CUDA kernels for general matrix multiplication using FP8 (8-bit floating point) precision, the fundamental compute operation that dominates both LLM training and inference. By reducing precision from the standard FP16 to FP8, these kernels roughly double throughput and halve memory bandwidth requirements while maintaining model quality through careful handling of the reduced dynamic range.

The kernels are specifically optimized for the matrix shapes and access patterns that occur in transformer model computation, including attention projections, feed-forward network layers, and the expert computations in MoE architectures. Rather than general-purpose FP8 GEMM implementations, DeepGEMM provides kernels tuned for the specific workloads that LLM serving requires, extracting performance that generic libraries leave on the table.

With over 6,300 GitHub stars, DeepGEMM completes DeepSeek's trilogy of open-source compute infrastructure alongside FlashMLA for attention and DeepEP for expert parallelism. Together these libraries provide the low-level compute primitives needed to train and serve large models with the efficiency that DeepSeek has demonstrated. The MIT license enables unrestricted use in both research and commercial inference deployments.

Pricing & Platform Specs

Pricing Summary

100% free and open source under the BSD-2-Clause license developed by DeepSeek-AI ($0 software cost). Clean, ultra-fast FP8 matrix multiplication (GEMM) library optimized for NVIDIA Hopper Tensor Memory Accelerator (TMA) and MoE architectures.

full pricing breakdown →

Supported Platforms

CUDA, NVIDIA GPUs (Hopper+ recommended)

Explore categories, tags & use cases

DeepSeek's optimized attention kernel for Multi-Head Latent Attention

FlashMLA is DeepSeek's MIT-licensed CUDA kernel library for optimized attention in DeepSeek-V3 and DeepSeek-V3.2-Exp style inference. It includes dense MLA decoding plus sparse attention kernels for DeepSeek Sparse Attention, with README-reported H800/CUDA metrics up to 3000 GB/s, 660 TFLOPS, and sparse 640/410 TFlops paths. It has 12.7K+ GitHub stars.

Open Source

DeepSeek's expert-parallel communication library for MoE model training

DeepEP is DeepSeek's open-source communication library optimized for expert-parallel training of Mixture-of-Experts models. It provides efficient GPU-to-GPU data routing for distributing tokens to expert networks across multiple devices during MoE model training and inference. Enables the distributed expert parallelism that powers DeepSeek's competitive model efficiency. Over 9,100 GitHub stars.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is DeepGEMM?

DeepGEMM is DeepSeek's open-source library of FP8 matrix multiplication CUDA kernels optimized for LLM inference and training on modern NVIDIA GPUs. It provides efficient GEMM operations using 8-bit floating point precision that reduce memory bandwidth requirements while maintaining model accuracy. Designed for integration into inference engines and training frameworks. Over 6,300 GitHub stars.

Is DeepGEMM free?

Yes — DeepGEMM is open source and free to use. 100% free and open source under the BSD-2-Clause license developed by DeepSeek-AI ($0 software cost). Clean, ultra-fast FP8 matrix multiplication (GEMM) library optimized for NVIDIA Hopper Tensor Memory Accelerator (TMA) and MoE architectures.

Is DeepGEMM open source?

Yes — DeepGEMM is open source.

Is DeepGEMM still maintained?

Yes — DeepGEMM is active. Its listing was last verified on September 6, 2026.

What are the best DeepGEMM alternatives?

The first editor-selected DeepGEMM alternatives are FlashMLA, DeepEP.