Skip to content
aicoolies logo
Mooncake logo

Mooncake

Disaggregated KV cache storage and transfer for LLM serving

Open-source infrastructure for disaggregated LLM serving that pools KV caches across prefill and decode workers, with high-performance transfer, distributed storage and integrations for vLLM and SGLang.

About Mooncake

Mooncake is an open-source infrastructure layer for large-scale LLM inference and training. Its KV-cache-centric disaggregated architecture separates prefill and decode work while pooling otherwise underused CPU memory, DRAM and SSD/NVMe resources across a GPU cluster. The project combines Transfer Engine for topology-aware, multi-NIC movement across heterogeneous networks and accelerators with Mooncake Store for distributed KV-cache and model-weight storage, placement, replication and eviction. It also includes components for elastic expert-parallel and process-group execution, and provides integrations for engines and systems including vLLM and SGLang. Mooncake is not a standalone model-serving engine: teams attach it beneath an engine when they need cross-node state transfer or a shared cache pool. It is also complementary to LMCache rather than a duplicate; Mooncake's official integration uses Mooncake as the transfer and storage backend while LMCache supplies cache-management and reuse behavior. KTransformers, despite sharing the kvcache-ai organization, is a separate heterogeneous CPU-GPU inference and fine-tuning framework. Mooncake is best suited to platform teams operating multi-node inference, long-context, multi-turn or agentic workloads; deployment complexity and vendor-reported performance gains should be validated on the team's own hardware, topology and traffic.

Pricing & Platform Specs

Pricing Summary

Free and open-source under the Apache-2.0 license. The project does not offer commercial software tiers; organizations provide their own GPU, CPU RAM, NVMe storage, and network infrastructure.

Supported Platforms

C++ infrastructure with Python packages plus Docker and Kubernetes deployment guides. Supports distributed DRAM/SSD/NVMe storage, TCP/RDMA/EFA-class transports, and integrations with vLLM, SGLang, LMCache and other serving stacks.

Explore categories, tags & use cases

Reusable KV cache infrastructure for scalable LLM inference

Open-source KV cache management layer that persists, offloads and reuses model key-value caches across requests and serving engines to reduce repeated prefill work and improve inference throughput.

Open Source

Kubernetes-native distributed LLM inference stack

llm-d is an open-source Kubernetes-native stack for distributed LLM inference with cache-aware routing and disaggregated serving. It separates prefill and decode stages across different GPU pools for optimal resource utilization, routes requests to nodes with warm KV caches, and integrates with vLLM as the serving engine. Apache-2.0 licensed with 2,900+ GitHub stars.

Open Source

Fast serving framework for LLMs and vision models

SGLang is an open-source serving framework for large language and vision-language models, designed for low latency and high throughput. It features RadixAttention for automatic KV cache reuse, compressed finite state machines for fast structured output generation, continuous batching, and tensor parallelism. With over 25,000 GitHub stars, it supports models like LLaMA, Mistral, Qwen, and Gemma on NVIDIA and AMD GPUs.

Open Source

High-throughput LLM serving engine

vLLM is an Apache-2.0 LLM inference and serving engine focused on high-throughput self-hosted model APIs. It combines PagedAttention, continuous batching, prefix caching, quantization options, OpenAI-compatible serving, structured outputs, metrics, Docker/Kubernetes deployment guidance and integrations with agent and LLM frameworks.

Open Source

Cloud-native control plane for scalable GenAI inference

Open-source Kubernetes-native building blocks for deploying, routing and scaling GenAI inference, including an LLM gateway, autoscaling, LoRA management and KV-cache offloading.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is Mooncake?

Open-source infrastructure for disaggregated LLM serving that pools KV caches across prefill and decode workers, with high-performance transfer, distributed storage and integrations for vLLM and SGLang.

Is Mooncake free?

Yes — Mooncake is open source and free to use. Free and open-source under the Apache-2.0 license. The project does not offer commercial software tiers; organizations provide their own GPU, CPU RAM, NVMe storage, and network infrastructure.

Is Mooncake open source?

Yes — Mooncake is open source.

Is Mooncake still maintained?

Yes — Mooncake is active. Its listing was last verified on August 26, 2026.

What are the best Mooncake alternatives?

The first editor-selected Mooncake alternatives are LMCache, llm-d, SGLang, and more.