Skip to content
aicoolies logo
LMCache logo

LMCache

Reusable KV cache infrastructure for scalable LLM inference

Open-source KV cache management layer that persists, offloads and reuses model key-value caches across requests and serving engines to reduce repeated prefill work and improve inference throughput.

About LMCache

LMCache is an open-source KV cache management layer for large-model inference. It turns the temporary key-value state produced during prefill into reusable infrastructure that can persist across requests, sessions and serving-engine instances. The current project supports tiered offloading to CPU memory, local storage and remote backends, cross-engine cache reuse, cache observability, pluggable storage and transport, and non-prefix reuse. The maintainers describe LMCache as engine-independent and vendor-neutral rather than as the vLLM engine itself; integrations span serving engines, hardware vendors and storage systems. Vendor-reported performance gains depend on the model, hardware and amount of shared context, so buyers should benchmark their own long-context, multi-turn and RAG workloads. LMCache is best suited to platform teams operating repeated-prompt or shared-context inference at enough scale to justify a separate cache layer.

Pricing & Platform Specs

Pricing Summary

Free and open-source under the Apache-2.0 license. Deployers manage their own infrastructure and compute resources when integrating LMCache into existing vLLM or SGLang clusters.

Supported Platforms

Installable with pip and deployable as an engine-independent cache daemon or integrated backend. Supports tiered CPU, local and remote storage, multiple serving engines and production observability.

Explore categories, tags & use cases

High-throughput LLM serving engine

vLLM is an Apache-2.0 LLM inference and serving engine focused on high-throughput self-hosted model APIs. It combines PagedAttention, continuous batching, prefix caching, quantization options, OpenAI-compatible serving, structured outputs, metrics, Docker/Kubernetes deployment guidance and integrations with agent and LLM frameworks.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is LMCache?

Open-source KV cache management layer that persists, offloads and reuses model key-value caches across requests and serving engines to reduce repeated prefill work and improve inference throughput.

Is LMCache free?

Yes — LMCache is open source and free to use. Free and open-source under the Apache-2.0 license. Deployers manage their own infrastructure and compute resources when integrating LMCache into existing vLLM or SGLang clusters.

Is LMCache open source?

Yes — LMCache is open source.

Is LMCache still maintained?

Yes — LMCache is active. Its listing was last verified on August 26, 2026.

What are the best LMCache alternatives?

The first editor-selected LMCache alternatives are vLLM.