Skip to content
aicoolies logo
llm-d logo

llm-d

Kubernetes-native distributed LLM inference stack

llm-d is an open-source Kubernetes-native stack for distributed LLM inference with cache-aware routing and disaggregated serving. It separates prefill and decode stages across different GPU pools for optimal resource utilization, routes requests to nodes with warm KV caches, and integrates with vLLM as the serving engine. Apache-2.0 licensed with 2,900+ GitHub stars.

About llm-d

llm-d addresses the operational complexity of running large language model inference at scale on Kubernetes. While individual serving engines like vLLM handle the mechanics of running models on GPUs, production deployments require an orchestration layer that manages routing, scheduling, scaling, and resource allocation across a fleet of GPU nodes. llm-d provides this orchestration through a Kubernetes-native architecture that uses custom resources and operators to declare inference topologies, with intelligent routing that considers KV cache state, GPU memory availability, and request characteristics when assigning work to nodes.

The disaggregated serving architecture separates the prefill stage (processing the input prompt) from the decode stage (generating output tokens) across different GPU pools. This separation enables significant efficiency gains because prefill is compute-intensive and benefits from high-bandwidth GPUs, while decode is memory-bandwidth-limited and can run on different hardware configurations. The cache-aware routing system tracks which prompts have been processed on which nodes, directing subsequent requests to nodes that already have relevant KV cache entries warm in GPU memory, avoiding redundant computation for conversations and repeated system prompts.

llm-d builds on vLLM as its serving engine while adding the cluster-level intelligence that transforms individual GPU servers into a coordinated inference platform. The project integrates with Kubernetes' native scaling mechanisms for automatic GPU allocation based on request volume, and supports mixed hardware configurations where different model sizes and quantization levels are served across heterogeneous GPU pools. With 2,900+ GitHub stars and an Apache-2.0 license, llm-d targets AI platform teams that need production-grade inference infrastructure beyond what a single vLLM instance provides.

Pricing & Platform Specs

Pricing Summary

Free and 100% open source under the Apache-2.0 license as a CNCF Sandbox project. llm-d delivers Kubernetes-native distributed LLM inference orchestration, prefill/decode disaggregation, prefix-cache aware routing, and multi-tiered KV-cache offloading on top of vLLM and SGLang with zero software licensing fees.

full pricing breakdown →

Supported Platforms

Kubernetes — Helm charts, requires GPU nodes with vLLM

Explore categories, tags & use cases

High-throughput LLM serving engine

vLLM is an Apache-2.0 LLM inference and serving engine focused on high-throughput self-hosted model APIs. It combines PagedAttention, continuous batching, prefix caching, quantization options, OpenAI-compatible serving, structured outputs, metrics, Docker/Kubernetes deployment guidance and integrations with agent and LLM frameworks.

Open Source

Fast serving framework for LLMs and vision models

SGLang is an open-source serving framework for large language and vision-language models, designed for low latency and high throughput. It features RadixAttention for automatic KV cache reuse, compressed finite state machines for fast structured output generation, continuous batching, and tensor parallelism. With over 25,000 GitHub stars, it supports models like LLaMA, Mistral, Qwen, and Gemma on NVIDIA and AMD GPUs.

Open Source

Run LLMs locally with one command

Tool for running large language models locally on your machine with a simple CLI interface. Download and run Llama 3, Mistral, Gemma, Phi, Code Llama, and dozens of other open-source models with a single command. Features model management, GPU acceleration (NVIDIA/AMD/Apple Silicon), OpenAI-compatible API server, Modelfile for customization, and multi-model switching. Ideal for offline AI development, privacy-sensitive use cases, and local testing. 120K+ GitHub stars.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is llm-d?

llm-d is an open-source Kubernetes-native stack for distributed LLM inference with cache-aware routing and disaggregated serving. It separates prefill and decode stages across different GPU pools for optimal resource utilization, routes requests to nodes with warm KV caches, and integrates with vLLM as the serving engine. Apache-2.0 licensed with 2,900+ GitHub stars.

Is llm-d free?

Yes — llm-d is open source and free to use. Free and 100% open source under the Apache-2.0 license as a CNCF Sandbox project. llm-d delivers Kubernetes-native distributed LLM inference orchestration, prefill/decode disaggregation, prefix-cache aware routing, and multi-tiered KV-cache offloading on top of vLLM and SGLang with zero software licensing fees.

Is llm-d open source?

Yes — llm-d is open source.

Is llm-d still maintained?

Yes — llm-d is active. Its listing was last verified on September 6, 2026.

What are the best llm-d alternatives?

The first editor-selected llm-d alternatives are vLLM, SGLang, Ollama.