vLLM Production Stack is the vLLM project's Kubernetes-native reference implementation for operating the existing vLLM inference engine across a cluster. It packages Helm charts and tutorials for moving from one engine instance to distributed deployment without changing the application's OpenAI-compatible interface. The stack adds a request router that can direct traffic by routing keys or sessions to improve KV-cache reuse, optional LMCache offload, service discovery, fault tolerance, autoscaling and an observability layer based on Prometheus and Grafana. Official tutorials cover minimal installations, persistent model weights and deployments on AWS, Google Cloud, Azure and Lambda Labs. This entry is deliberately separate from the live vLLM engine page: vLLM performs model inference, while Production Stack assembles the Kubernetes deployment, routing and operations layer around multiple vLLM instances. It is a reference system rather than a fully managed endpoint, so teams remain responsible for Kubernetes, GPU capacity, storage, networking, model licenses and production hardening. Small single-instance deployments may need only vLLM itself.

vLLM Production Stack
Official Kubernetes and Helm reference stack built on the vLLM inference engine
Official vLLM reference implementation for scaling the existing inference engine on Kubernetes with Helm, request routing, KV-cache offload, autoscaling and Prometheus/Grafana observability.
About vLLM Production Stack
Pricing & Platform Specs
Pricing Summary
Free and open-source under the Apache-2.0 license. Deployers incur standard underlying cloud compute and GPU hardware infrastructure costs for self-hosted Kubernetes clusters.
Supported Platforms
Kubernetes and Helm reference stack with vLLM serving engines, request routing, optional LMCache KV offload, autoscaling, service discovery and Prometheus/Grafana observability. Current verified release: vllm-stack-0.1.12.
Explore categories, tags & use cases
Categories
Alternatives
High-throughput LLM serving engine
vLLM is an Apache-2.0 LLM inference and serving engine focused on high-throughput self-hosted model APIs. It combines PagedAttention, continuous batching, prefix caching, quantization options, OpenAI-compatible serving, structured outputs, metrics, Docker/Kubernetes deployment guidance and integrations with agent and LLM frameworks.
Kubernetes-native model inference platform
KServe is an open-source Kubernetes-native platform for deploying and managing ML model inference at scale. It provides standardized inference protocols, autoscaling including scale-to-zero, canary rollouts, A/B testing, and multi-model serving. KServe supports all major ML frameworks including TensorFlow, PyTorch, scikit-learn, XGBoost, and LLM runtimes like vLLM and Triton through pluggable serving runtimes.
Kubernetes-native distributed LLM inference stack
llm-d is an open-source Kubernetes-native stack for distributed LLM inference with cache-aware routing and disaggregated serving. It separates prefill and decode stages across different GPU pools for optimal resource utilization, routes requests to nodes with warm KV caches, and integrates with vLLM as the serving engine. Apache-2.0 licensed with 2,900+ GitHub stars.
Cloud-native control plane for scalable GenAI inference
Open-source Kubernetes-native building blocks for deploying, routing and scaling GenAI inference, including an LLM gateway, autoscaling, LoRA management and KV-cache offloading.
Kubernetes operator for serving AI inference workloads
KubeAI is an Apache-2.0 Kubernetes operator for deploying and scaling AI inference workloads, including LLMs, embeddings, reranking, and speech-to-text. It gives platform teams OpenAI-compatible endpoints, model proxy/controller primitives, model caching, scale-from-zero behavior, and cluster-native resource management for self-hosted inference on Kubernetes.
Community experience
Sources & verification
- Sources checked
- Content verified
Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.
FAQ
What is vLLM Production Stack?
Official vLLM reference implementation for scaling the existing inference engine on Kubernetes with Helm, request routing, KV-cache offload, autoscaling and Prometheus/Grafana observability.
Is vLLM Production Stack free?
Yes — vLLM Production Stack is open source and free to use. Free and open-source under the Apache-2.0 license. Deployers incur standard underlying cloud compute and GPU hardware infrastructure costs for self-hosted Kubernetes clusters.
Is vLLM Production Stack open source?
Yes — vLLM Production Stack is open source.
Is vLLM Production Stack still maintained?
Yes — vLLM Production Stack is active. Its listing was last verified on August 26, 2026.
What are the best vLLM Production Stack alternatives?
The first editor-selected vLLM Production Stack alternatives are vLLM, KServe, llm-d, and more.