Skip to content
aicoolies logo
GPUStack logo

GPUStack

Open-source GPU control plane for scalable AI model serving

Open-source GPU cluster manager that configures vLLM, SGLang, TensorRT-LLM or custom engines, serves models through compatible APIs, and provisions SSH-accessible GPU instances across on-premises, Kubernetes and cloud environments.

About GPUStack

GPUStack is an open-source control plane for AI model serving and GPU instance provisioning. It manages GPU clusters across on-premises, Kubernetes and cloud environments, schedules accelerator capacity, and configures pluggable inference engines such as vLLM, SGLang and TensorRT-LLM rather than replacing those engines. The platform exposes compatible model APIs and adds authentication, access control, load balancing, monitoring, token and request metering, automated recovery, performance-oriented engine settings, and on-demand SSH-accessible GPU instances. Its current documentation lists support paths for NVIDIA, AMD, Ascend, Hygon, MThreads, Iluvatar, MetaX, Cambricon and T-Head accelerators, while worker nodes remain Linux-only. Maintainer claims about Day-0 model support and performance gains depend on the selected engine, model and hardware, so operators should validate compatibility and throughput against their own fleet. GPUStack is best suited to platform teams running multiple models, clusters or tenants; a single local model or a team seeking a fully managed serverless endpoint will usually need less operational machinery.

Pricing & Platform Specs

Pricing Summary

Free and open source under Apache-2.0. Cluster operators provide their own bare-metal or cloud GPU infrastructure.

Supported Platforms

Docker-deployable control plane with Linux GPU workers, multi-cluster scheduling, pluggable vLLM/SGLang/TensorRT-LLM engines, compatible model APIs, monitoring, metering and SSH-accessible GPU instances. Current verified release: v2.2.3.

Explore categories, tags & use cases

Cloud-native control plane for scalable GenAI inference

Open-source Kubernetes-native building blocks for deploying, routing and scaling GenAI inference, including an LLM gateway, autoscaling, LoRA management and KV-cache offloading.

Open Source

Kubernetes operator for serving AI inference workloads

KubeAI is an Apache-2.0 Kubernetes operator for deploying and scaling AI inference workloads, including LLMs, embeddings, reranking, and speech-to-text. It gives platform teams OpenAI-compatible endpoints, model proxy/controller primitives, model caching, scale-from-zero behavior, and cluster-native resource management for self-hosted inference on Kubernetes.

Open Source

Open-source control plane for AI workloads across multi-cloud GPU infrastructure

dstack is an open-source platform that orchestrates AI training and inference workloads across heterogeneous GPU infrastructure spanning multiple clouds, Kubernetes clusters, and bare-metal servers. It abstracts away cloud-specific APIs so teams define GPU requirements declaratively and dstack automatically provisions the cheapest available resources from AWS, GCP, Azure, Lambda, or on-premises hardware.

Open Source

Run AI workloads on any cloud with automatic cost optimization

SkyPilot is an open-source framework for running LLMs, AI, and batch jobs on any cloud with automatic cost optimization. It supports AWS, GCP, Azure, Lambda Cloud, and more, automatically selecting the cheapest available GPUs and managing spot instance preemption. Features include multi-cloud job scheduling, managed spot jobs with automatic recovery, and cluster autoscaling with 6,000+ GitHub stars.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is GPUStack?

Open-source GPU cluster manager that configures vLLM, SGLang, TensorRT-LLM or custom engines, serves models through compatible APIs, and provisions SSH-accessible GPU instances across on-premises, Kubernetes and cloud environments.

Is GPUStack free?

Yes — GPUStack is open source and free to use. Free and open source under Apache-2.0. Cluster operators provide their own bare-metal or cloud GPU infrastructure.

Is GPUStack open source?

Yes — GPUStack is open source.

Is GPUStack still maintained?

Yes — GPUStack is active. Its listing was last verified on August 26, 2026.

What are the best GPUStack alternatives?

The first editor-selected GPUStack alternatives are AIBrix, KubeAI, Dstack, and more.