Explore / Category guide
Model Providers
Discover the top Model Providers in 2026. Compare architecture, pricing tiers, performance benchmarks, and open-source developer alternatives.
Category overviewAbout Model ProvidersRead guideClose guide
Where do the weights actually run? That single question sorts this page, and it is worth answering before you compare anything else, because it determines your cost curve, your latency floor and your data-residency story all at once.
The catalogue reflects the split almost exactly: 31 of the 61 catalogued entries (50.8%) are open source. Roughly half of this shelf is a vendor you send tokens to; the other half is software you install. Few categories divide that evenly, and the even division is not an accident — it is the state of the market in 2026.
On the rented side, published per-token pricing is now precise enough to model before you commit. The Anthropic API (90/100, Anthropic API) lists Haiku 4.5 at $1/$5, Sonnet 4.6 at $3/$15 and Opus 4.7 at $5/$25 per million input/output tokens as of 16 August 2026 — a fivefold spread between the cheapest and most expensive model from a single vendor, which usually moves a bill more than switching vendors does. On the consumer side Claude scores 93/100 (Claude, verified 17 August 2026) and ChatGPT 92/100 (ChatGPT); both appear in 12 of the 589 published comparisons, so the head-to-head questions are already answered on the comparison pages rather than here. Gemini (86/100) lists a free tier alongside AI Pro at $19.99/mo and AI Ultra at $249.99/mo. Note that both Gemini and DeepSeek (90/100) carry a telemetry-concerns tag in our catalogue — a flag on data handling, not on quality.
On the owned side, the choice is throughput versus convenience. vLLM (91/100, free and open source) is the serving layer for production GPU workloads where utilisation is the constraint; our review is explicit that it is infrastructure, not an application framework. Ollama (88/100) is the opposite trade — free, in 11 published stacks and 11 comparisons, and it exposes an OpenAI-compatible endpoint so most client code does not change. LM Studio (84/100) is the same idea with a desktop GUI instead of a terminal. Together AI (89/100) sits between the two camps: someone else's H100s at $6.49/hr, running open weights you chose.
One historical entry to ignore as an option: Google Bard was discontinued in February 2024 and is kept only as a graveyard record. It is already excluded from the 60 tools this page counts.
20 of the 61 entries carry a scored review. Start with the shortlist, then read the pricing summary on the tool page — provider pricing in this category moves faster than anything else on the site.

showing 15 of 63 tools
NVIDIA's real-time persona-driven voice dialogue model
PersonaPlex is NVIDIA's open-source, full-duplex speech-to-speech conversational AI model that enables persona control through text-based role prompts and audio-based voice conditioning. Built on the Moshi architecture, it produces natural, low-latency spoken interactions with consistent persona across conversations. The model supports multiple pre-packaged voice embeddings for both natural and varied speaking styles, making it suitable for building interactive voice agents and assistants.
Container-native local AI model serving with Podman
RamaLama is an open-source tool that containerizes AI model inference using Podman or Docker, eliminating host system configuration complexity. It auto-detects GPUs (NVIDIA, AMD, Intel, Apple Silicon), pulls models from HuggingFace, Ollama, and OCI registries, and runs them in isolated rootless containers with read-only mounts and network isolation. Developed under the Containers project (Red Hat ecosystem), it brings familiar container workflows to local LLM serving.
Edge AI deployment SDK for heterogeneous SoCs
Roofline AI is a contact-sales edge AI deployment toolkit built around an MLIR- and IREE-based compiler. Its SDK compiles models ahead of time, a lightweight C runtime executes them across CPUs, GPUs, and NPUs, and a performance dashboard tracks latency, throughput, memory use, and model coverage across devices.
Fast serving framework for LLMs and vision models
SGLang is an open-source serving framework for large language and vision-language models, designed for low latency and high throughput. It features RadixAttention for automatic KV cache reuse, compressed finite state machines for fast structured output generation, continuous batching, and tensor parallelism. With over 25,000 GitHub stars, it supports models like LLaMA, Mistral, Qwen, and Gemma on NVIDIA and AMD GPUs.
Multi-agent model API that orchestrates frontier models behind one OpenAI-compatible endpoint
Sakana Fugu is a hosted model-provider API that exposes a learned multi-agent system as one OpenAI-compatible model. It dynamically routes coding, code review, research, and reasoning tasks across a frontier-model pool, with Fugu for lower-latency work and Fugu Ultra for harder workloads where answer quality matters more than cost or speed.
NVIDIA's LLM inference optimization and acceleration library
TensorRT-LLM is NVIDIA's open-source library for optimizing LLM inference on NVIDIA GPUs. It provides kernel fusion, quantization (FP8, INT4, INT8), KV cache optimization, and in-flight batching to maximize throughput. Supports multi-GPU and multi-node setups with tensor and pipeline parallelism, and integrates with Triton Inference Server for production deployment of models like LLaMA, GPT, Mistral, and Qwen.
Hugging Face's open-source inference server for embeddings, rerankers, and classifiers
Text Embeddings Inference is Hugging Face's Apache-2.0 server for high-throughput embedding, reranking, and sequence-classification models. TEI packages token-based dynamic batching, optimized Transformers kernels, Safetensors loading, OpenAI-compatible embedding endpoints, Prometheus metrics, and configurable OpenTelemetry tracing in deployable CPU and GPU images.
Hugging Face's production LLM serving framework
Text Generation Inference (TGI) is Hugging Face's production-ready serving framework for large language models. It features flash attention, continuous batching, tensor parallelism, quantization via GPTQ/AWQ/EETQ, and Safetensors support. Powers Hugging Face's Inference API and Inference Endpoints, with an OpenAI-compatible API and Docker deployment. Supports LLaMA, Mistral, Falcon, and other popular model architectures.
NVIDIA's optimized AI model serving platform
Triton Inference Server is NVIDIA's open-source inference serving platform that deploys AI models from TensorRT, PyTorch, ONNX, TensorFlow, OpenVINO, Python, and more across cloud, data center, and edge environments. It supports dynamic batching, model ensembles, concurrent model execution on GPUs and CPUs, and real-time, streaming, and batch inference patterns. Includes Model Analyzer for profiling and Model Navigator for automated optimization.
Local model inference engine with OpenAI-compatible API and web UI
Xinference is a local inference engine that runs LLMs, embedding models, image generation, and audio models with an OpenAI-compatible API. It provides a web dashboard for model management, supports vLLM, llama.cpp, and transformers backends, and handles multi-GPU deployment automatically. Supports 100+ models including Qwen, Llama, Mistral, and DeepSeek with over 9,200 GitHub stars.
GLM Coding Plan by Z.AI — subscription access to GLM-5.3 for agents and IDEs
Z.AI’s GLM Coding Plan is a fixed-quota coding subscription for AI IDEs and coding agents, powered by Zhipu’s GLM-5.3 and GLM-5.3-Flash (not a Kimi plan).
Serverless AI inference for generative media at scale
fal.ai is a serverless AI inference platform providing ultra-low-latency APIs for generating images, videos, audio, and 3D models. With 600+ production-ready models and native Python and JavaScript SDKs, it eliminates GPU management while delivering 30-50% lower costs than alternatives. Automatic scaling with no cold starts and real-time streaming support make it ideal for interactive AI applications.
High-performance local LLM inference in C/C++
llama.cpp is the foundational C/C++ library with 75K+ GitHub stars powering local LLM inference on consumer hardware. Provides optimized CPU and GPU inference for quantized models in GGUF format. Supports LLaMA, Mistral, Phi, Gemma, and most open-weight families. Features 2-8 bit quantization for reduced memory, multi-GPU support, context extension, grammar-constrained output, and an OpenAI-compatible API server. The engine behind Ollama and LM Studio.
Kubernetes-native distributed LLM inference stack
llm-d is an open-source Kubernetes-native stack for distributed LLM inference with cache-aware routing and disaggregated serving. It separates prefill and decode stages across different GPU pools for optimal resource utilization, routes requests to nodes with warm KV caches, and integrates with vLLM as the serving engine. Apache-2.0 licensed with 2,900+ GitHub stars.
Official Python SDK for the xAI API
The xAI Python SDK is the official Python client for the xAI API, giving developers a direct way to build Grok-powered apps without relying on community proxies or unofficial wrappers. It supports synchronous and asynchronous Python clients for chat completions, streaming responses, function/tool calling, and multimodal workflows, making it a clean fit for backend services, agents, notebooks, and developer tools that need programmatic xAI access.