Skip to content
aicoolies logo
llama.cpp logo

llama.cpp

High-performance local LLM inference in C/C++

llama.cpp is the foundational C/C++ library with 75K+ GitHub stars powering local LLM inference on consumer hardware. Provides optimized CPU and GPU inference for quantized models in GGUF format. Supports LLaMA, Mistral, Phi, Gemma, and most open-weight families. Features 2-8 bit quantization for reduced memory, multi-GPU support, context extension, grammar-constrained output, and an OpenAI-compatible API server. The engine behind Ollama and LM Studio.

About llama.cpp

llama.cpp is the foundational library for running LLMs on consumer hardware. With 75K+ stars, it powers Ollama, LM Studio, and many local AI applications.

Optimized for CPU (AVX/AVX2/AVX-512), Apple Silicon (Metal), NVIDIA (CUDA), and AMD (ROCm). GGUF format quantization from 2-bit to 8-bit reduces memory while maintaining quality.

Supports LLaMA, Mistral, Phi, Gemma, Qwen, and virtually all open-weight models. Features context extension, grammar-constrained output, batch processing, and speculative decoding.

Built-in HTTP server provides OpenAI-compatible API for seamless local inference. Continuously optimized by a large community.

Pricing & Platform Specs

Pricing Summary

Free and 100% open-source LLM inference engine under the MIT license with zero software licensing fees or subscription costs. Executes locally and privately across Apple Silicon Metal, NVIDIA CUDA, AMD ROCm, Vulkan, and CPU SIMD hardware with zero cloud dependencies, including an OpenAI-compatible REST server (llama-server) for zero-cost self-hosted deployments.

full pricing breakdown →

Supported Platforms

CPU, CUDA, Metal, ROCm, any OS

Explore categories, tags & use cases

Categories

Cross-platform on-device AI inference SDK

RunAnywhere SDK is a production-ready toolkit for running AI models entirely on-device across iOS, macOS, Android, Web, React Native, and Flutter. It provides a unified C++ core with platform-specific bindings for LLM text generation via llama.cpp, vision-language models, Whisper speech-to-text, Piper text-to-speech, and on-device image generation. All processing stays local with zero cloud dependency, ensuring privacy and low latency for mobile and edge AI applications.

freemiumOpen Source

Cross-platform on-device AI model runtime

Nexa SDK enables running frontier LLMs and multimodal models locally across PC, mobile, IoT, and wearables with automatic hardware acceleration for GPU, NPU, and CPU. It supports Qwen, Gemma, Llama, DeepSeek models with Python/C++ desktop SDKs, Android/iOS mobile SDKs, and Docker for edge deployment. Includes an OpenAI-compatible API server with chat and function calling support.

Open Source

Side-by-Side Comparisons

Ollama logo
Ollama
vs
llama.cpp logo
llama.cpp

Ollama vs llama.cpp — Local LLM Wrapper vs the Inference Engine It Wraps

Ollama and llama.cpp both let you run open-weight models on your own hardware, but they sit at different layers of the stack. llama.cpp is the C/C++ inference engine that started the local-LLM movement and quietly powers a huge slice of the ecosystem. Ollama is the Go-based developer wrapper that hides the rough edges and turned local models into a one-line install for everyone else.

Ollamallama.cpp

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is llama.cpp?

llama.cpp is the foundational C/C++ library with 75K+ GitHub stars powering local LLM inference on consumer hardware. Provides optimized CPU and GPU inference for quantized models in GGUF format. Supports LLaMA, Mistral, Phi, Gemma, and most open-weight families. Features 2-8 bit quantization for reduced memory, multi-GPU support, context extension, grammar-constrained output, and an OpenAI-compatible API server. The engine behind Ollama and LM Studio.

Is llama.cpp free?

Yes — llama.cpp is open source and free to use. Free and 100% open-source LLM inference engine under the MIT license with zero software licensing fees or subscription costs. Executes locally and privately across Apple Silicon Metal, NVIDIA CUDA, AMD ROCm, Vulkan, and CPU SIMD hardware with zero cloud dependencies, including an OpenAI-compatible REST server (llama-server) for zero-cost self-hosted deployments.

Is llama.cpp open source?

Yes — llama.cpp is open source.

Is llama.cpp still maintained?

Yes — llama.cpp is active. Its listing was last verified on September 6, 2026.

What are the best llama.cpp alternatives?

The first editor-selected llama.cpp alternatives are RunAnywhere SDK, Nexa SDK.