Skip to content
aicoolies logo
Google Research logo

gemma.cpp

Lightweight C++ inference for Google Gemma models

gemma.cpp is Google's standalone C++ inference engine built specifically for running Gemma language models without Python or CUDA dependencies. It provides optimized CPU inference using SIMD instructions and Highway library, supports Gemma 2 and Gemma 3 models, and runs on x86 and ARM architectures. Designed for embedded systems, edge devices, and server deployments needing minimal overhead.

About gemma.cpp

gemma.cpp is Google DeepMind's purpose-built inference engine that strips away the overhead of Python runtimes and heavy ML frameworks to run Gemma models with maximum efficiency on CPU hardware. Unlike general-purpose inference engines like llama.cpp that support many model architectures, gemma.cpp is optimized specifically for the Gemma model family, enabling architecture-specific optimizations that would not be possible in a generic framework. The result is faster inference with lower memory usage for Gemma-specific deployments.

The engine leverages Google's Highway library for portable SIMD operations, automatically selecting the best instruction set available on the target CPU — AVX-512, AVX2, SSE4, or NEON for ARM. This makes it suitable for deployment across x86 servers, Apple Silicon Macs, Raspberry Pi devices, and other ARM hardware without code changes. gemma.cpp supports the complete Gemma model family including Gemma 2 and Gemma 3 variants, with quantized model formats that reduce memory requirements while preserving quality.

With over 6,800 GitHub stars and Google's direct maintenance, gemma.cpp serves developers who need to deploy Gemma models in environments where Python is unavailable, undesirable, or too slow. Use cases include embedded systems, mobile applications via native code, IoT edge devices, and high-throughput server deployments. The Apache-2.0 license and Google's active development ensure the engine stays current with new Gemma model releases and architectural improvements.

Pricing & Platform Specs

Pricing Summary

100% free and open source under the Apache-2.0 license ($0 software cost). Minimal, zero-dependency C++ inference engine with Google Highway portable SIMD acceleration for Google Gemma and Gemma 2 models.

full pricing breakdown →

Supported Platforms

x86 and ARM CPUs — no Python or CUDA required

Explore categories, tags & use cases

Categories

High-performance local LLM inference in C/C++

llama.cpp is the foundational C/C++ library with 75K+ GitHub stars powering local LLM inference on consumer hardware. Provides optimized CPU and GPU inference for quantized models in GGUF format. Supports LLaMA, Mistral, Phi, Gemma, and most open-weight families. Features 2-8 bit quantization for reduced memory, multi-GPU support, context extension, grammar-constrained output, and an OpenAI-compatible API server. The engine behind Ollama and LM Studio.

Open Source

Run LLMs locally with one command

Tool for running large language models locally on your machine with a simple CLI interface. Download and run Llama 3, Mistral, Gemma, Phi, Code Llama, and dozens of other open-source models with a single command. Features model management, GPU acceleration (NVIDIA/AMD/Apple Silicon), OpenAI-compatible API server, Modelfile for customization, and multi-model switching. Ideal for offline AI development, privacy-sensitive use cases, and local testing. 120K+ GitHub stars.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is gemma.cpp?

gemma.cpp is Google's standalone C++ inference engine built specifically for running Gemma language models without Python or CUDA dependencies. It provides optimized CPU inference using SIMD instructions and Highway library, supports Gemma 2 and Gemma 3 models, and runs on x86 and ARM architectures. Designed for embedded systems, edge devices, and server deployments needing minimal overhead.

Is gemma.cpp free?

Yes — gemma.cpp is open source and free to use. 100% free and open source under the Apache-2.0 license ($0 software cost). Minimal, zero-dependency C++ inference engine with Google Highway portable SIMD acceleration for Google Gemma and Gemma 2 models.

Is gemma.cpp open source?

Yes — gemma.cpp is open source.

Is gemma.cpp still maintained?

Yes — gemma.cpp is active. Its listing was last verified on September 6, 2026.

What are the best gemma.cpp alternatives?

The first editor-selected gemma.cpp alternatives are llama.cpp, Ollama.