Skip to content
aicoolies logo
LiteRT-LM logo

LiteRT-LM

Google's production on-device LLM inference framework

LiteRT-LM is Google's official open-source framework for running large language models on-device across Android, iOS, Web, Desktop, and Raspberry Pi. Already deployed in Chrome and Pixel hardware, it provides production-grade on-device LLM inference with 1.4K+ GitHub stars. Apache 2.0 licensed.

About LiteRT-LM

LiteRT-LM is Google's production-ready, open-source inference framework for deploying Large Language Models on edge devices. It is the same engine that powers Gemini Nano across Google products including Chrome, Chromebook Plus, Pixel Watch, and Android system features. The framework supports cross-platform deployment on Android, iOS, Web, Desktop, and IoT devices like Raspberry Pi, with hardware acceleration through GPU and NPU accelerators for maximum on-device performance.

Beyond basic text generation, LiteRT-LM provides built-in support for multi-modal inputs including vision and audio, enabling developers to build sophisticated on-device AI applications. The framework also includes function calling and tool use APIs, allowing agentic workflows where on-device models can invoke external tools and services without cloud roundtrips. A streaming API delivers token-by-token output for responsive user experiences, while the underlying C++ interface gives advanced developers full control over custom inference pipelines.

LiteRT-LM works with Gemini Nano models available through Google AI Edge and supports quantized model formats optimized for constrained hardware. The framework handles model loading, memory management, and hardware-specific optimizations automatically, abstracting away the complexity of running LLMs on diverse device configurations. As an open-source project maintained by Google AI Edge, it receives regular updates aligned with Android and Chrome platform releases. For developers building privacy-sensitive, offline-capable, or latency-critical AI features, LiteRT-LM provides a battle-tested path from prototyping to production-scale edge deployment.

Pricing & Platform Specs

Pricing Summary

100% free and open source under the Apache-2.0 license ($0 software cost). LiteRT-LM by Google AI Edge is a high-performance on-device LLM inference engine supporting Android, iOS, WebGPU, and desktop with native NPU/GPU acceleration.

full pricing breakdown →

Supported Platforms

Android, iOS, Web, Desktop, Raspberry Pi

Explore categories, tags & use cases

High-performance local LLM inference in C/C++

llama.cpp is the foundational C/C++ library with 75K+ GitHub stars powering local LLM inference on consumer hardware. Provides optimized CPU and GPU inference for quantized models in GGUF format. Supports LLaMA, Mistral, Phi, Gemma, and most open-weight families. Features 2-8 bit quantization for reduced memory, multi-GPU support, context extension, grammar-constrained output, and an OpenAI-compatible API server. The engine behind Ollama and LM Studio.

Open Source

Cross-platform high-performance ML inference engine

ONNX Runtime is Microsoft's open-source inference engine for machine learning models in ONNX format. It delivers cross-platform acceleration via execution providers for NVIDIA CUDA, TensorRT, DirectML, CoreML, OpenVINO, and more. Supports training acceleration, quantization, and GenAI workloads. Used in production across Windows, Azure, Office 365, and thousands of applications with pip-installable Python and native C++/C#/Java APIs.

Open Source

PyTorch on-device AI for mobile and edge devices

ExecuTorch is PyTorch's official solution for deploying AI models on mobile, embedded, and edge devices. It features a 50KB base runtime, 12+ hardware backends including Apple CoreML, Qualcomm QNN, ARM, and Vulkan, and native PyTorch export without format conversions. Powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, and Ray-Ban Smart Glasses, supporting LLMs, vision, speech, and multimodal models.

Open Source

Google's lightweight ML framework for mobile and embedded

TensorFlow Lite is Google's lightweight ML framework for deploying models on mobile and embedded devices. It supports quantization, GPU/NPU delegation, and runs on Android, iOS, Linux, and microcontrollers. Provides pre-trained models, model conversion tools from TensorFlow and JAX, and hardware acceleration via GPU, Hexagon DSP, and CoreML delegates. Powers on-device ML in billions of Google app installations.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is LiteRT-LM?

LiteRT-LM is Google's official open-source framework for running large language models on-device across Android, iOS, Web, Desktop, and Raspberry Pi. Already deployed in Chrome and Pixel hardware, it provides production-grade on-device LLM inference with 1.4K+ GitHub stars. Apache 2.0 licensed.

Is LiteRT-LM free?

Yes — LiteRT-LM is open source and free to use. 100% free and open source under the Apache-2.0 license ($0 software cost). LiteRT-LM by Google AI Edge is a high-performance on-device LLM inference engine supporting Android, iOS, WebGPU, and desktop with native NPU/GPU acceleration.

Is LiteRT-LM open source?

Yes — LiteRT-LM is open source.

Is LiteRT-LM still maintained?

Yes — LiteRT-LM is active. Its listing was last verified on September 6, 2026.

What are the best LiteRT-LM alternatives?

The first editor-selected LiteRT-LM alternatives are llama.cpp, ONNX Runtime, ExecuTorch, and more.