aicoolies logo
LiteRT-LM logo

Best LiteRT-LM Alternatives

4 editor-verified alternatives · LiteRT-LM overview →

source: tools.alternatives · stored order · active records only; review scores are annotations and never change membership or order

llama.cpp logo
1

llama.cpp

open sourceexplicit relation

llama.cpp is the foundational C/C++ library with 75K+ GitHub stars powering local LLM inference on consumer hardware. Provides optimized CPU and GPU inference for quantized models in GGUF format. Supports LLaMA, Mistral, Phi, Gemma, and most open-weight families. Features 2-8 bit quantization for reduced memory, multi-GPU support, context extension, grammar-constrained output, and an OpenAI-compatible API server. The engine behind Ollama and LM Studio.

Free and open-source
ONNX Runtime logo
2

ONNX Runtime

open sourceexplicit relation

ONNX Runtime is Microsoft's open-source inference engine for machine learning models in ONNX format. It delivers cross-platform acceleration via execution providers for NVIDIA CUDA, TensorRT, DirectML, CoreML, OpenVINO, and more. Supports training acceleration, quantization, and GenAI workloads. Used in production across Windows, Azure, Office 365, and thousands of applications with pip-installable Python and native C++/C#/Java APIs.

Free and open-source (MIT license)
ExecuTorch logo
3

ExecuTorch

open sourceexplicit relation

ExecuTorch is PyTorch's official solution for deploying AI models on mobile, embedded, and edge devices. It features a 50KB base runtime, 12+ hardware backends including Apple CoreML, Qualcomm QNN, ARM, and Vulkan, and native PyTorch export without format conversions. Powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, and Ray-Ban Smart Glasses, supporting LLMs, vision, speech, and multimodal models.

Free and open-source (BSD-3-Clause)
TensorFlow Lite logo
4

TensorFlow Lite

open sourceexplicit relation

TensorFlow Lite is Google's lightweight ML framework for deploying models on mobile and embedded devices. It supports quantization, GPU/NPU delegation, and runs on Android, iOS, Linux, and microcontrollers. Provides pre-trained models, model conversion tools from TensorFlow and JAX, and hardware acceleration via GPU, Hexagon DSP, and CoreML delegates. Powers on-device ML in billions of Google app installations.

Free and open-source (Apache 2.0)

Open-source LiteRT-LM alternatives

llama.cpp, ONNX Runtime, ExecuTorch, TensorFlow Litesee all open-source developer tools.

FAQ

What is the best LiteRT-LM alternative?

llama.cpp tops our editor-verified list of 4 LiteRT-LM alternatives.

Are there open-source LiteRT-LM alternatives?

Yes — llama.cpp, ONNX Runtime, ExecuTorch, and more are open source.

4 Best LiteRT-LM Alternatives (2026) — aicoolies