Skip to content
aicoolies logo
MLC LLM logo

MLC LLM

Run LLMs natively on any device with ML compilation

MLC LLM is an open-source engine for deploying large language models natively across diverse platforms using machine learning compilation. It runs models on NVIDIA/AMD GPUs, Apple Silicon, mobile devices, and browsers via WebGPU without cloud dependencies. Features include OpenAI-compatible API, quantization support, and optimized backends for CUDA, Metal, Vulkan, and WebAssembly.

About MLC LLM

MLC LLM uses machine learning compilation technology to deploy large language models natively on virtually any hardware platform. Rather than relying on framework-specific runtimes, it compiles models into optimized native code for the target platform using Apache TVM's compiler infrastructure. This approach enables running LLMs on NVIDIA GPUs via CUDA, AMD GPUs via ROCm/Vulkan, Apple Silicon via Metal, mobile devices via Android NDK and iOS, and even web browsers via WebGPU — all from the same model definition.

The project provides pre-compiled model libraries for popular architectures including Llama, Mistral, Gemma, Phi, and Qwen, along with tools for compiling custom models. It offers an OpenAI-compatible REST API server for drop-in replacement in existing applications, chat CLI for interactive use, and Python/JavaScript/Swift APIs for embedding in applications. Quantization support includes group quantization and mixed-precision modes to reduce memory requirements while maintaining generation quality.

MLC LLM is open-source under Apache 2.0, developed by the MLC AI community with roots in CMU research. It distinguishes itself from tools like llama.cpp by using compiler-based optimization rather than hand-tuned kernels, which enables automatic optimization for new hardware targets. The project maintains active development with regular model updates and platform support improvements, making it a strong choice for developers who need to deploy LLMs across heterogeneous hardware without maintaining separate deployment paths for each platform.

Pricing & Platform Specs

Pricing Summary

100% free and open source under the Apache-2.0 license ($0 software licensing costs). MLC LLM is a universal machine learning compiler and runtime built on Apache TVM Unity that compiles and runs large language models natively across server GPUs, Apple Silicon, mobile devices (iOS/Android), and web browsers via WebGPU with zero commercial software fees.

full pricing breakdown →

Supported Platforms

Python/CLI — GPU, CPU, mobile, browser via WebGPU

Explore categories, tags & use cases

Categories

PyTorch on-device AI for mobile and edge devices

ExecuTorch is PyTorch's official solution for deploying AI models on mobile, embedded, and edge devices. It features a 50KB base runtime, 12+ hardware backends including Apple CoreML, Qualcomm QNN, ARM, and Vulkan, and native PyTorch export without format conversions. Powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, and Ray-Ban Smart Glasses, supporting LLMs, vision, speech, and multimodal models.

Open Source

Google's lightweight ML framework for mobile and embedded

TensorFlow Lite is Google's lightweight ML framework for deploying models on mobile and embedded devices. It supports quantization, GPU/NPU delegation, and runs on Android, iOS, Linux, and microcontrollers. Provides pre-trained models, model conversion tools from TensorFlow and JAX, and hardware acceleration via GPU, Hexagon DSP, and CoreML delegates. Powers on-device ML in billions of Google app installations.

Open Source

Intel's open-source AI inference optimization toolkit

OpenVINO is Intel's open-source toolkit for optimizing and deploying AI inference across CPUs, GPUs, and NPUs. It supports models from PyTorch, TensorFlow, ONNX, and TFLite, providing graph optimizations, quantization, and hardware-specific acceleration. The toolkit includes a GenAI API for LLM deployment and runs on Intel, ARM, and x86 platforms for edge, desktop, and cloud inference workloads.

Open Source

Cross-platform high-performance ML inference engine

ONNX Runtime is Microsoft's open-source inference engine for machine learning models in ONNX format. It delivers cross-platform acceleration via execution providers for NVIDIA CUDA, TensorRT, DirectML, CoreML, OpenVINO, and more. Supports training acceleration, quantization, and GenAI workloads. Used in production across Windows, Azure, Office 365, and thousands of applications with pip-installable Python and native C++/C#/Java APIs.

Open Source

Run LLMs locally with one command

Tool for running large language models locally on your machine with a simple CLI interface. Download and run Llama 3, Mistral, Gemma, Phi, Code Llama, and dozens of other open-source models with a single command. Features model management, GPU acceleration (NVIDIA/AMD/Apple Silicon), OpenAI-compatible API server, Modelfile for customization, and multi-model switching. Ideal for offline AI development, privacy-sensitive use cases, and local testing. 120K+ GitHub stars.

Open Source

High-performance local LLM inference in C/C++

llama.cpp is the foundational C/C++ library with 75K+ GitHub stars powering local LLM inference on consumer hardware. Provides optimized CPU and GPU inference for quantized models in GGUF format. Supports LLaMA, Mistral, Phi, Gemma, and most open-weight families. Features 2-8 bit quantization for reduced memory, multi-GPU support, context extension, grammar-constrained output, and an OpenAI-compatible API server. The engine behind Ollama and LM Studio.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is MLC LLM?

MLC LLM is an open-source engine for deploying large language models natively across diverse platforms using machine learning compilation. It runs models on NVIDIA/AMD GPUs, Apple Silicon, mobile devices, and browsers via WebGPU without cloud dependencies. Features include OpenAI-compatible API, quantization support, and optimized backends for CUDA, Metal, Vulkan, and WebAssembly.

Is MLC LLM free?

Yes — MLC LLM is open source and free to use. 100% free and open source under the Apache-2.0 license ($0 software licensing costs). MLC LLM is a universal machine learning compiler and runtime built on Apache TVM Unity that compiles and runs large language models natively across server GPUs, Apple Silicon, mobile devices (iOS/Android), and web browsers via WebGPU with zero commercial software fees.

Is MLC LLM open source?

Yes — MLC LLM is open source.

Is MLC LLM still maintained?

Yes — MLC LLM is active. Its listing was last verified on September 6, 2026.

What are the best MLC LLM alternatives?

The first editor-selected MLC LLM alternatives are ExecuTorch, TensorFlow Lite, OpenVINO, and more.