Skip to content
aicoolies logo
ExecuTorch logo

ExecuTorch

PyTorch on-device AI for mobile and edge devices

ExecuTorch is PyTorch's official solution for deploying AI models on mobile, embedded, and edge devices. It features a 50KB base runtime, 12+ hardware backends including Apple CoreML, Qualcomm QNN, ARM, and Vulkan, and native PyTorch export without format conversions. Powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, and Ray-Ban Smart Glasses, supporting LLMs, vision, speech, and multimodal models.

About ExecuTorch

ExecuTorch is PyTorch's unified edge AI runtime, enabling deployment of AI models from smartphones to microcontrollers with a remarkably small 50KB base footprint. Developed collaboratively by Meta, Arm, Apple, and Qualcomm, it reached version 1.0 in late 2025, marking its transition from experimental to production-stable. The key innovation is direct model export from PyTorch — no ONNX, TFLite, or intermediate format conversions needed — preserving model semantics and eliminating the error-prone conversion step that plagues other edge deployment workflows.

The framework supports 12+ hardware backends through a delegate system, including Apple CoreML and Metal for iOS, Qualcomm QNN for Snapdragon, ARM Ethos-U for microcontrollers, Vulkan for cross-platform GPU, and XNNPACK for optimized CPU inference. Built-in quantization via torchao supports 8-bit, 4-bit, and dynamic quantization for reducing model size and inference latency. ExecuTorch already powers billions of on-device inferences at Meta across Instagram, WhatsApp, Quest 3 VR, and Ray-Ban Meta Smart Glasses.

ExecuTorch supports deploying LLMs like Llama 3.2, Qwen 3, and Phi-4-mini, along with vision, speech, and multimodal models on edge devices. It provides developer tools including ETDump profiler, ETRecord inspector, and selective build to strip unused operators. The project is fully open-source under BSD-3-Clause, installable via pip, and integrates with Hugging Face through Optimum-ExecuTorch for transformer model deployment. For mobile developers, it offers native Swift/Objective-C APIs for iOS and Android Studio integration.

Pricing & Platform Specs

Pricing Summary

100% free and open source under the BSD-3-Clause license ($0 software cost). PyTorch ExecuTorch is Meta's official on-device AI inference engine for mobile, embedded systems, and bare-metal microcontrollers with zero licensing fees.

full pricing breakdown →

Supported Platforms

Python/C++ — iOS, Android, Linux, embedded, microcontrollers

Explore categories, tags & use cases

Categories

Google's lightweight ML framework for mobile and embedded

TensorFlow Lite is Google's lightweight ML framework for deploying models on mobile and embedded devices. It supports quantization, GPU/NPU delegation, and runs on Android, iOS, Linux, and microcontrollers. Provides pre-trained models, model conversion tools from TensorFlow and JAX, and hardware acceleration via GPU, Hexagon DSP, and CoreML delegates. Powers on-device ML in billions of Google app installations.

Open Source

Intel's open-source AI inference optimization toolkit

OpenVINO is Intel's open-source toolkit for optimizing and deploying AI inference across CPUs, GPUs, and NPUs. It supports models from PyTorch, TensorFlow, ONNX, and TFLite, providing graph optimizations, quantization, and hardware-specific acceleration. The toolkit includes a GenAI API for LLM deployment and runs on Intel, ARM, and x86 platforms for edge, desktop, and cloud inference workloads.

Open Source

Cross-platform high-performance ML inference engine

ONNX Runtime is Microsoft's open-source inference engine for machine learning models in ONNX format. It delivers cross-platform acceleration via execution providers for NVIDIA CUDA, TensorRT, DirectML, CoreML, OpenVINO, and more. Supports training acceleration, quantization, and GenAI workloads. Used in production across Windows, Azure, Office 365, and thousands of applications with pip-installable Python and native C++/C#/Java APIs.

Open Source

Run LLMs natively on any device with ML compilation

MLC LLM is an open-source engine for deploying large language models natively across diverse platforms using machine learning compilation. It runs models on NVIDIA/AMD GPUs, Apple Silicon, mobile devices, and browsers via WebGPU without cloud dependencies. Features include OpenAI-compatible API, quantization support, and optimized backends for CUDA, Metal, Vulkan, and WebAssembly.

Open Source

First commercially viable 1-bit LLMs that are 14x smaller and 8x faster

PrismML Bonsai delivers the first commercially viable 1-bit large language models with 8B, 4B, and 1.7B parameter variants. The 8B model runs in just 1GB of RAM versus 16GB for standard FP16 models, achieving 44 tokens per second on iPhone. Backed by $16.25M from Khosla Ventures and released under Apache 2.0, Bonsai makes capable LLMs practical for edge devices and resource-constrained environments.

paid

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is ExecuTorch?

ExecuTorch is PyTorch's official solution for deploying AI models on mobile, embedded, and edge devices. It features a 50KB base runtime, 12+ hardware backends including Apple CoreML, Qualcomm QNN, ARM, and Vulkan, and native PyTorch export without format conversions. Powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, and Ray-Ban Smart Glasses, supporting LLMs, vision, speech, and multimodal models.

Is ExecuTorch free?

Yes — ExecuTorch is open source and free to use. 100% free and open source under the BSD-3-Clause license ($0 software cost). PyTorch ExecuTorch is Meta's official on-device AI inference engine for mobile, embedded systems, and bare-metal microcontrollers with zero licensing fees.

Is ExecuTorch open source?

Yes — ExecuTorch is open source.

Is ExecuTorch still maintained?

Yes — ExecuTorch is active. Its listing was last verified on September 6, 2026.

What are the best ExecuTorch alternatives?

The first editor-selected ExecuTorch alternatives are TensorFlow Lite, OpenVINO, ONNX Runtime, and more.