Skip to content
aicoolies logo
MLX logo

MLX-VLM

Run and fine-tune Vision Language Models locally on Mac

Open-source Python package for running and fine-tuning Vision Language Models locally on Mac using Apple's MLX framework. Supports multimodal inference with images, audio, and video across Qwen, DeepSeek, Phi, and Gemma architectures. Features OpenAI-compatible API server, Gradio chat UI, and KV cache optimization. 3.8K+ GitHub stars.

About MLX-VLM

MLX-VLM is a comprehensive toolkit for running and fine-tuning Vision Language Models on Apple Silicon Macs using Apple's MLX framework. It supports inference across a wide range of VLM architectures including LLaVA, Qwen2-VL, Pixtral, Phi-3 Vision, and many more, delivering fast and memory-efficient processing without requiring cloud GPU resources. Beyond static image understanding, MLX-VLM also handles video analysis tasks such as captioning, summarization, and temporal reasoning with compatible models, making it a versatile multimodal inference engine for macOS.

The toolkit provides multiple interfaces for different workflows: a Python API for programmatic integration, a CLI for quick inference tasks, a Gradio-based chat UI for interactive exploration, and a FastAPI server for serving models over HTTP. MLX-VLM also supports LoRA and QLoRA fine-tuning, allowing developers and researchers to adapt any supported model to custom datasets directly on-device. This on-device fine-tuning capability eliminates the need for cloud GPU rentals during prototyping and experimentation phases, making it particularly valuable for teams working with proprietary or sensitive visual data.

Built entirely on Apple's MLX array framework, MLX-VLM leverages unified memory architecture and Metal GPU acceleration to maximize performance on M-series chips. The project is open-source under the MIT license and actively maintained, with regular updates adding support for new model architectures as they emerge. Installation is straightforward via pip, and quantized model variants (4-bit, 8-bit) are available through Hugging Face for reduced memory usage. MLX-VLM fills a critical gap for Apple Silicon users who want local, private multimodal AI capabilities without the latency and cost of cloud-based inference services.

Pricing & Platform Specs

Pricing Summary

100% free and open source under the MIT license ($0 software cost). Native Apple Silicon Metal-accelerated inference and LoRA fine-tuning engine for Vision-Language Models (Qwen2-VL, PaliGemma, Pixtral, Phi-3-Vision).

full pricing breakdown →

Supported Platforms

macOS (Apple Silicon)

Explore categories, tags & use cases

Run LLMs locally with one command

Tool for running large language models locally on your machine with a simple CLI interface. Download and run Llama 3, Mistral, Gemma, Phi, Code Llama, and dozens of other open-source models with a single command. Features model management, GPU acceleration (NVIDIA/AMD/Apple Silicon), OpenAI-compatible API server, Modelfile for customization, and multi-model switching. Ideal for offline AI development, privacy-sensitive use cases, and local testing. 120K+ GitHub stars.

Open Source

Run local LLMs with an intuitive desktop GUI and OpenAI-compatible API server.

Free desktop application by Element Labs for discovering, downloading, and running open-source LLMs locally. Features a curated Hugging Face model browser, side-by-side model comparison, parameter tuning, and an OpenAI-compatible API server on localhost:1234. Powered by llama.cpp with Metal acceleration for Apple Silicon.

freemium

Free, open-source local AI inference engine

LocalAI is an open-source local AI inference engine with 44K+ GitHub stars that runs LLMs, image generation, audio transcription, and embeddings entirely on consumer hardware without GPU requirements. Provides an OpenAI API-compatible REST endpoint as a drop-in replacement, supporting 1000+ models including LLaMA, Mistral, and Phi families. Features include text-to-speech, speech-to-text, function calling, constrained grammar output, and multi-modal capabilities all running locally.

Open Source

Run LLMs as a single portable executable file

Llamafile by Mozilla packages a complete LLM — model weights, inference engine, and OpenAI-compatible API server — into a single executable file that runs on Mac, Windows, Linux, FreeBSD, and OpenBSD with no installation. Built on llama.cpp and Cosmopolitan Libc for cross-platform portability, it delivers GPU-accelerated inference when available and falls back to optimized CPU execution. Supports GGUF models with a built-in web chat UI and REST API for integration.

Open Source

High-throughput LLM serving engine

vLLM is an Apache-2.0 LLM inference and serving engine focused on high-throughput self-hosted model APIs. It combines PagedAttention, continuous batching, prefix caching, quantization options, OpenAI-compatible serving, structured outputs, metrics, Docker/Kubernetes deployment guidance and integrations with agent and LLM frameworks.

Open Source

Run open-source LLMs on your phone, fully offline and private

Google AI Edge Gallery is an open-source mobile app that lets you download and run large language models like Gemma directly on Android and iOS devices with zero cloud dependency. Built on MediaPipe and LiteRT, it features AI chat with reasoning mode, multimodal image analysis, real-time audio transcription, and autonomous agent skills—all running entirely on-device for complete privacy. A reference implementation for developers building offline-first AI experiences.

freeOpen Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is MLX-VLM?

Open-source Python package for running and fine-tuning Vision Language Models locally on Mac using Apple's MLX framework. Supports multimodal inference with images, audio, and video across Qwen, DeepSeek, Phi, and Gemma architectures. Features OpenAI-compatible API server, Gradio chat UI, and KV cache optimization. 3.8K+ GitHub stars.

Is MLX-VLM free?

Yes — MLX-VLM is open source and free to use. 100% free and open source under the MIT license ($0 software cost). Native Apple Silicon Metal-accelerated inference and LoRA fine-tuning engine for Vision-Language Models (Qwen2-VL, PaliGemma, Pixtral, Phi-3-Vision).

Is MLX-VLM open source?

Yes — MLX-VLM is open source.

Is MLX-VLM still maintained?

Yes — MLX-VLM is active. Its listing was last verified on September 6, 2026.

What are the best MLX-VLM alternatives?

The first editor-selected MLX-VLM alternatives are Ollama, LM Studio, LocalAI, and more.