Skip to content
aicoolies logo
PrismML Bonsai logo

PrismML Bonsai

First commercially viable 1-bit LLMs that are 14x smaller and 8x faster

PrismML Bonsai delivers the first commercially viable 1-bit large language models with 8B, 4B, and 1.7B parameter variants. The 8B model runs in just 1GB of RAM versus 16GB for standard FP16 models, achieving 44 tokens per second on iPhone. Backed by $16.25M from Khosla Ventures and released under Apache 2.0, Bonsai makes capable LLMs practical for edge devices and resource-constrained environments.

About PrismML Bonsai

PrismML emerged from stealth in March 2026 with Bonsai, a family of 1-bit language models that achieve dramatic efficiency gains without proportional quality loss. The 8B parameter model requires only 1GB of memory compared to 16GB for a standard FP16 Llama 3 8B, representing a 14x reduction in model size. Inference runs at 44 tokens per second on an iPhone and scales even faster on desktop hardware. This efficiency breakthrough makes capable language models practical for mobile devices, IoT endpoints, and any environment where compute and memory are constrained.

The technical approach uses a novel 1-bit quantization architecture with FP16 scale factors applied every 128 bits. PrismML provides custom forks of llama.cpp and MLX optimized for 1-bit inference, along with demo code, Colab notebooks, and developer integration documentation. The models are available on HuggingFace under the Apache 2.0 license. AnythingLLM integrated Bonsai models on launch day, demonstrating immediate ecosystem adoption and compatibility with existing local LLM infrastructure.

With $16.25M in funding from Khosla Ventures, Cerberus, and Google, PrismML has the backing to develop the 1-bit quantization toolchain into a comprehensive platform. The Bonsai 8B, 4B, and 1.7B models provide different capability-efficiency tradeoffs for various deployment scenarios. The 355-point Hacker News Show HN launch and positive reception on r/LocalLLaMA confirm strong community interest in edge-efficient LLMs. For developers building on-device AI experiences, Bonsai represents the most practical path to running capable models without cloud dependencies.

Pricing & Platform Specs

Pricing Summary

Commercial edge AI compression and on-device runtime platform. Custom enterprise pricing based on deployment volume, target edge silicon architectures (Apple Silicon, ARM, edge GPUs), proprietary compiler access, and enterprise SLA support.

full pricing breakdown →

Supported Platforms

Custom llama.cpp/MLX forks; HuggingFace; runs on iPhone, desktop, edge

Explore categories, tags & use cases

Run LLMs natively on any device with ML compilation

MLC LLM is an open-source engine for deploying large language models natively across diverse platforms using machine learning compilation. It runs models on NVIDIA/AMD GPUs, Apple Silicon, mobile devices, and browsers via WebGPU without cloud dependencies. Features include OpenAI-compatible API, quantization support, and optimized backends for CUDA, Metal, Vulkan, and WebAssembly.

Open Source

PyTorch on-device AI for mobile and edge devices

ExecuTorch is PyTorch's official solution for deploying AI models on mobile, embedded, and edge devices. It features a 50KB base runtime, 12+ hardware backends including Apple CoreML, Qualcomm QNN, ARM, and Vulkan, and native PyTorch export without format conversions. Powers Meta's on-device AI across Instagram, WhatsApp, Quest 3, and Ray-Ban Smart Glasses, supporting LLMs, vision, speech, and multimodal models.

Open Source

Run LLMs as a single portable executable file

Llamafile by Mozilla packages a complete LLM — model weights, inference engine, and OpenAI-compatible API server — into a single executable file that runs on Mac, Windows, Linux, FreeBSD, and OpenBSD with no installation. Built on llama.cpp and Cosmopolitan Libc for cross-platform portability, it delivers GPU-accelerated inference when available and falls back to optimized CPU execution. Supports GGUF models with a built-in web chat UI and REST API for integration.

Open Source

Side-by-Side Comparisons

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is PrismML Bonsai?

PrismML Bonsai delivers the first commercially viable 1-bit large language models with 8B, 4B, and 1.7B parameter variants. The 8B model runs in just 1GB of RAM versus 16GB for standard FP16 models, achieving 44 tokens per second on iPhone. Backed by $16.25M from Khosla Ventures and released under Apache 2.0, Bonsai makes capable LLMs practical for edge devices and resource-constrained environments.

Is PrismML Bonsai free?

No — PrismML Bonsai is a paid tool. Commercial edge AI compression and on-device runtime platform. Custom enterprise pricing based on deployment volume, target edge silicon architectures (Apple Silicon, ARM, edge GPUs), proprietary compiler access, and enterprise SLA support.

Is PrismML Bonsai still maintained?

Yes — PrismML Bonsai is active. Its listing was last verified on September 6, 2026.

What are the best PrismML Bonsai alternatives?

The first editor-selected PrismML Bonsai alternatives are MLC LLM, ExecuTorch, Llamafile.