Nine Times Smaller, 98.2% as Capable: PrismML's Ternary Bonsai 2 27B Squeezes a 27B Model Into 5.9 GB
PrismML compresses Qwen3.8 27B from 53.8 GB to 5.9 GB using ternary weights at ~1.76 bits each, retaining 98.2% of aggregate benchmark performance — and the Apache-2.0 weights now run on an 8 GB consumer GPU.
For most of the past decade, running a 27-billion-parameter model meant one thing: a datacenter GPU with at least 54 GB of memory, or a quantization scheme that quietly lobotomized the model on the way down. This week, a Caltech-spawned startup called Prism ML published a release that makes that trade look increasingly obsolete. Ternary Bonsai 2 27B takes Alibaba’s Qwen3.8 27B — a full multimodal, 262K-context frontier-class small model — and compresses it from roughly 53.8 GB in FP16 down to 5.9 GB, while retaining 98.2% of its aggregate benchmark performance.
That is not a rounding error away from lossless. It is a 9x reduction in memory footprint with a measured capability loss of less than two points, released under Apache 2.0 with weights, GGUF packings, custom kernels for CUDA and Apple’s MLX, and a public whitepaper.
What Ternary Bonsai 2 27B actually is
The core idea is aggressively simple: every weight in the language model is forced to one of three values — −1, 0, or +1. Instead of storing 16-bit floating-point numbers, you are storing what is effectively a trit, which packs down to roughly 1.76 effective bits per weight. Prism ML applies this end to end across the language model, not just to selected layers.
Three engineering choices keep the model from falling apart under that constraint. First, FP16 group-wise scaling: every group of 128 weights shares a single FP16 scale factor, so the network retains dynamic range even when individual weights are ternary. Second, a blockwise Hadamard rotation — inspired by Meta’s SpinQuant rotation techniques — is applied before ternary assignment, spreading outliers across dimensions so no single weight has to encode an extreme value. Third, a tiny sliver of the network — about 26.2 million parameters, or 0.0976% of the total — is deliberately kept in higher precision to protect recurrent state paths and normalization layers, the parts of a transformer most sensitive to low-bit collapse.
The result is a model that keeps the deployment promises of a 2B-class model with the behavior of a 27B-class one: a 262K-token context window, multimodal text-and-image input, and the same thinking-mode reasoning behavior as its full-precision parent.
The benchmark picture
Across a suite spanning reasoning, math, coding, instruction following, vision, and agentic tool use, Ternary Bonsai 2 27B scores 83.9 overall against Qwen3.8 27B’s 85.4 — and, notably, ahead of the previous-generation Qwen3.6 27B’s 83.6. In other words, the compressed model beats the full-precision model from one generation earlier.
The per-category breakdown from Prism ML’s published table:
- Agentic & tool calling (τ²-bench, BFCLv3): 77.57 vs 79.74 for FP16
- Coding (HumanEval+, LiveCodeBench v6, MBPP+, BigCodeBench): 81.58 vs 82.17
- Instruction following (IFBench, IFEval): 82.66 vs 81.25 — ahead of full precision
- Knowledge & reasoning (MMLU-Redux, GPQA Diamond): 83.95 vs 86.66
- Math (AIME 2025/2026, GSM8K, MATH-500): 96.57 vs 97.06
- Vision (CharXiv, A-OKVQA, OCRBench v2): 78.59 vs 81.64
The distribution matters more than the aggregate. Agentic coding loops, tool-use chains, and long-horizon tasks are exactly where quantized models historically fall over, because small per-step errors compound across dozens of steps. Bonsai 2’s tightest retention gaps are in coding and math — the categories that compound. Vision shows the widest gap (about three points), which is consistent with what the quantization community has seen for years: vision towers hate low bits more than language layers do.
Throughput and energy
Compression is not just about fitting. Prism ML reports up to 143 tokens/second on an RTX 5090 and 46.8 tokens/second on an M5 Max, thanks to custom low-bit kernels rather than generic dequantize-then-multiply paths. On an RTX 4090, the model consumes 0.714 mWh per token — which the company measures as 40% more energy-efficient than an 8B model running in full precision. A model nine times larger by parameter count, sipping less power than an 8B FP16 baseline, inverts the usual efficiency calculus.
The community has already stress-tested the claim that this runs on consumer hardware. On Hacker News, users report running it on an RTX 3070 with 8 GB of VRAM using the GGUF release, which ships two packings with custom ternary hybrid-attention kernels for llama.cpp. Reddit’s r/LocalLLaMA has active threads probing it for solo agentic coding and computer-use workflows — the exact workloads Prism ML is targeting with its demo of coding agents in Cline running locally on a 5090.
Why this matters beyond local inference
The first Bonsai 27B release two months ago landed at about 95% retention. Closing the gap from 95% to 98.2% in a single generation is the difference between “impressive demo” and “practical default.” Below roughly two points of aggregate degradation, compression stops being a compromise for most real applications — and starts being a deployment unlock.
The economics cascade up the stack. If a 27B model fits where a 3B model used to, then laptops, phones, and edge boxes get frontier-adjacent capability with no network round-trip, no per-token API bill, and no data leaving the device. In datacenters, the same math means more users per GPU, more requests per watt, and larger effective models inside a fixed memory envelope. It also makes hybrid architectures genuinely viable: a local ternary model handling sensitive or high-frequency work, escalating selectively to cloud frontier models only when the task demands it.
There is a strategic wrinkle worth noting: the base model is Alibaba’s Qwen3.8 27B, but the compression breakthrough comes from a US startup backed by Khosla Ventures, Cerberus, and Google, with continuing support from Samsung. The frontier of “intelligence density” — how much capability you deliver per gigabyte and per milliwatt — is shaping up as its own competitive axis, distinct from raw capability leaderboards. Prism ML is betting that the question increasingly asked of any model will be not just “how smart is it,” but “how much of that intelligence fits in my budget.”
For developers, the barrier to checking the claims is low: Apache 2.0 weights are on Hugging Face, GGUF builds work with llama.cpp today, MLX builds cover Mac, iPhone, and iPad, and the whitepaper documents the compression and evaluation methodology in full. The era where running a serious model locally meant buying a 24 GB GPU may quietly be ending — one trit at a time.
Sources
- [1] https://prismml.com/news/bonsai-2-27b
- [2] https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
- [3] https://news.ycombinator.com/item?id=49746618
- [4] https://www.reddit.com/r/LocalLLM/comments/1wldy67/ternary_bonsai_2_27b_qwen38_for_solo_agentic/
- [5] https://atomic.chat/blog/guides/how-to-run-bonsai-2-locally
- [6] https://aiweekly.co/ai-news-today/edition/2026-09-22