Nine Times Smaller, 98.2% as Smart: PrismML's Ternary Bonsai 2 Puts a Full 27B Model in 5.9 GB
The Caltech spinout's ternary Bonsai 2 27B compresses Qwen3.8 27B into a 5.9 GB footprint while keeping 98.2% of its benchmark average — and beating conventional 2-bit quantization by twelve points where it matters most.
For most of the past decade, running a 27-billion-parameter model meant reserving roughly 54 GB of high-bandwidth memory just for the weights. PrismML, the Caltech spinout that emerged from stealth in March with a $16.25 million seed round led by Khosla Ventures, has just published its strongest argument yet that this trade-off is optional. Ternary Bonsai 2 27B, released September 18 under the Apache 2.0 license, stores a full 27B-class model in 5.9 GB — about a ninth of the FP16 original — while retaining 98.2% of its benchmark average across fourteen evaluations.
What the release actually is
Bonsai 2 27B is a compressed build of Alibaba’s Qwen3.8 27B, the strongest open-weight model in that size class. PrismML’s approach constrains the language model’s weights to ternary values — −1, 0, or +1 — with FP16 group-wise scaling, yielding 1.76 effective bits per weight. The technique descends from the BitNet line of research, but PrismML’s differentiator is that it applies the low-bit representation end to end across the language model and then measures, in public and per benchmark, exactly how much capability survives.
The headline numbers come from an evaluation run with EvalScope and vLLM on NVIDIA H100s in thinking mode, where the model’s full reasoning path is exercised. Across fourteen benchmarks spanning knowledge, math, coding, instruction following, tool calling, and vision, FP16 Qwen3.8 27B averages 86.32. Bonsai 2 27B averages 84.78 — the 98.2% retention figure. The model supports a 262K-token context window, accepts text and image input, and runs on NVIDIA GPUs via CUDA or Apple silicon via MLX with custom low-bit kernels.
Where the capability lives
The aggregate score is not the interesting part. What matters is the distribution, because conventional quantization does not degrade uniformly — it collapses selectively, precisely on the benchmarks that demand sustained chains of reasoning.
PrismML’s model card makes the comparison explicit. Qwen3.8-27B quantized with IQ2_XXS, a standard sub-4-bit method, scores a healthy-looking 88.93 on MMLU-Redux — which is why casual testing misses the damage. But on AIME26 it falls to 57.5, and on LiveCodeBench to 56.4. Bonsai 2 27B scores 95.83 and 90.07 on those same benchmarks. On the fourteen-benchmark average, the 9.4 GB IQ2_XXS build manages 72.59, meaning the 5.9 GB ternary model outscores it by more than twelve points while being a third smaller.
By skill category, the picture is consistent: math drops only from 97.06 to 96.57, coding is actually level with the FP16 baseline (89.07 to 89.42), and instruction following comes in slightly ahead (81.25 to 82.66). The residual gap concentrates in knowledge and reasoning (85.55 to 79.86) and vision (71.36 to 66.19). Notably, the previous Bonsai report showed the same selective-collapse pattern on a second model family, Gemma-4-31B — evidence that the failure mode belongs to conventional low-bit methods, not to one base model.
Throughput and energy
Compression buys more than memory. PrismML reports up to 143 tokens per second on an RTX 5090 and 46.8 tok/s on an Apple M5 Max; the model card adds roughly 28 tok/s on a standard laptop, ~130 tok/s on an RTX 5090 in serving configurations, and ~30 tok/s even on a 72 W datacenter L4. On an RTX 4090, decode costs 0.714 mWh per token — about 40% less energy than running an 8B model at full precision.
PrismML frames this as “intelligence density”: usable capability per stored gigabyte. By that measure Bonsai 2 achieves over 2.3x the density of the densest conventional build of the same model and roughly 9x that of FP16.
The honest caveats
Two limitations deserve attention. First, the company itself concedes that unpacking ternary weights costs arithmetic: the dense PTQ1_0 packing is faster on Ada-class and smaller accelerators, but batch-1 decode is actually slower on Ampere, Hopper, and Blackwell parts, where memory bandwidth is not the bottleneck. Second, community scrutiny is already doing its job — the Hacker News thread on the release includes users testing edge cases like poetry recitation against conventional GGUF quants, and Reddit commenters are asking for exactly the per-benchmark breakdowns the model card now provides. A 98.2% average is not 98.2% everywhere; vision and deep knowledge work still pay a visible tax.
There is also strategic context worth knowing. PrismML emerged from Caltech research led by CEO Babak Hassibi, with backing from Khosla Ventures, Cerberus, compute grants from Google, and continuing support from Samsung. In July, reports surfaced that Apple had held talks to acquire the company, precisely because extreme compression is the key to frontier-class models on iPhones. Whether or not an acquisition happens, the direction is clear.
Why it matters
The frontier conversation of the past month has been dominated by compute budgets in the hundreds of billions. Bonsai 2 is a reminder that the opposite lever — doing dramatically more with dramatically less — is compounding just as fast. If a 27B model now fits alongside your photos on a laptop, the economics of agent loops change: local models can handle sensitive or high-frequency steps and escalate selectively to the cloud, privacy-sensitive deployments stop requiring a datacenter, and the same math that shrinks a phone model also lets a datacenter serve more users per GPU.
The question PrismML poses to the field is worth taking seriously: not just how capable a model is, but how much useful intelligence it delivers within a fixed memory, compute, and power budget. On this release’s evidence, the answer is moving faster than most people assumed.
Sources
- [1] https://prismml.com/news/bonsai-2-27b
- [2] https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
- [3] https://www.marktechpost.com/2026/09/18/prismml-releases-ternary-bonsai-2-27b-a-5-9-gb-apache-2-0-model-retaining-98-2-of-qwen3-8-27b-performance/
- [4] https://www.hpcwire.com/2026/04/03/prismml-emerges-from-stealth-with-1-bit-llm-family/
- [5] https://news.ycombinator.com/item?id=49746765