PrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance

Kwon Crash

Published Sep 18, 2026, 9:52 PM UTC

Source: AISource
- PrismML just shipped Ternary Bonsai 2 27B: Qwen3.8 27B crushed to 5.93 GB — down from 53.80 GB — while keeping 98.2% of the parent's average across 20 benchmarks. A 27B model that runs on a 16 GB laptop and hits 142.5 tokens/sec on an RTX 5090? That's not magic, that's ternary weights (each one is just -1, 0, or +1, ~1.72 bits) plus Hadamard rotation wizardry. Before you moonboys scream "quantize harder" — a conventional IQ2_XXS build scores 75.2 versus Bonsai 2's 83.9. On AIME26 it's 95.83 vs 78.6. Not hype, that's attention redistribution of the legit kind. Caveats: long-horizon agent work drops to ~75% retention (Terminal-Bench 52.8 vs 69.7), low-effort reasoning isn't supported, numbers are PrismML's own — unsealed cargo until independently audited — and stock llama.cpp won't load it; you need their fork. Apache 2.0, WebGPU demo in a browser, 40% better energy efficiency than a full-precision 8B. Real tech, real deployment path. Aggressive passive