PrismML compresses Qwen3.8 27B to 5.93 GB with ternary weights
TL;DR
- PrismML improved retention from 95% to 98.2% across two model generations, with math and coding benchmarks essentially matching full-precision Qwen3.8 27B.
- The 5.9GB footprint is independently verifiable via arithmetic: 27B parameters at 1.76 bits yields approximately 5.94GB, confirming the compression claim.
- Vision capability loads as a separate 0.63GB file, meaning text-only workloads carry no multimodal memory penalty at runtime.
PrismML's Ternary Bonsai 2 27B takes Alibaba's Qwen3.8 27B and squeezes it from 53.80 GB down to 5.93 GB by holding every language-model weight to one of three values: minus one, zero, or plus one. According to MarkTechPost, the ternary scheme lands the model at roughly 1.72 bits per weight, with only 26.2 million parameters, or 0.0976% of the total, kept in higher precision for recurrent state paths and normalization. "Each group of 128 weights shares 1 FP16 scale," the writeup notes, and a blockwise Hadamard rotation inspired by SpinQuant is applied before ternary assignment.
On 20 benchmarks the compressed model averages 83.9 against the parent's 85.4, or 98.2% retention. Math holds up best at 99.5% and coding at 99.3%.
The weak spot is agentic work. On Terminal-Bench 2.1, Ternary Bonsai 2 scores 52.8 against the full-precision Qwen3.8 27B's 69.7, a gap of nearly seventeen points that the retention headline glosses over.
Throughput is the other selling point: 142.5 tokens per second on an RTX 5090 at 0.582 mWh per token, 46.8 on an M5 Max, 27.7 on an M5 Pro. The Apache 2.0 weights sit on Hugging Face, with MLX packs for Apple Silicon and a WebGPU browser demo, though loading them requires "PrismML's llama.cpp fork" because stock llama.cpp cannot parse the PTQ1_0 and PQ2_0 formats. It arrives amid a busy week of open-weight releases we've been logging in open-source AI.
What others are reporting
-
PrismML Read →
First-party post with full benchmark table, throughput numbers (143 tokens/sec on RTX 5090), and explicit generation-over-generation retention comparison (95% to 98.2%).
Compression becomes a deployment unlock: nearly the same capability, in a footprint that can run in far more places.
-
PR Newswire Read →
Official press release with on-record CEO quote; frames the release around closing the performance gap to full-precision models rather than a throughput story.
Bonsai 27B proved that powerful models do not have to be confined to cloud infrastructure. — Babak Hassibi, PrismML CEO
-
Progressive Robot Read →
Independently verified the 9x compression claim via arithmetic, identified where capability loss concentrates (vision, knowledge), and flagged unverified energy and partnership claims.
If capable models run on the device you already own, the economics, privacy and latency of everyday AI all shift at once.
-
OpenRouter Read →
Confirms the model is live on OpenRouter's API marketplace at launch with full tool-calling and streaming support, making it accessible without local hardware.
Originally reported by marktechpost.com
Read the original article →Original headline: PrismML Ships Ternary Bonsai 2 27B: 5.9GB Model Retains 98.2% of Qwen3.8 27B Under Apache 2.0