PrismML’s Tiny LLM Unlocks an Astonishing Shift in Everyday AI

Quick Reads
- PrismML’s tiny LLM, Bonsai 2 27B, compresses Qwen3.8 27B down to 5.9 GB.
- The compression delivers a 9x to 10x reduction in memory versus the original model.
- Bonsai 2 matches 98% of Qwen’s benchmark scores, up from 95% in March.
- The Caltech-founded startup has raised $22.25 million in seed funding so far.
- PrismML plans to compress hundred-billion-parameter models within the next few months.
PrismML’s tiny LLM ambitions are turning heads across Silicon Valley, and the reasons run deeper than hype. The Caltech-founded startup argues that powerful reasoning models don’t need massive size to perform well. Rumors even link the company to Apple, though its CEO declined to confirm any talks. Khosla Ventures, Cerberus Capital, and Caltech back the startup alongside its seed investors.
On Thursday, PrismML released Bonsai 2 27B, its newest model, according to the company’s official announcement. The model compresses Qwen3.8 27B, a popular open-source model from Alibaba, down to just 5.9 GB. That shrinkage delivers a nine-to-tenfold reduction in memory compared with the original. Consequently, the compressed model can now fit on a regular PC and possibly a high-end smartphone.
Bonsai 2 matches 98% of Qwen’s aggregate benchmark scores. That marks a notable jump from the 95% the original Bonsai model hit in March. That earlier model has already racked up more than 11 million downloads. PrismML’s smaller models have added another 2.6 million downloads on top of that. Hassibi believes compression always costs some performance, though the gap rarely matters in real-world use.
PrismML’s tiny LLM approach relies on shrinking a model’s underlying weights, the information it learns during training. Standard weights typically need 16 bits each to store. PrismML’s ternary method reduces that to just three values: +1, −1, or 0. This dramatically cuts the space each model needs to run.
PrismML isn’t alone in this race. Multiverse Computing, a well-funded Spanish rival, is also chasing compression breakthroughs. Still, Hassibi insists PrismML’s models lose almost no performance against their originals. The startup plans to apply its technique to models in the hundred-billion-parameter range within months, as TechCrunch’s original report explains. Hassibi expects larger models to compress even more easily without losing intelligence. Stoica, a PrismML adviser and Databricks co-founder, calls the shift a turning point for everyday users. He says the technology will soon put private, free intelligence directly on people’s own devices.





