The startup’s latest release, Bonsai 2 27B, compresses Alibaba’s Qwen3.8 27B model into a 5.9 GB package. According to CEO Babak Hassibi, the model captures 98% of the original’s benchmark performance, a jump from the 95% achieved by their initial release in March. While perfect parity remains elusive, the company argues that a 2% variance is negligible for most practical applications. PrismML achieves this efficiency by utilizing 'ternary' weights, which simplify the standard 16-bit data representation down to three values: +1, -1, or 0.
Backed by investors including Khosla Ventures and Cerberus Capital, the firm is already looking toward larger horizons. Hassibi plans to apply this compression to models in the several-hundred-billion-parameter range, where he expects to retain even more intelligence. Advisor Ion Stoica, a co-founder of Databricks, emphasizes that the real value lies in local execution. By moving advanced AI from the cloud to the device, users gain private, cost-free intelligence that runs entirely on hardware they already own.

Comments (0)
No comments yet. Be the first!