Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

PrismML's Bonsai 2 27B squeezes a 27B open-source model into 5.9GB

Caltech-spun AI lab PrismML released Bonsai 2 27B on Thursday, compressing Alibaba's open-source Qwen3.8 27B down to 5.9GB, roughly a 9x to 10x memory reduction that fits on a PC and possibly a high-end smartphone. The company says the new model matches 98% of Qwen's aggregate benchmark scores, up from 95% for the first Bonsai, on a $22.25 million seed round.

Published
PrismML 发布 Bonsai 2 27B:27B 开源模型被压到 5.9GB,可跑在 PC 上
Image source: techcrunch.com

PrismML, an AI lab founded by a group of Caltech researchers, released Bonsai 2 27B on Thursday. The model compresses Qwen3.8 27B, a widely used open-source model from Alibaba, down to 5.9GB.

That is roughly a 9x to 10x reduction in memory versus the original, small enough to fit on a PC and possibly a high-end smartphone. The bet behind the release is that capable, high-performing reasoning models do not, in fact, have to be large.

Accuracy is what PrismML is selling. According to the company, Bonsai 2 matches 98% of Qwen's aggregate benchmark scores, up from 95% for the first Bonsai, which shipped in March. The original model has been downloaded more than 11 million times, and PrismML's even smaller models another 2.6 million, the company says — the trend from one release to the next is the signal it wants people to see.

The technique is about how a model stores what it learned. In a normal model each weight needs 16 bits; PrismML uses "ternary" weights that simplify each one down to +1, -1, or 0. Fewer bits per weight means a smaller footprint.

The company itself is early. It is led by CEO Babak Hassibi, a Caltech professor and an expert in compression technologies, and counts Ion Stoica as an advisor — Stoica co-founded Databricks and directs Berkeley's Sky Computing Lab, which has spun out projects from Letta to SGLang. PrismML is backed by Khosla Ventures, Cerberus Capital, and Caltech, and has raised just a $22.25 million seed round.

It is far from alone in the compression race. Multiverse Computing, founded by a well-known professor from Spain's Donostia International Physics Center, works on the same problem and has raised substantially more money. Hassibi argues PrismML is unique because its models lose virtually no performance compared with the originals, but he also says compression will always have some impact, and whether it ever reaches 100% benchmark parity remains to be seen.

Whether that parity actually matters is a separate question. LLMs are not so accurate in their uncompressed form, and benchmarks are not so perfectly reflective of real tasks, that a 2% degradation would likely change how a model performs in actual use — the surrounding harness matters a lot for accuracy too. The real question is whether on-device models get close enough in daily work that the default place to run a 27B-class model moves from the cloud to the device.

For now, the thing to watch is whether the gap keeps narrowing in the next Bonsai release, and whether the on-device experience matches the benchmark story once real hardware and real workloads are involved.

Why it matters

If ternary-weight compression holds up in real workloads, the default place to run a capable reasoning model could shift from cloud APIs to the device itself, changing both cost structure and privacy boundaries.

PrismMLBonsai 2Open SourceOn-device AI
Back to realtime news

Nearby Updates

All

09/18, 06:14

The FAA's plan to fix air traffic? $875M worth of AI

The FAA is preparing to launch SMART, an AI-based cloud platform meant to help air traffic controllers manage workloads and route flights more safely, according to The Wall Street Journal. The program from Air Space Intelligence will cost $875 million over 12 years and rolls out in the Washington, D.C. metropolitan area first.

09/18, 07:25

Crusoe raises $3.9B at a $30.9B valuation to build mega data centers and modular AI factories

Data center developer Crusoe said Thursday it raised $3.9 billion in a Series F round that lifts its valuation to $30.9 billion, co-led by Atreides Management, Mubadala Capital and Valor Equity Partners. The capital will fund existing projects, including a large Texas site used by OpenAI, plus truck-transportable modular AI factories the company calls Spark.

09/18, 04:55

GitLab ships version 19.4 with expanded AI agent tools

GitLab has released version 19.4, an update whose focus is expanded AI agent tooling inside the platform, according to a report by Investing.com. The release shows code hosting and DevOps platforms folding agent capabilities into the product core rather than offering them only as outside add-ons.

09/18, 04:05

Report: US government website used a Chinese AI search tool the FBI said copied Anthropic

According to whbl.com, a US government website used an AI search tool from China that the FBI has said copied Anthropic's technology. The report links two threads usually discussed separately: allegations that Chinese AI products copied American models, and whether such tools are already running on public-facing government services.