ASX 2009,005.90
-14.20(-0.16%)
NIKKEI65,020.94
+806.46(+1.26%)
NIFTY 5023,897.70
+24.25(+0.10%)
HSI25,650.87
+427.66(+1.74%)
SHANGHAI3,930.116
-11.972(-0.30%)
Trending:US MarketsAI & SiliconUSA Jobs DeskFed PolicyCybersecurityGov & LawEntertainmentSports Wire
TECH/AI/AAPL

PrismML hopes its tiny LLM could change how we all use AI

If AI lab PrismML isn't on your radar yet, it should be.

By Nexvoro Tech Wire
PUBLISHED THU, SEP 17, 2026 10:45 PM UTC6 MIN READ
CNBC Market Tracker • NASDAQ:AAPL
REAL-TIME QUOTE
Apple Inc
$234.12-0.98 (-0.42%)
Volume: 68.4M
52-Wk Range: $138.80 - 271.00

KEY POINTS

  • Primary coverage dispatched via TechCrunch.
  • Signals noteworthy shifts in sector dynamics and operational developments.
  • Comprehensive factual details verified from official publication records.
  • Objective, non-partisan journalistic standards preserved.
PrismML hopes its tiny LLM could change how we all use AI
PHOTO VIA TECHCRUNCHNEXVORO EDITORIAL WIRE

Primary Journalistic Dispatch & Direct Reporting

Disrupt 2026: OpenAI, Anthropic, Replit, and more take over 6 industry stages. 25% off tickets now

If AI lab PrismML isn't on your radar yet, it should be - not because it's raised gobs of money (it hasn't yet, just a $22.25 million seed round), but because of the technical minds involved and the potentially industry-changing tech it's developing.

PrismML is betting that capable, high-performing, reasoning large language models don't, in fact, have to be large.

In-Depth Developments & Factual Context

It is making reasoning models so small they can fit on PCs and smartphones. (It's even rumored to be in talks with Apple , though CEO Babak Hassibi declined to comment on that to TechCrunch.)

On Thursday, PrismML released Bonsai 2 27B , its latest in a family of models, which compresses Qwen3.8 27B, a widely used open-source model from Alibaba, down to 5.9 GB. That's small enough to fit on a PC and, possibly, a high-end smartphone. It's a 9x to 10x reduction in memory versus the original.

PrismML was founded by a group of Caltech researchers and is led by Hassibi, a Caltech professor and an expert in compression technologies. The startup also counts Ion Stoica as an advisor. Stoica is a co-founder of Databricks (and other companies) and the director of Berkeley's famed Sky Computing Lab, which has birthed many technologies and startups, from Letta to SGLang .

Industry Impact & Strategic Analysis

PrismML is also backed by investors Khosla Ventures, Cerberus Capital, and Caltech.

This startup is certainly not the only company working on LLM compression tech. Multiverse Computing, founded by a well-known professor from Spain's Donostia International Physics Center, is another. (And Multiverse Computing has raised gobs of cash .)

But Hassibi says that PrismML's compression tech is unique because its LLMs have lost virtually no performance compared with the originals. Bonsai 2 matches 98% of Qwen's aggregate benchmark scores. That's up from the first Bonsai, released a couple of months ago in March, that matched 95%. That original model has already been downloaded over 11 million times, and PrismML's even smaller models have been downloaded another 2.6 million times, the company says.

Forward Outlook & Market Perspective

So this shows that PrismML's compression results have improved from one release to the next. Whether it could ever get to 100% benchmark performance parity is a question that remains to be seen. Compression will likely always have some impact, Hassibi says.

Still, perfect benchmark parity is fairly academic anyway. LLMs are not so accurate in their uncompressed form, and benchmarks not so perfectly reflective of actual tasks, that a 2% degradation would likely meaningfully affect how a model performs in actual use. (Plus, the surrounding software - the harness a model runs inside of - matters a lot when it comes to accuracy , too.)

PrismML says it achieves this by shrinking the "weights" that make up a model - weights are, essentially, the information a model learns and stores during training. Normally, each weight requires 16 bits. PrismML's approach, called "ternary" weights, simplifies that down to three: +1, −1, or 0. With far smaller values to store for each weight, the model takes up dramatically less space. (For a deeper dive on the compression technique, here's the project's GitHub page .)

The startup's next goal is to apply this compression technique to even bigger models. "The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there," Hassibi told TechCrunch.

As model size grows, he added, "There is more room to be able to compress them without losing the intelligence. So I would just say, as a general trend, for larger models, it's easier to get to 100%."

Stoica tells us that he's excited for this tech because it's making it possible for advanced models to run on users' devices. "You are going to have intelligence at your fingertips, and it's going to be free because it's going to run on the device you already bought. It's also going to be private, because you're not going to send it to the cloud."

Reporting synthesized and verified under Nexvoro.tech editorial guidelines. Full primary records referenced via TechCrunch.

Sponsored / Google AdSense SlotResponsive Leaderboard 728x90 / 970x250 (article-mid-story)
Reporting synthesized under Nexvoro.tech Editorial Standards • Referenced via TechCrunch
Verified Dispatch
Related Tickers:#AI#US NEWS#TECHCRUNCH

More Coverage in AI

View Topic Desk →
OpenAI Urges Global AI Standards and Safety Guardrails Amid Rising Industry Anxiety Over Recursive Self-Improvement
AI
AI17H AGO

OpenAI Urges Global AI Standards and Safety Guardrails Amid Rising Industry Anxiety Over Recursive Self-Improvement

As debate intensifies over the rapid acceleration of artificial intelligence, OpenAI has proposed a comprehensive framework for international safety standards, focusing heavily on alignment research and recursive self-improvement. The move follows recent high-profile departures and escalating concerns from industry insiders regarding humanity's long-term control over advanced frontier models.

CNBC World & Geopolitics6 min read
Inside the White House: How Nvidia CEO Jensen Huang Became President Trump's Most Trusted AI Ally
AI
AISEP 20

Inside the White House: How Nvidia CEO Jensen Huang Became President Trump's Most Trusted AI Ally

As Washington fiercely debates artificial intelligence oversight, Nvidia CEO Jensen Huang has emerged as President Donald Trump's top confidant, successfully pushing back against growing regulatory pressures. While rival tech executives advocate for government slowdowns, the head of the world's most valuable chipmaker is charting a rapid course for American tech supremacy.

CNBC Top News7 min read