The startup, PrismML, said it has shrunk down Qwen 3.6, an open-source large language model developed by Chinese internet giant Alibaba, to run on an iPhone 17 Pro. The model has 27 billion parameters, which are roughly similar to the synapses in a brain and can help determine the complexity of the data a model can process. In contrast, most models that run on mobile phones have only a few billion parameters active at a time.
The largest AI models, which can measure in the trillions of parameters, are still far too big to run on mobile devices. But the model PrismML has working on an iPhone is capable of tasks like complex chat, reasoning, fully autonomous agents and software coding, the startup said. The open-source model will be available for download next week on Tuesday.
In an interview, Babak Hassibi, CEO of PrismML, predicted that the vast majority of AI will eventually be processed on devices.
“Imagine a world, maybe three years from now, where 95% of the intelligence that you need is available to you locally, on your phone, on your laptop, on your appliances, and it’s really on the last maybe 5% of high-end stuff that you’ll need to go to the cloud,” Hassibi said. “I think that’s how people are seeing the way forward.”
Shrinking models to run on devices, he added, “fundamentally changes the economics of AI.”
PrismML uses a mathematical trick to shrink the Qwen 3.6 model to a fraction of its original size. Shrinking models typically results in worse performance, but the company claims its technique for miniaturizing AI model sizes doesn’t hinder their performance. PrismML has compressed the size of Qwen 3.6 to less than 4 gigabytes, down from around 54.
PrismML is a spinoff of the California Institute of Technology, where Hassibi, a professor of electrical engineering at the school, and his co-founders conducted the mathematical research used in the startup’s technology. Caltech owns the patents behind the technology but licenses them exclusively to PrismML.
PrismML plans to continue shrinking larger AI models, even at the scale of a trillion parameters, which will bring it into the realm of cutting-edge models such as OpenAI’s GPT and Anthropic’s Claude, said Hassibi.
As part of the new Siri announcement, Apple said some of the iPhone’s new AI capabilities would run on devices. One new on-device Apple model has 20 billion parameters but uses a so-called sparse architecture, in which only 1 billion to 4 billion parameters are active at a time. In the case of PrismML’s on-device model, all 27 billion parameters are active at the same time.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.