An advertisement for the new iPhone Air is displayed as customers enter the Apple Store for the release of new iPhone 17 models on September 19, 2025 in New York.
Angela Weiss | AFP | Getty Images
Apple is in talks with a small Silicon Valley company that says it can shrink powerful artificial intelligence models to the point where they run directly on an iPhone, the startup’s CEO told CNBC.
PrismML, a California Institute of Technology spinout backed by Khosla Ventures, publicly released compressed versions of Alibaba’s open-source Qwen model on Tuesday. The company said it reduced the model from about 54GB to less than 4GB, allowing all 27 billion of its parameters to run on an iPhone 15 or newer.
PrismML CEO Babak Hassibi told CNBC that Apple and other companies have evaluated the startup’s models and measured their speed, power efficiency and performance on devices.
“They’re really evaluating our technology right now,” Hassibi said of Apple.
He called the talks very early and said it was still unclear where they would lead but that “things are progressing well.”
Apple did not immediately respond to a request for comment.
The Information previously reported on the PrismML breakthrough.
The release comes a day after Apple released the iOS 27 public beta, giving iPhone owners full access for the first time to the company’s long-delayed overhaul of Siri. Apple is trying to make Siri more competitive with assistants OpenAI And Anthropocene At the same time, more personal information and AI processing remains on the device.
The company’s approach could address one of the key hurdles facing Apple’s AI strategy. The most powerful models typically require too much memory and processing power to run on a smartphone.
Apple can send complex requests to cloud-based models, but running more AI directly on the iPhone would reduce the delay associated with sending data to a remote server, reduce cloud computing costs and support the company’s privacy policy. It would also allow certain features to work without an internet connection.

Carolina Milanesi, president and principal analyst at Creative Strategies, said smaller models could allow Apple to bring more sophisticated features to the iPhone, including computational photography, video generation and health or fitness tools that rely on sensitive personal data.
“The more you can do on the device, the better it is,” she said, pointing to health and medication data that users would prefer to keep private.
PrismML said it shrinks AI models by dramatically simplifying the way their internal information is stored – reducing each 16-bit value to just one or three possible values. This significantly reduces the memory requirement for storing and operating the model.
Hassibi compared it to the chip industry’s transition from 8-bit to 4-bit computing, but takes it a step further.
The startup said the compressed models use between 10 and 15 times less memory, generate responses six to eight times faster and use three to six times less energy than traditional versions running on existing hardware.
However, Hassibi acknowledged that there is a compromise. PrismML’s models typically lose a few percentage points of overall performance, with factual memory weakening ahead of skills like reasoning, math and coding, he said.
PrismML releases two compressed versions of the model for free. They are designed to run on everyday devices including iPhones, MacBooks and Nvidia PCs.
The technology emerged from Hassibi’s research group at Caltech. The university owns the underlying patents and licenses them exclusively to PrismML. In March, the company raised a $16.25 million seed round backed by Khosla Ventures and other investors.
Hassibi said GoogleGemma’s open source model is next in the pipeline, followed by much larger models, including those from frontier labs that now generally require data center hardware.
PrismML says the technology could ultimately extend well beyond phones and laptops to robotics, autonomous systems and other products that need to make decisions quickly without relying on a cloud connection.
“It is very important that intelligence is local and can work quickly,” he said.

Table of Contents
Apple’s on-device advantage
Apple already runs parts of its AI system locally, including translations, some summaries and features that are closely tied to personal information. More complex requests are routed to Apple’s private cloud infrastructure or external models.
Horace Dediu, founder of Asymco, said Apple is likely trying to keep the majority of common Siri interactions on-device, reserving the most demanding tasks to the cloud.
The advantage is not simply in using less memory, but in incorporating a more powerful model within the same physical limitations.
“They’re trying to figure out how big the model is and how clever the model is that they can fit on the device,” Dediu said. Storing common requests locally reduces latency, improves data protection, and potentially reduces licensing and cloud costs.
Apple could have an advantage in using these models because the company co-develops the iPhone’s chips and software, giving it more granular control over how the AI runs on the device.
However, analysts warned that PrismML’s claims have yet to be proven outside of controlled demonstrations.
Tarun Pathak, research director at Counterpoint Research, said the model’s performance on long prompts, battery consumption when multitasking and reliability across millions of requests will be critical.
“The ultimate test will involve millions of queries, thousands of device combinations and robust testing at scale,” Pathak said.
Phil Solis, who leads IDC research on client processors, said power consumption may be the biggest unanswered question. A model capable of being used frequently – or continuously running in the background for agent-like tasks – can drain a phone’s battery, even if it uses less memory.

What it means for chip demand
PrismML’s release also comes amid intense debate over whether improvements in AI efficiency could ultimately reduce demand for memory chips and expensive data center infrastructure.
Storage has become one of the biggest constraints and costs in consumer electronics and AI servers. Morgan Stanley Estimates suggest Apple’s average dynamic random access memory cost per bit could increase by about 190% year-over-year in fiscal 2027, with NAND costs increasing by about 180%. NAND is typically used in flash drives and solid-state drives.
The company expects Apple to increase the starting price of comparable iPhone 18 models by about $200 to protect margins.
PrismML said its approach could enable the execution of a cloud model that typically requires eight GPUs to run on one, while allowing models that previously required a server to be moved to phones and laptops.
This could reduce the amount of memory or computing capacity required for a particular AI task. However, this does not necessarily mean that overall chip demand will decline.
Gil Luria, an analyst at DA Davidson, said shrinking models won’t eliminate the need for processors or memory. It could simply move more of those chips out of data centers and into phones and other devices.
“It’s not that you won’t need the chip,” Luria said. “You’re still going to need the GPU, and you’re still going to need the memory.”
He added that running AI on individual devices may actually be less efficient than using shared data center infrastructure because chips in phones may remain idle most of the time.
Efficiency breakthroughs can also lead to more usage rather than less spending, as cheaper and faster AI enables new products and encourages consumers to use models more frequently.
Still, the market has quickly discounted anything suggesting that AI may require less memory than expected. micron Shares plunged in March after Google released its TurboQuant paper on reducing memory usage without affecting model performance, although the stock later recovered.
The public release of PrismML gives everyday users and investors the opportunity to test whether the claimed gains hold up outside of the lab. And for Apple, running more powerful AI directly on the iPhone could help the company improve Siri without sacrificing the privacy and hardware integration that sets its products apart.
“The combination of cloud and on-device AI can deliver a richer, more efficient and more privacy-focused AI experience,” said Counterpoint’s Pathak. “Complex tasks are moved to the cloud, while sensitive, latency-critical and privacy-sensitive tasks are carried out on the device.”
REGARD: Apple is suing OpenAI for stealing trade secrets: What you should know
Choose CNBC as your preferred source on Google and never miss a moment from the most trusted name in business news.
https://www.cnbc.com/2026/07/14/apple-prismml-ai-compression-iphone.html
