Apple is reportedly in preliminary discussions with PrismML, a startup backed by Khosla Ventures, to integrate advanced model compression technology into its ecosystem. This move aims to bring sophisticated artificial intelligence directly onto the iPhone, bridging the gap between mobile hardware limitations and the massive computational requirements of modern Large Language Models (LLMs).

Solving the On-Device AI Constraint

As Apple rolls out iOS 18 and a redesigned Siri, the company faces a significant technical hurdle: high-performance AI models typically require massive amounts of memory and processing power that exceed standard smartphone capacities. By integrating PrismML’s technology, Apple seeks to move processing from the cloud to the device itself.

This shift is crucial for Apple’s strategic positioning. On-device AI reduces latency, lowers expensive cloud computing costs, and reinforces Apple's core privacy promise by ensuring sensitive user data remains on the handset. Furthermore, it allows key AI features to function even without an active internet connection.

How PrismML Shrinks Massive AI Models

PrismML, which originated from research at the California Institute of Technology, utilizes a sophisticated method of simplifying how a model's internal values are stored. While traditional methods might reduce figures from 16 bits to 4 bits, PrismML takes this a step further, reducing each figure to as few as one or three possible values.

The efficiency gains are substantial. PrismML recently demonstrated this by compressing Alibaba's open-source Qwen model from a massive 54 GB down to under 4 GB. This compression allows all 27 billion parameters of the model to operate on an iPhone 15 or newer. According to CEO Babak Hassibi, these compressed models offer several performance advantages:

  • Memory Efficiency: Uses up to 15x less memory.
  • Speed: Runs six to eight times faster.
  • Power Consumption: Consumes up to six times less energy.

While Hassibi noted a modest drop in performance—specifically regarding factual recall—the models maintain high proficiency in reasoning and coding abilities.

Implications for the Semiconductor and Hardware Markets

The potential adoption of such compression technology by a giant like Apple could shift the landscape for hardware manufacturers. Currently, analysts like Morgan Stanley predict that Apple’s memory costs could rise sharply by fiscal 2027 to support AI, which might lead to higher iPhone price tags.

However, if PrismML’s technology becomes a standard, the demand for massive memory modules in smartphones might be mitigated, even as the need for high-speed, efficient local processing increases. As PrismML moves toward compressing other major models like Google's Gemma, the goal remains clear: making high-level intelligence local, fast, and accessible on consumer hardware.

Key Takeaways

  • Massive Compression: PrismML can shrink a 54 GB AI model to under 4 GB, making it viable for high-end smartphones like the iPhone 15.
  • Efficiency Gains: The technology promises up to 15x less memory usage and 6x less energy consumption, which is vital for maintaining battery life during AI tasks.
  • Privacy and Speed: Moving AI from the cloud to the device will reduce latency and enhance user privacy, a cornerstone of Apple's product philosophy.

Risks and open questions

PrismML’s results are promising but still early-stage.

What to watch next

  • Official confirmation: Apple has not announced a partnership, so any press release or iOS 18 beta note about on-device LLMs will be a key signal.

Bottom line: PrismML’s compression could let iPhones host genuinely large language models without blowing up memory or draining the battery, delivering faster, private AI experiences.