Climbing the Memory Wall: How Apple, Google and Qualcomm are Driving On-Device AI Mainstream
- The memory wall from the bandwidth and cost perspective has been one of the key constraints in running LLMs efficiently and natively on the device.
- Google’s Gemma 4 changes this using advanced quantization techniques to reduce memory requirements and expand edge AI deployment.
- Apple, with its latest AFM 3 Core Advanced, uses a sparse model architecture that leverages NAND flash to reduce DRAM dependency.
- As token consumption continues to rise, a hybrid AI architecture is the way to go with its balanced on-device and cloud-based AI processing.
The industry is facing challenges posed by the ongoing memory crisis, as modern foundation models require substantial memory just to load before inference begins, while larger parameter models consume significantly higher memory capacity for running the model weights. Most devices run within fixed memory envelopes, and a model that demands more memory will create bottlenecks and reduce the device’s processing speed.
According to the latest Counterpoint data, GenAI smartphone penetration reached 36% by the end of 2025, but it is not expected to reach the mainstream or mass market anytime soon. As the smartphone industry transitions from image, text and voice-based GenAI interactions toward multimodal and agentic AI applications, memory continues to be one of the key constraints. Memory bandwidth growth is lagging behind NPU performance improvements, and the rising DRAM costs are limiting the deployment of advanced on-device AI use cases to mainstream devices.

To address this challenge, both Apple and Google, around their recent developer conferences, have discussed and explored ways to use various techniques to fit LLMs in lower memory configurations.
Google’s unit DeepMind has introduced the latest Gemma 4 with edge-optimized quantization techniques that significantly reduce model size and memory requirements, enabling AI models to run more efficiently on resource-constrained devices. Google has been working closely with compute vendors such as Qualcomm to test this and bring it to the market.
Google tackles memory bottleneck with quantization
Gemma 4 Models’ Memory Requirements

Gemma 4 incorporates quantization into the training, which means the model doesn’t need larger memory for execution, and the 2-billion-parameter model can even run on 1GB memory. The benefit of quantization is that it will lead to wider deployment of on-device GenAI across devices, especially mid-range smartphones, reducing the need for cloud-based execution of the model. Gemma 4 can execute basic tasks on the device, whereas complex tasks can be offloaded to the cloud.
Google’s strategy can be compared to Qualcomm’s introduction of INT2 inferencing support. However, these two technologies operate at different layers of the AI stack. Gemma 4 is a model-level optimization, while Qualcomm’s INT2 is a hardware-level capability that allows Snapdragon NPUs to execute 2-bit models more efficiently.
In the smartphone industry, memory has emerged as a bigger bottleneck than compute in enabling advanced AI capabilities. Gemma 4 addresses this challenge by reducing memory requirements, paving the way for more on-device and agentic AI experiences. This is great as it expands the AI-driven TAM for Google and the entire Android ecosystem. It will be keenly watched how app developers deploy this in their applications and dynamically switch based on device capabilities.
On the other hand, Apple, which has hundreds of millions of devices with capable Apple Silicon compute, is constrained by a relatively low memory footprint when running AI natively and efficiently. Apple is also working on fitting and running LLMs, leveraging memory creatively. For example, Apple’s new model architecture is designed in a way that the model weights are stored in NAND and only the active parameters are dynamically loaded into DRAM for execution.
Apple’s WWDC 2026 moment: Turning AI into system experience
At WWDC 2026, Apple released the third generation of the Apple Foundation Model (AFM), a family of five purpose-built AI models, optimized for on-device and cloud-based deployments. Among them, two models are optimized for on-device capabilities, whereas three are optimized for cloud-based deployments, running on Apple’s Private Cloud Compute, which offers utmost privacy to the users. The newer generation of AFM follows an innovative architecture, where instead of DRAM, the model is stored in NAND, enabling a larger 20-billion-parameter model to run on consumer hardware.
The AFM family consists of five models:
- AFM 3 Core: It’s a 3-billion-parameter model capable of running on the device.
- AFM 3 Core Advanced: This is a 20-billion-parameter model with multimodal capabilities and running on the device. The model uses a sparse architecture, activating 1 billion to 4 billion parameters at a time depending on the user request.
- AFM 3 Cloud: This is a cloud-based model optimized for speed, efficiency and performance.
- ADM 3 Cloud (Image): It is a cloud-based model for image generation and editing, offering Image Playground and other capabilities.
- AFM 3 Cloud Pro: It is a cloud-based model for complex tasks such as agentic use cases and complex reasoning.
AFM 3 Core, AFM 3 Core Advanced, AFM 3 Cloud and ADM 3 Cloud (Image) are optimized and purpose-built for Apple Silicon, whereas AFM 3 Cloud Pro Apple has partnered with Google Cloud and NVIDIA to extend Apple’s Private Cloud Compute to NVIDIA’s GPU on the Google Cloud platform.
AFM 3 Core Advanced Model Architecture

One of the major features of AFM 3 Core Advanced is that it utilizes NAND alongside DRAM. Traditional LLMs require all weights to reside in DRAM, which limits the DRAM usage for other purposes. AFM 3 Core Advanced introduces a newer architecture built on Instruction Following Pruning (IFP), developed by Apple researchers. Here, the full model is stored in NAND, and dynamically loads only the active parameters into DRAM, based on a predetermined number of active parameters tailored to each specific use case. This ensures reduced DRAM latency and an edge over competitors.
Apple has not announced external benchmark comparisons for its third-generation AFM models, which limits the ability to verify the newer AFM models' performance comparison. Users are becoming more accustomed to benchmark comparison and model capabilities, and the challenges with previously-launched Apple Intelligence mean users will take time to change their perception of Siri AI. Further, the developer uptake of these models will decide the success and adoption of Siri AI. So, it is not limited to basic AI use cases and is therefore more useful; otherwise, Apple users will continue to spend more time in OpenAI, Google or Anthropic ecosystems.
Implications
The AI industry is finally moving towards hybrid architecture, where the cloud and on-device AI combination can serve a more complete, efficient and privacy-centric AI experience, depending on the scenarios. Complex tasks will be offloaded to the cloud, whereas sensitive, latency-sensitive and privacy-relevant tasks will be executed on the device.
Google has also structured Gemma 4 as the foundation for Gemini Nano, which will make the model deeply integrated into the Android ecosystem, and more users will be able to run Gemini Nano on their devices, especially mid-tier devices, as currently only flagship devices can support Gemini Nano.
Both Google’s Gemma 4 and Apple’s third-generation AFM will lead to wider adoption of on-device GenAI capability. Gemma 4 expands the addressable devices for AI capabilities by reducing the memory bottleneck, whereas Apple's utilization of NAND will make AI available at lower DRAM. Apple’s IFP architecture, introduced with AFM 3 Core Advanced, will help to greatly reduce the DRAM size required for on-device GenAI capabilities, and will enable more devices to offer on-device LLMs with lower DRAM. The ongoing memory crisis has led to a significant increase in DRAM prices and led to memory shortage.
Apple’s decision to run AFM 3 Cloud Pro only on Google Cloud and NVIDIA GPUs indicates that its silicon can run most of the AI-related tasks, showcasing its increasing power and capability.
Finally, in this AI race, Google is definitely running with a more vertical and advanced hybrid AI architecture and solutions, whereas Apple is starting this long journey as it moves to GenAI and, in the future, potentially (re)architects for native Agentic AI, building upon the OpenClaw success across its Macs.
Category
Industry
Smartphone
Service
Standard
Report Type
Report
Time Period
Other
Receive our insightful weekly newsletter and stay ahead of the competition.
Author
Soumen Mandal
Soumen is a Senior Analyst tracking IoT, Automotive and Telecommunication ecosystem at Counterpoint Research. He is interested in IoT applications, connections, components, electric vehicles, connected cars, autonomous vehicles, semiconductors, shared mobility, services and emerging technologies. He started his career as an Energy Analyst with Manikaran Power Ltd. He has experience working with DISCOMs and SLDCs in the Indian power and energy industry. He is currently based in Gurgaon. He holds an Electrical Engineering degree and an MBA in Marketing & Finance.