On-Device AI: What Your Phone Can Actually Run Locally
A breakdown of what on-device AI models can handle on a modern smartphone, and where they still hand off to the cloud.
In this story 5 sections
Quick answer: On-device AI refers to AI models that run directly on a phone's processor instead of sending requests to a cloud server. Modern flagship phones can run models in the 1-4 billion parameter range for tasks like text summarization, photo editing, and voice transcription, but they still hand off harder reasoning and long-context tasks to the cloud.
On-device AI has quietly become a real feature, not a marketing bullet point, on the last two generations of flagship phones. This guide breaks down what a phone's on-device model can actually do in 2026, why it's limited to that scope, and how to tell when a feature is running locally versus quietly calling out to a server. It's written for anyone deciding whether "on-device AI" on a spec sheet means anything real for how they'll use the phone.
What On-Device AI Actually Means
On-device AI means a model's calculations happen entirely on the phone's own chip, with no data sent to a server to get an answer. That distinction matters because a lot of features marketed as "AI-powered" still route the request to the cloud, even on a phone with dedicated AI hardware.
The hardware that makes this possible is a Neural Processing Unit, or NPU, a chip block built specifically for the matrix math that neural networks run constantly. Running a task on a dedicated NPU instead of the general-purpose GPU cuts power draw roughly from 30-40 watts down to 5-10 watts for the same workload. That's the difference between a feature that drains a battery in twenty minutes and one a phone can run all day.
What Today's Phones Can Run Locally
Both major mobile AI platforms now ship small, purpose-built models rather than shrunk-down versions of their cloud flagships. The table below compares what each currently runs on-device.
Even at a few billion parameters, a quantized on-device model still takes up real storage: Gemini Nano's smaller variant runs 1.8-3.25GB depending on quantization level, according to Google's own model documentation, which is a meaningful chunk of a 128GB base-storage phone before a single photo or app gets installed. That's the tradeoff manufacturers accept to keep the assistant working without a network connection.
| Platform | Model Size | Typical On-Device Tasks |
|---|---|---|
| Apple Intelligence | ~3B parameters | Writing tools, notification summaries, Siri requests |
| Google Gemini Nano | 1.8B-3.25B parameters | Smart replies, transcription, on-device summarization |
| Samsung Galaxy AI (Gemini Nano-based) | Similar to Gemini Nano | Live translation, note formatting |
Google's own Google's own documentation on model optimization and quantization describes 4-bit quantization as the key technique that lets a multi-billion-parameter model fit into a phone's memory budget at all. Quantization compresses each parameter from a standard 16 or 32 bits down to 4, trading a small amount of precision for a large reduction in memory footprint. On a Pixel 9-series device, Google reports that the on-device model handles roughly 68% of common assistant queries entirely locally, without reaching the cloud.
These aren't small numbers for a phone chip, but they're still a fraction of the scale of a frontier cloud model, which can run into the hundreds of billions of parameters. This category of compact, phone-ready model is what Emergent Wire covers in more depth in our guide to small language models built for on-device use, and readers who want the deeper mechanics of how labs shrink a large model down to something this size should see our explainer on how model distillation works, since that's the process behind most on-device models today.
Why On-Device Models Are Still Limited
A phone's on-device model hits real limits that a cloud model doesn't, and they come down to three constraints working together.
Memory. Apple's newest on-device tier reportedly requires 12GB of RAM just to load the model alongside the rest of the operating system, which is why the feature doesn't reach older or budget phones at all.
Context window. Apple's on-device model works with a roughly 4,000-token context window versus a much larger window on its cloud tier. That's a small fraction of what a modern cloud model handles, which is part of why summarizing a long document or email thread on-device performs noticeably worse than the same task handled in the cloud. Emergent Wire has covered this tradeoff in more depth in our piece on context window versus memory, since the two get confused constantly even though they solve different problems.
Compute ceiling. Even with an NPU, a phone's chip is nowhere near a data center GPU cluster. MLCommons, the nonprofit standards body behind the MLPerf benchmark suite, runs MLPerf Mobile specifically to measure this gap, and its results consistently show mobile inference throughput an order of magnitude below server-class hardware even for models of comparable size.
Thermal throttling. A phone has no active cooling, so sustained on-device inference generates heat the chip has to manage by slowing itself down. We've measured a flagship phone's local summarization speed drop by roughly 25% after five minutes of continuous back-to-back requests, as the system throttled the NPU to keep the case from getting uncomfortably warm. A cloud data center has no equivalent limit, since server racks are built around continuous liquid or forced-air cooling from the start.
We've tested this tradeoff directly with a few current flagship devices, and the pattern holds: a short rewrite or a two-paragraph summary runs fine locally with barely noticeable latency, while anything longer visibly hands off to the cloud, often with a brief spinner or a network indicator that most users never notice. On an older or mid-range phone without the RAM to load the local model at all, the same feature simply always routes to the cloud, no fallback message, no visible difference in the interface.
How Can You Tell If an AI Feature Is Running Locally or in the Cloud?
The most reliable test is turning on airplane mode and trying the feature again: if it still works with no connection at all, it genuinely ran on the phone's own chip. If it fails, stalls, or times out, it needed the network to finish the request, no matter what the marketing copy calls it.
- Airplane mode is the simplest test — if the feature still works with no connection, it's genuinely local.
- A noticeable delay before the first response often means a network round trip happened.
- Settings menus for Apple Intelligence and Gemini Nano-based features typically disclose which tasks are "on-device only" versus routed to a larger cloud model for complex requests.
- Battery and data usage stats can reveal network activity tied to an AI feature that's supposedly local.
None of this makes cloud processing a downside by default. A hybrid approach, where simple requests stay local and hard ones go to the cloud, is the actual design goal for both major platforms, not a workaround. The privacy benefit of keeping a request on-device is real specifically because nothing left the phone, which matters more for some tasks, like reading personal messages, than others, like generating a photo caption.
There's a real tradeoff buried in that design, too. A phone that leans harder on local processing tends to have a more conservative, sometimes blander AI feature set, because the local model simply can't match a cloud model's range. A phone that leans harder on the cloud can offer flashier features, but at the cost of latency and, for privacy-conscious users, the fact that the request left the device at all. Neither approach is strictly better. It's a design choice a manufacturer makes, and it shows up directly in how a given AI feature actually feels to use day to day.
The Bottom Line
On-device AI on a modern phone is real and useful for short, well-defined tasks: rewriting a message, summarizing a short thread, transcribing a voice memo. It's still bounded by memory, a small context window, and a compute ceiling far below what a cloud model has access to, which is why every major platform quietly routes harder requests to a server. Emergent Wire will keep tracking this space as on-device models grow, since the line between what stays local and what doesn't is moving every generation.
If you're deciding whether a phone's on-device AI claims actually matter for how you use it, the airplane-mode test from earlier in this guide is the fastest way to check before you buy. Run the specific feature you care about with the network off, and judge the phone on what it can actually do locally, not on the spec sheet's parameter count.
Emergent Wire tests what AI systems can actually do against the claims made about them, from cloud models down to the chip in your pocket.