In 2026 a phone app that wants a language feature, such as summarising a note, rewriting a message or pulling structured fields out of a receipt, has a real choice to run it on the device. Both major mobile platforms now ship a system model that apps can call, and open runtimes can execute a bundled small model on the CPU, GPU or NPU. The choice is no longer whether it is possible but which path, for which users, with what fallback.
This article is about the engineering around that choice. Hardware internals are covered in NPU acceleration, and generic edge deployment in edge deployment architecture. Here we start from the physics that limits phones, compare the OS-provided and bundled paths with real APIs, work through memory budgets, delivery and routing, and finish with a benchmarking method and a checklist. API details were checked against platform documentation at the end of September 2026; they change often, so confirm them against current docs before you ship.
The physics: memory, bandwidth and heat
Generating each token requires reading essentially all of the model's weights from memory once. On a phone that makes decode speed a memory-bandwidth problem, not a compute problem. A useful upper bound is tokens per second at most equal to memory bandwidth divided by bytes read per token. Take a 3-billion-parameter model quantized to 4 bits: about 1.5 GB of weights. On a phone whose memory system sustains, say, 60 GB/s for this workload, the bound is about 40 tokens per second, and real throughput is lower because of KV-cache reads, activation traffic and the fact that other apps and the display share the same memory. Halving the bits per weight roughly doubles the bound, which is why aggressive quantization matters more on phones than anywhere else; SLM quantization covers the methods.
Prefill, processing the prompt, is different: it is compute bound and parallel, so NPUs and GPUs shine there, and time to first token scales with prompt length. Two further limits dominate product behaviour. The operating system kills apps that use too much memory, and the threshold depends on device and what else is running, so a model that loads on a test phone can be killed on a user's phone with ten tabs open. And phones throttle when hot: a benchmark that runs for ten seconds measures the burst, not the sustained speed a user sees after a few minutes of use.
Two paths: the system model or your own
| OS-provided model | Bundled model | |
|---|---|---|
| Examples | Apple Foundation Models framework; Android ML Kit GenAI APIs on Gemini Nano via AICore | ExecuTorch, LiteRT-LM, llama.cpp, MLC LLM running a model you ship |
| Download cost to your app | None; the model is shared by the system | Hundreds of MB to a few GB per install |
| Device coverage | Only devices and regions where the system model is available | Any device with enough memory, at varying speed |
| Model choice | Fixed by the platform; updated on the platform's schedule | Yours, including fine-tunes and custom tokenizers |
| Customisation | Instructions, guided output, tools; adapters where supported | Full: fine-tuning, quantization choices, sampling |
| Behaviour stability | Can change with an OS update | Pinned to the version you ship |
| Limits | Context and quota limits set by the platform | Your memory and thermal budget |
The pragmatic default in 2026 is to use the system model where it is available and good enough for the task, and to bundle a model only for tasks that need a specific model, need coverage on devices without a system model, or cannot tolerate behaviour changes arriving with OS updates.
Apple: the Foundation Models framework
Apple's on-device model is about 3 billion parameters, uses KV-cache sharing between blocks to cut memory and time to first token, and is compressed with 2-bit quantization-aware training, according to Apple's 2025 foundation models technical report. The Swift Foundation Models framework exposes it with sessions, guided generation into Swift types, and tool calling. Always check availability first: the model is absent on unsupported hardware, when Apple Intelligence is turned off, or while the model is still downloading.
import FoundationModels
@Generable
struct ReceiptFields {
@Guide(description: "Merchant name as printed")
var merchant: String
@Guide(description: "Total amount in the receipt's currency")
var total: Double
}
func extract(_ receiptText: String) async throws -> ReceiptFields? {
let model = SystemLanguageModel.default
guard case .available = model.availability else {
return nil // route to the bundled or cloud path
}
let session = LanguageModelSession(
instructions: "Extract fields from the receipt. Do not invent values.")
let response = try await session.respond(to: receiptText,
generating: ReceiptFields.self)
return response.content
}Guided generation constrains decoding to the structure of the type, which removes a whole class of parse failures; the mechanism is explained in guided decoding. The context window was 4,096 tokens per session at launch, covering instructions, prompt and output together, and iOS 26.4 added a contextSize property and a token counting method, so read the limit at runtime instead of hardcoding it. Long inputs must be chunked, and the framework raises an error when a session exceeds its context, which your code should catch and handle by summarising history or starting a new session.
Android: ML Kit GenAI and the Prompt API
On Android, AICore is a system service that runs Gemini Nano on supported devices, and the ML Kit GenAI APIs sit on top of it. There are task APIs for summarisation, proofreading, rewriting and image description, and a general Prompt API. Because the model is shared by the system, an app does not ship it, but it may need to trigger a download the first time.
val generativeModel = Generation.getClient()
suspend fun summarise(note: String): String? {
when (generativeModel.checkStatus()) {
FeatureStatus.UNAVAILABLE -> return null // route elsewhere
FeatureStatus.DOWNLOADABLE -> {
generativeModel.download().collect { status ->
if (status is DownloadStatus.DownloadCompleted) Log.d(TAG, "model ready")
}
}
else -> Unit
}
val response = generativeModel.generateContent(
"Summarise this note in three bullet points:" + note)
return response.candidates.firstOrNull()?.text
}The documentation sets limits you must design around: keep input under about 4,000 tokens, avoid use cases that need more than about 4K tokens of output, the API requires Android API level 26 or higher, it is not supported on devices with an unlocked bootloader, and AICore enforces an inference quota per app, so batch jobs that loop over hundreds of items will hit errors. Device support is limited to specific recent models; treat availability as a runtime fact, never a build-time assumption. A streaming variant, generateContentStream, returns chunks for responsive UIs.
Bundling your own model
Bundle a model when you need one the platform does not provide, a fine-tune for your domain, or broad device coverage. The runtimes differ mainly in how they reach accelerators. ExecuTorch, PyTorch's on-device runtime, reached 1.0 and exports models ahead of time into a program with delegates for backends such as Vulkan GPUs and Qualcomm NPUs. LiteRT-LM is Google's orchestration layer for running language models on LiteRT across Android, iOS, web and desktop, with GPU and NPU acceleration. llama.cpp runs GGUF files on CPU with Metal, Vulkan and other GPU backends, and is the fastest path to a working prototype; GGUF runtimes explains that format. NPU delegates give the best energy per token but cover fewer operators and devices, so profile which parts of your graph fall back to the CPU.
Model size is the first decision. On current phones, models from roughly 0.5 to 4 billion parameters at 4 bits are the practical range for interactive features; a 1 to 2 billion parameter model fine-tuned for one task often beats a general 3 to 4 billion model on that task while using half the memory. Deliver weights separately from the app binary so the store download stays small, and fetch them on first use of the feature, preferably on unmetered networks, with resumable downloads, a checksum check before loading and a version in the path so updates never corrupt a working copy. Platform mechanisms for large asset delivery exist on both stores; use them where they fit your distribution.
The memory budget, worked
Memory is weights plus KV cache plus runtime overhead. The KV cache per token is 2 (keys and values) times layers times KV heads times head dimension times bytes per element. For an illustrative 3B model with 28 layers, 8 KV heads of dimension 128 and a 16-bit cache, that is 2 x 28 x 8 x 128 x 2 bytes, about 115 KB per token, so a 4,096-token context costs about 470 MB on top of about 1.5 GB of 4-bit weights. Quantizing the cache to 8 bits halves that. The KV cache covers the optimisations in depth.
def mobile_budget(params_b, weight_bits, layers, kv_heads, head_dim,
ctx_tokens, kv_bytes=2, overhead_mb=300):
weights_mb = params_b * 1e9 * weight_bits / 8 / 1e6
kv_mb = 2 * layers * kv_heads * head_dim * kv_bytes * ctx_tokens / 1e6
return round(weights_mb), round(kv_mb), round(weights_mb + kv_mb + overhead_mb)
print(mobile_budget(3.0, 4, 28, 8, 128, 4096)) # (1500, 470, 2270)
print(mobile_budget(3.0, 4, 28, 8, 128, 4096, kv_bytes=1)) # (1500, 235, 2035)
print(mobile_budget(1.5, 4, 28, 2, 128, 4096)) # (750, 117, 1167)Memory-mapping the weights file lets the OS page weights in and out rather than counting them all as private dirty memory, which improves survival under pressure, but pages evicted under pressure must be re-read, and decode speed collapses when that happens repeatedly. Budget for the second-worst device you intend to support, release the model when the feature's screen closes, and handle the low-memory callback by cancelling generation cleanly.
Routing: one task contract, several engines
Real apps rarely have one path. A router decides per request which engine runs it, based on availability, the input's token count against each engine's context, battery and thermal state, and the user's privacy settings. The key design rule is that every engine honours the same task contract: the same prompt template or instructions, the same output schema, the same evaluation set and the same telemetry fields. Without that, the feature behaves differently by device and nobody can tell why.
def route(task, text, device, policy):
tokens = estimate_tokens(text)
if device.os_model_available and tokens + task.max_output <= device.os_model_context:
return "os_model"
if device.bundled_model_ready and device.thermal_state in ("nominal", "fair") \
and not device.low_power_mode:
return "bundled"
if policy.cloud_allowed and device.online:
return "cloud"
if tokens > device.os_model_context and device.os_model_available:
return "os_model_chunked" # map-reduce over chunks
return "unavailable" # show a clear, non-AI fallbackLog the chosen route, token counts, time to first token, total latency and whether output validation passed, without logging user content. Those fields answer most support questions and show when an OS update changes behaviour.
Benchmark on real devices, for minutes
Emulators and desktop numbers tell you nothing about thermal throttling or memory pressure. Build a small device lab covering your top device classes by user count, including at least one older, lower-memory phone. For each engine and model, measure time to first token at your real prompt lengths, decode tokens per second, peak memory, energy per request from the platform profiler, and quality on a fixed evaluation set of a few hundred real task examples. Run each benchmark in a loop for ten minutes and report the last minute separately from the first, because sustained throughput after throttling is what users experience in long sessions. Re-run the whole suite on each OS update, since system models and drivers change underneath you.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Feature missing for most users | System model unavailable on their devices | Router with bundled or non-AI fallback; measure coverage |
| App killed mid-generation | Weights plus KV cache exceed the memory the OS allows | Smaller model, shorter context, 8-bit cache, release on exit |
| Fast in testing, slow in use | Thermal throttling after sustained load | Benchmark sustained; cap output; defer batch work to charging |
| Context errors on long inputs | Input exceeds the session context | Count tokens first; chunk and summarise |
| Quota errors in batch jobs | Per-app inference quota in the system service | Throttle, batch in the background, surface progress |
| Behaviour changed overnight | OS updated the system model | Evaluation set in CI against current OS builds; pin a bundled fallback |
| Corrupt model after update | Interrupted download overwrote the working copy | Versioned paths, checksum before load, atomic switch |
What to do next
- Define each language feature as a task contract: instructions, output schema, 200 or more evaluation examples, and telemetry fields.
- Try the system model on both platforms first, checking availability at runtime and handling download states.
- Measure device coverage from your analytics; decide whether uncovered users get a bundled model, cloud, or a non-AI path.
- If bundling, pick the smallest model that passes your evaluation, quantize to 4 bits, and compute the memory budget for your weakest supported device.
- Ship weights separately with resumable, checksummed, versioned downloads.
- Build the router and log route, tokens, latency and validation results without user content.
- Benchmark ten-minute sustained runs on real devices and re-run on every OS update.
- Read platform limits such as context size at runtime, and re-check the platform docs each release cycle.