Running a language model on a phone or laptop is not a smaller version of calling a server. On a server, inference is a stateless request: the model is already loaded, the GPU is dedicated, and nobody minds if a process restarts. Inside an app, the model shares memory with everything else the user is running, the operating system can suspend or kill the process at any moment, the chip slows down as it heats up, and the person holding the device expects the keyboard to stay responsive while tokens stream in.
This article is about the engine inside the app: the code between your UI and the inference library. It covers loading and warm-up, the prefill and decode loop and where it runs, streaming and cancellation, what to do when the context window fills, saving and restoring the KV cache, and how to respond to memory pressure, backgrounding and heat. The budget arithmetic for weights and KV cache is covered in SLMs on mobile in 2026 and the bytes-per-token bound on decode speed in SLM inference optimization, so this page cites them and spends its words on behaviour over time.
The parts of an on-device engine
Whatever runtime you use, the same components appear. A weights file, usually quantized to 4 or 8 bits, holds the model. A tokenizer turns text into integer ids and back. A compute backend runs the matrix multiplications on the CPU's vector units, the GPU, or an NPU, often splitting layers between them. A KV cache stores the attention keys and values for every token already processed, so each new token only needs its own computation. A sampler turns the final logits into a chosen token. A session ties a KV cache to a conversation.
The design decision that matters most is ownership. Exactly one worker thread or task should own the model, the backend and every KV cache. The UI sends it requests and receives tokens through a queue. If two threads can call into the engine, you will eventually get two generations interleaving writes into one cache, and the corruption shows up as fluent nonsense rather than a crash.
Two phases and one constraint
Generation has two phases. Prefill processes the whole prompt in parallel; it is limited by compute and its cost grows with prompt length. Decode produces one token at a time; each step reads essentially all the weights once, so it is limited by memory bandwidth. As a rough guide from the linked arithmetic, a 3-billion-parameter model at 4 bits needs about 1.5 GB for weights plus roughly 115 KB of 16-bit KV cache per token, and its decode speed cannot exceed usable memory bandwidth divided by the bytes read per token.
A device adds a third constraint: sustained performance. Phones run at peak clock briefly, then throttle to stay within their thermal envelope, so a ten-second benchmark measures the burst while real users live in the throttled state. Plan with numbers measured over minutes.
Loading and warm-up
Most runtimes memory-map the weights file instead of reading it into allocated memory. Mapping returns almost instantly, but no data has been read yet. The first forward pass touches every weight page, and each first touch is a page fault that reads from flash storage. That is why the first token after a cold start can take several seconds while later ones are fast: you are paying for I/O, not arithmetic. Mapped, unmodified pages are also clean memory that the OS can drop under pressure and fault back in later, which is better for survival than private allocations but means a later generation can suddenly slow down.
Treat loading as a measured step with three timestamps, and run the warm-up when the user shows intent, such as opening the screen that uses the model, not at app launch.
import time
def load_and_warm(engine_factory, model_path, log):
t0 = time.monotonic()
engine = engine_factory.map_weights(model_path) # mmap, near-instant
t1 = time.monotonic()
engine.init_backend() # buffers, kernels, shader compile
t2 = time.monotonic()
engine.generate_tokens("Hello", max_tokens=1) # touches every weight page once
t3 = time.monotonic()
log("model_load", map_ms=(t1 - t0) * 1e3, backend_ms=(t2 - t1) * 1e3,
first_token_ms=(t3 - t2) * 1e3, model=model_path)
return engineThe engine methods belong to a hypothetical interface; map them onto your runtime. Log the three durations separately because their fixes differ: slow mapping means storage, slow init means cacheable shader or graph compilation, and slow first token means page faults.
The generation loop on a worker
The loop below shows the shape every on-device engine needs: prefill in chunks so cancellation is never more than one chunk away, a decode loop that checks a cancellation flag every step, tokens pushed to the UI through a thread-safe queue, and explicit stop conditions. Text is decoded incrementally because one token can be part of a multi-byte character, so you emit only complete text.
import queue, threading
class Cancelled(Exception):
pass
def run_generation(engine, session, prompt_ids, params, out: queue.Queue, cancel: threading.Event):
try:
# Prefill in chunks: bounded latency to notice cancellation, bounded scratch memory.
for i in range(0, len(prompt_ids), params.prefill_chunk):
if cancel.is_set():
raise Cancelled()
engine.prefill(session, prompt_ids[i:i + params.prefill_chunk])
detok = engine.incremental_detokenizer()
for step in range(params.max_new_tokens):
if cancel.is_set():
raise Cancelled()
if session.n_tokens >= session.n_ctx:
out.put(("stop", "context_full"))
return
logits = engine.decode_step(session)
tok = engine.sample(logits, params) # temperature, top-p, grammar mask
if tok in params.stop_ids:
out.put(("stop", "eos"))
return
engine.append(session, tok)
text = detok.push(tok) # "" until a full character is ready
if text:
out.put(("text", text))
out.put(("stop", "max_tokens"))
except Cancelled:
session.truncate_to_last_committed() # leave the cache consistent
out.put(("stop", "cancelled"))
except MemoryError:
out.put(("error", "out_of_memory"))Three details are easy to get wrong. Cancellation must leave the KV cache consistent: roll a half-applied prefill back to the last committed position, or the next turn attends to a fragment the user never sent. The UI should render once per frame, not once per token. And the worker should use only the performance cores; adding efficiency cores often slows decode, because each step waits for the slowest thread.
When the context window fills
A session's KV cache has a fixed capacity, the context length you allocated. Long chats reach it. You have four options, and you should pick one deliberately rather than let the runtime pick for you.
| Policy | What happens | Cost | Use when |
|---|---|---|---|
| Stop | Report context_full; the user starts a new chat | None | Short tasks, extraction, tools |
| Shift | Keep the system prompt, drop the oldest turns, keep generating | Cheap if the runtime can shift positions; otherwise a re-prefill | Casual chat where old turns matter less |
| Summarise | Generate a summary of old turns and rebuild the session from system prompt plus summary plus recent turns | One extra generation and one re-prefill | Assistants where earlier facts must persist |
| Retrieve | Store turns outside the cache and re-insert only relevant ones per request | Embedding and search per turn | Long-lived notes or document chat |
Whatever you choose, never drop the system prompt or the start of the conversation silently. Models rely on the first tokens heavily, and evicting them tends to degrade output much more than evicting middle turns. When shifting, free a margin of several hundred tokens at once rather than one token per step, and log every time a policy fires, because the summarise path quietly doubles the work for long sessions.
Saving and restoring KV state
Prefill is the most expensive part of resuming a conversation. If a user returns to a chat with 3,000 tokens of history, re-prefilling all of it costs seconds on a phone. Most runtimes can serialise a session's KV cache and position to bytes. Writing that snapshot when the app moves to the background, and reading it back on return, turns a multi-second rebuild into a file read.
A snapshot is valid only for the exact model file, quantization, context length and runtime version that produced it, so store a header with a weights hash and those parameters and discard mismatches. Write to a temporary file and rename it, so a killed process never leaves a truncated snapshot that looks valid. At roughly 115 KB per token a 3,000-token session exceeds 300 MB, which may be slower to write than to recompute; measure both, and cap what you keep on disk.
A cheaper and almost always worthwhile form of reuse is a fixed prefix. If every request starts with the same system prompt and tool descriptions, prefill it once after load and copy or fork that state for each new session. This is the on-device version of prefix caching.
Living with the operating system
Mobile operating systems reclaim memory by killing processes. iOS terminates apps that exceed their memory limit or that are in the background when memory runs low, and Android's low-memory killer does the same, preferring background processes. Both platforms deliver warnings first: iOS sends memory warnings to the app, and Android delivers trim-memory callbacks. Treat those signals as instructions: cancel any running generation, save the session snapshot if one is due, and release the model if the screen using it is not visible. A model you can reload in two seconds is better than a process the user finds restarted.
Backgrounding is the second hazard. On iOS, apps generally cannot submit GPU work while in the background, and on both platforms background execution time is limited, so cancel GPU generation on leaving the foreground and make it resumable: save the prompt and completed output so the conversation continues on return.
Heat is the third. iOS exposes a thermal state through ProcessInfo.thermalState and Android exposes thermal status through PowerManager. When the device reports a serious state, shorten responses, lower the maximum output length, pause background work such as embedding indexing, or route the request to a smaller model.
A worked session, end to end
Follow one user through a writing-assistant feature using an illustrative 3B model at 4 bits on a recent phone. The numbers are plausible orders of magnitude for planning, not measurements of a specific device.
- The user opens the editor. The app maps the weights, initialises the backend and runs a one-token warm-up. Mapping takes milliseconds; the warm-up takes a couple of seconds because it faults in about 1.5 GB of pages. The fixed 400-token system prompt is prefilled once and kept as a reusable prefix.
- The user asks for a rewrite of a 1,200-token paragraph. The session forks the prefix state, prefills the paragraph in chunks of 256 tokens, and streams the first word after well under a second of prefill on a fast phone. Decode runs on the worker using the performance cores; the UI renders once per frame.
- Halfway through, the user taps stop. The worker sees the flag at the next decode step, emits cancelled, and the cache stays consistent at the last appended token.
- The user switches to another app. The memory warning arrives a minute later. The session snapshot, about 230 MB for 2,000 tokens at 16 bits, is written to a temporary file and renamed, and the model is released.
- The user returns. The app remaps the weights and restores the snapshot after checking its header. The next request continues the conversation without re-prefilling 2,000 tokens. The device is warm, the thermal state reports elevated, and the engine lowers the maximum output length for the next response.
Failure modes
- Slow first token every time. The model is released and remapped too eagerly, or the warm-up runs at launch and its pages are evicted before use. Log the three load durations and tie warm-up to user intent.
- Decode that slows down mid-session. Weight pages were evicted under memory pressure and are being faulted back in, or the device has throttled. Record tokens per second per step, page-fault counts where the platform exposes them, and the thermal state.
- Fluent nonsense after cancel or resume. The KV cache was left half written, or a snapshot from a different model or runtime version was restored. Roll back on cancel and validate snapshot headers.
- Process killed during generation. The budget was planned for a flagship, or several sessions held caches at once. Keep one active session and cancel on memory warnings.
- Silent quality loss in long chats. The runtime's default context shift dropped the system prompt. Pick a policy and log when it fires.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Context length | Long context: fewer resets, much more KV memory and slower prefill | Short context plus summarise or retrieve: smaller footprint, extra work and some loss |
| Backend | GPU or NPU: faster prefill, foreground only on some platforms, shader compile cost | CPU: predictable and available in more states, slower prefill |
| Resume | Snapshot KV to disk: fast resume, large files, versioning burden | Re-prefill from text: simple and robust, seconds of compute |
| Model lifetime | Keep loaded: instant responses, high kill risk in background | Release when hidden: reload cost, process survives |
To go deeper on the components the engine depends on, read the KV cache and quantization for small models.
What to do next
- Put the model behind a single worker that owns the backend and every session, and make the UI talk to it only through a request and token queue.
- Instrument loading as three separate durations and move warm-up to the moment the user shows intent.
- Make prefill chunked and decode cancellable every step, and test that a cancelled session produces sensible output on the next turn.
- Choose a context-full policy per feature, keep the system prompt pinned, and log every time the policy fires.
- Prototype KV snapshot save and restore with a validated header, then measure write time against re-prefill time on your slowest supported device.
- Wire memory warnings, backgrounding and thermal state into cancellation, release and output-length decisions, and run a 10-minute sustained test before shipping.