TensorFlow Lite was built for small vision and audio models with fixed input shapes: load a file, feed a tensor, read a tensor. A language model breaks every one of those assumptions. Its input length varies from one token to thousands, it runs in a loop that feeds its own output back in, and it carries a growing cache of attention keys and values between steps. Yet the same runtime, now called LiteRT, is one of the main ways to run LLMs on phones and laptops, using the CPU, the mobile GPU and increasingly the NPU.
This article explains how that works from the runtime up: what is in a model file, how delegates hand parts of the graph to the GPU, how an LLM is cut into prefill and decode signatures with the KV cache as explicit tensors, how a PyTorch model is converted and quantized, and how the LiteRT-LM runtime drives the loop in an app. A worked memory budget shows what fits on a device. For the NPU hardware itself see edge NPUs, in depth; for app lifecycle and product choices see on-device LLM inference and SLMs on mobile in 2026.
TFLite, LiteRT, LiteRT-LM: what is what
The names changed recently, and older tutorials use the old ones. Google renamed TensorFlow Lite to LiteRT in 2024; the file format and much of the API carried over, and the TensorFlow Lite name still appears in packages and documentation. For LLMs specifically, the pieces are:
| Piece | Role | Status |
|---|---|---|
| LiteRT (formerly TFLite) | Runtime: loads .tflite flatbuffers, runs ops on CPU, GPU or NPU | Current |
| litert-torch (formerly AI Edge Torch) | Python library that converts PyTorch models, with a generative API for transformers | Current, pip package litert-torch |
| LiteRT-LM | LLM runtime: tokenizer, sampling, KV cache management, conversation API; reads .litertlm bundles | Current, recommended path |
| MediaPipe LLM Inference API | Earlier high-level LLM API for Android, iOS and web; reads .task bundles | Maintenance only; Google recommends migrating to LiteRT-LM |
| NNAPI delegate | Android's old hardware abstraction for accelerators | NNAPI deprecated in Android 15 |
The layering matters for debugging. When generation is slow, the cause sits in one of three places: the converted graph, the delegate that runs it, or the loop around it. Each has its own tools.
How the runtime executes a model
A .tflite file is a FlatBuffer containing the operator graph, tensor metadata and the constant weights. Because FlatBuffers can be read in place, the runtime memory-maps the file: weights are paged in from storage as ops touch them rather than copied into heap, which keeps load time and resident memory down for multi-gigabyte files.
The interpreter executes the graph op by op on the CPU by default, using the XNNPACK library for optimized float and quantized kernels. A delegate is a plug-in that claims the subgraphs it can run on other hardware. The GPU delegate compiles supported ops into GPU programs, using OpenCL or OpenGL ES compute shaders on Android and Metal on Apple devices, and typically computes in 16-bit floating point. Vendor NPU delegates and accelerator plug-ins do the same for neural accelerators.
Partitioning is where performance is won or lost. If one op in the middle of a layer is unsupported by the delegate, the graph splits: the GPU runs part of the layer, copies tensors back to CPU memory, the CPU runs the odd op, and the result is copied back. In a transformer that repeats the pattern in every layer, those copies can cost more than the math. A converted LLM is fast on GPU only when the whole decoder block maps to supported ops, which is why the generative converter re-authors models from a set of known-good building blocks instead of tracing arbitrary PyTorch.
Prefill, decode and the KV cache as tensors
The interpreter wants fixed shapes, so an LLM is exported as several entry points over one shared set of weights, called signatures. The generative API exports at least two. A prefill signature takes a batch of prompt tokens with their positions and fills the KV cache for all of them in one pass. A decode signature takes one token and its position, attends over the cache, appends that token's keys and values, and returns logits. Prefill signatures can be exported for several lengths, named prefill_{SEQ-LENS}, and the runtime chooses the one closest to the input; a 37-token prompt runs through a 64-token prefill with padding.
The KV cache is not hidden state inside the runtime. It is a set of tensors, sized for a maximum context length chosen at conversion, passed into each call and updated by it. That has consequences: memory for the full context is allocated whether you use it or not, the maximum context is fixed in the file, and the loop around the graph owns the cache's lifetime. The loop looks like this:
# Pseudocode for what an LLM runtime does around the converted graph
cache = allocate_kv(layers, kv_heads, max_len, head_dim) # fixed at conversion
sig = pick_prefill(len(prompt_ids)) # e.g. prefill_64
logits, cache = sig(tokens=pad(prompt_ids, sig.len), positions=range(sig.len), kv=cache)
pos = len(prompt_ids)
tok = sample(logits[pos - 1], temperature, top_k)
while tok != EOS and pos < max_len:
emit(tok)
logits, cache = decode(tokens=[tok], positions=[pos], kv=cache)
pos += 1
tok = sample(logits, temperature, top_k)Prefill and decode stress the hardware differently. Prefill multiplies a whole block of tokens against every weight matrix, so it is compute-bound and benefits most from the GPU or NPU. Decode multiplies one token against every weight, so each step must stream all weights from memory, and its speed is bound by memory bandwidth. That is why weight quantization speeds up decode almost in proportion to the bytes saved, and why time to first token and tokens per second should be measured separately.
Worked example: a memory budget
Take a hypothetical 2-billion-parameter decoder with 26 layers, 4 key-value heads and a head dimension of 256, similar in shape to small open models, and a 4,096-token maximum context.
| Item | Calculation | Size |
|---|---|---|
| Weights, 16-bit | 2.0e9 x 2 bytes | 4.0 GB |
| Weights, int8 | 2.0e9 x 1 byte | 2.0 GB |
| Weights, int4 | 2.0e9 x 0.5 bytes | 1.0 GB plus scales |
| KV per token, 16-bit | 2 x 26 x 4 x 256 x 2 bytes | 104 KiB |
| KV cache at 4,096 tokens | 104 KiB x 4,096 | 416 MiB |
On a phone with 8 GB of RAM, of which the operating system and other apps hold much, the 16-bit model is impractical alongside an app, int8 is tight, and int4 weights plus the cache use roughly 1.5 GB, which is workable. Halving the context to 2,048 saves 208 MiB, which can be the difference between running and being killed by the low-memory killer when the app goes to the background. Decode speed follows weight bytes: if the device sustains, say, 30 GB/s of effective bandwidth, the ceiling for int4 is around 30 tokens per second and for int8 around 15, before compute and overheads. Measure on real devices; these are bounds, not predictions. Quantization trade-offs are covered in SLM edge quantization.
Converting a PyTorch model
Conversion with litert-torch follows a fixed workflow: re-author the model with the generative API's transformer building blocks (or start from one of its examples, which include Gemma, TinyLlama and others), load the trained weights, check that the re-authored model matches the original's outputs, then convert with a quantization recipe and export prefill and decode signatures. The sketch below shows the shape; exact module paths and arguments live in the repository's examples, which change between releases, so copy from there rather than from here.
# Shape of a litert-torch generative conversion (pip install litert-torch).
# Sketch only: take exact imports and arguments from the repo's example for your model.
import litert_torch
model = build_reauthored_model(checkpoint_dir) # from the generative API examples
assert_close(model, reference_hf_model, sample_prompts) # verify before converting
quant_config = quant_recipes.full_int8_dynamic_recipe() # one of the provided recipes
edge_model = (
litert_torch.signature("prefill_64", model, prefill_inputs(64))
.signature("prefill_256", model, prefill_inputs(256))
.signature("decode", model, decode_inputs())
.convert(quant_config=quant_config)
)
edge_model.export("model_q8.tflite")The verification step is not optional. Re-authoring can introduce subtle mismatches in rotary position handling, normalization epsilon or attention masking that still produce fluent text but different answers. Compare logits on a fixed prompt set before and after re-authoring, and again after quantization, with a task-level eval on top.
The .tflite file alone cannot generate text: it lacks the tokenizer, special tokens, chat template and sampling defaults. Those are packed with it into a bundle, .litertlm for LiteRT-LM or .task for the older MediaPipe API. Ship and version the bundle as one unit so a tokenizer can never be paired with the wrong weights.
Running it with LiteRT-LM
LiteRT-LM exposes an engine and conversations. On Android and the JVM the Kotlin API looks like this, per Google's getting-started guide:
// build.gradle.kts: pin an exact version rather than latest.release
// implementation("com.google.ai.edge.litertlm:litertlm-android:<version>")
val engineConfig = EngineConfig(
modelPath = "/data/local/tmp/model.litertlm",
backend = Backend.GPU(), // or Backend.CPU(), Backend.NPU()
cacheDir = context.cacheDir.path // optional: speeds up the second load
)
val engine = Engine(engineConfig)
engine.initialize() // slow: run off the main thread
engine.createConversation().use { conversation ->
conversation.sendMessageAsync("Summarise this note: ...")
.collect { chunk -> appendToUi(chunk.toString()) }
}Google's guide notes that initialize() can take up to around 10 seconds, so it belongs on a background thread behind a loading state, ideally started before the user asks a question. The optional cache directory lets the runtime cache data between loads, which the guide says speeds up the second load. Keep one engine per process and close it when the feature is no longer in use; two engines double the memory.
Choosing a backend
| Backend | Good at | Watch for |
|---|---|---|
| CPU (XNNPACK) | Works everywhere; strong int8 and int4 kernels; predictable | Slower prefill; competes with the UI thread for cores |
| GPU | Fast prefill; frees CPU | First-load compilation time; fp16 precision; driver differences between devices |
| NPU | Best performance per watt where supported | Device-specific support and model preparation |
Choose per device class, not once. Keep a CPU fallback for every model, test on the oldest device you support, and record which backend actually ran in your telemetry, because a silent fallback from GPU to CPU looks like a performance regression with no code change.
Failure modes
- Partial delegation. One unsupported op splits every layer between GPU and CPU, and decode gets slower than CPU alone. Check the delegate's log of claimed ops.
- Killed in the background. Weights plus a full-length cache exceed what the OS lets a background app keep. Shrink context, use int4 weights, release the engine on background.
- Slow first answer. Initialization and GPU program compilation happen when the user taps send. Warm up early and set a cache directory.
- Fluent but wrong after conversion. A re-authoring mismatch in positions or masking. Compare logits against the reference before shipping.
- Context overflow. The conversation grows past the converted maximum length. Truncate or summarise history before you reach it.
- Thermal throttling. Sustained generation heats the device and clocks drop. Benchmark over minutes, not single prompts.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| int4 weights | Half the memory and roughly double decode speed of int8 | Some quality loss; needs evaluation |
| Longer max context | Longer conversations | Cache memory allocated up front |
| Several prefill lengths | Less padding waste | Larger file, more signatures to test |
| Bundled model vs OS-provided model | Your choice of model and version | Download size and update burden |
What to do next
- Pick a model from the litert-torch examples that fits your memory budget using the calculation above.
- Convert with an int8 recipe first, verify logits against the reference, then try int4 and compare task quality.
- Package the model, tokenizer and template as one versioned .litertlm bundle.
- Integrate LiteRT-LM with initialization off the main thread and a cache directory.
- Measure time to first token and decode tokens per second separately on CPU and GPU, on your oldest supported device.
- Log the backend that actually ran, and keep a CPU fallback.
- Set the maximum context to what your feature needs, not the largest the model allows.
- If you still use the MediaPipe LLM Inference API, plan the move to LiteRT-LM.