TensorFlow Lite was built for small vision and audio models with fixed input shapes: load a file, feed a tensor, read a tensor. A language model breaks every one of those assumptions. Its input length varies from one token to thousands, it runs in a loop that feeds its own output back in, and it carries a growing cache of attention keys and values between steps. Yet the same runtime, now called LiteRT, is one of the main ways to run LLMs on phones and laptops, using the CPU, the mobile GPU and increasingly the NPU.

This article explains how that works from the runtime up: what is in a model file, how delegates hand parts of the graph to the GPU, how an LLM is cut into prefill and decode signatures with the KV cache as explicit tensors, how a PyTorch model is converted and quantized, and how the LiteRT-LM runtime drives the loop in an app. A worked memory budget shows what fits on a device. For the NPU hardware itself see edge NPUs, in depth; for app lifecycle and product choices see on-device LLM inference and SLMs on mobile in 2026.

TFLite, LiteRT, LiteRT-LM: what is what

The names changed recently, and older tutorials use the old ones. Google renamed TensorFlow Lite to LiteRT in 2024; the file format and much of the API carried over, and the TensorFlow Lite name still appears in packages and documentation. For LLMs specifically, the pieces are:

PieceRoleStatus
LiteRT (formerly TFLite)Runtime: loads .tflite flatbuffers, runs ops on CPU, GPU or NPUCurrent
litert-torch (formerly AI Edge Torch)Python library that converts PyTorch models, with a generative API for transformersCurrent, pip package litert-torch
LiteRT-LMLLM runtime: tokenizer, sampling, KV cache management, conversation API; reads .litertlm bundlesCurrent, recommended path
MediaPipe LLM Inference APIEarlier high-level LLM API for Android, iOS and web; reads .task bundlesMaintenance only; Google recommends migrating to LiteRT-LM
NNAPI delegateAndroid's old hardware abstraction for acceleratorsNNAPI deprecated in Android 15

The layering matters for debugging. When generation is slow, the cause sits in one of three places: the converted graph, the delegate that runs it, or the loop around it. Each has its own tools.

How the runtime executes a model

A .tflite file is a FlatBuffer containing the operator graph, tensor metadata and the constant weights. Because FlatBuffers can be read in place, the runtime memory-maps the file: weights are paged in from storage as ops touch them rather than copied into heap, which keeps load time and resident memory down for multi-gigabyte files.

The interpreter executes the graph op by op on the CPU by default, using the XNNPACK library for optimized float and quantized kernels. A delegate is a plug-in that claims the subgraphs it can run on other hardware. The GPU delegate compiles supported ops into GPU programs, using OpenCL or OpenGL ES compute shaders on Android and Metal on Apple devices, and typically computes in 16-bit floating point. Vendor NPU delegates and accelerator plug-ins do the same for neural accelerators.

Partitioning is where performance is won or lost. If one op in the middle of a layer is unsupported by the delegate, the graph splits: the GPU runs part of the layer, copies tensors back to CPU memory, the CPU runs the odd op, and the result is copied back. In a transformer that repeats the pattern in every layer, those copies can cost more than the math. A converted LLM is fast on GPU only when the whole decoder block maps to supported ops, which is why the generative converter re-authors models from a set of known-good building blocks instead of tracing arbitrary PyTorch.

Prefill, decode and the KV cache as tensors

Prompt tokenslength 37prefill_64 signaturepad 37 to 64, one passlogitsSamplertoken 1decode signature1 token per callnext tokenlogitsSamplertoken 2, 3, ...KV cache tensorslayers x 2 x kv_heads x max_len x head_dimread/writeread/writeBoth signatures share one weight buffer;the cache is an input and an output of each call
An LLM in LiteRT: a prefill signature processes the prompt in one pass, then the decode signature runs once per generated token, both reading and updating the KV cache tensors.

The interpreter wants fixed shapes, so an LLM is exported as several entry points over one shared set of weights, called signatures. The generative API exports at least two. A prefill signature takes a batch of prompt tokens with their positions and fills the KV cache for all of them in one pass. A decode signature takes one token and its position, attends over the cache, appends that token's keys and values, and returns logits. Prefill signatures can be exported for several lengths, named prefill_{SEQ-LENS}, and the runtime chooses the one closest to the input; a 37-token prompt runs through a 64-token prefill with padding.

The KV cache is not hidden state inside the runtime. It is a set of tensors, sized for a maximum context length chosen at conversion, passed into each call and updated by it. That has consequences: memory for the full context is allocated whether you use it or not, the maximum context is fixed in the file, and the loop around the graph owns the cache's lifetime. The loop looks like this:

# Pseudocode for what an LLM runtime does around the converted graph
cache = allocate_kv(layers, kv_heads, max_len, head_dim)       # fixed at conversion
sig = pick_prefill(len(prompt_ids))                            # e.g. prefill_64
logits, cache = sig(tokens=pad(prompt_ids, sig.len), positions=range(sig.len), kv=cache)
pos = len(prompt_ids)
tok = sample(logits[pos - 1], temperature, top_k)
while tok != EOS and pos < max_len:
    emit(tok)
    logits, cache = decode(tokens=[tok], positions=[pos], kv=cache)
    pos += 1
    tok = sample(logits, temperature, top_k)

Prefill and decode stress the hardware differently. Prefill multiplies a whole block of tokens against every weight matrix, so it is compute-bound and benefits most from the GPU or NPU. Decode multiplies one token against every weight, so each step must stream all weights from memory, and its speed is bound by memory bandwidth. That is why weight quantization speeds up decode almost in proportion to the bytes saved, and why time to first token and tokens per second should be measured separately.

Worked example: a memory budget

Take a hypothetical 2-billion-parameter decoder with 26 layers, 4 key-value heads and a head dimension of 256, similar in shape to small open models, and a 4,096-token maximum context.

ItemCalculationSize
Weights, 16-bit2.0e9 x 2 bytes4.0 GB
Weights, int82.0e9 x 1 byte2.0 GB
Weights, int42.0e9 x 0.5 bytes1.0 GB plus scales
KV per token, 16-bit2 x 26 x 4 x 256 x 2 bytes104 KiB
KV cache at 4,096 tokens104 KiB x 4,096416 MiB

On a phone with 8 GB of RAM, of which the operating system and other apps hold much, the 16-bit model is impractical alongside an app, int8 is tight, and int4 weights plus the cache use roughly 1.5 GB, which is workable. Halving the context to 2,048 saves 208 MiB, which can be the difference between running and being killed by the low-memory killer when the app goes to the background. Decode speed follows weight bytes: if the device sustains, say, 30 GB/s of effective bandwidth, the ceiling for int4 is around 30 tokens per second and for int8 around 15, before compute and overheads. Measure on real devices; these are bounds, not predictions. Quantization trade-offs are covered in SLM edge quantization.

Converting a PyTorch model

Conversion with litert-torch follows a fixed workflow: re-author the model with the generative API's transformer building blocks (or start from one of its examples, which include Gemma, TinyLlama and others), load the trained weights, check that the re-authored model matches the original's outputs, then convert with a quantization recipe and export prefill and decode signatures. The sketch below shows the shape; exact module paths and arguments live in the repository's examples, which change between releases, so copy from there rather than from here.

# Shape of a litert-torch generative conversion (pip install litert-torch).
# Sketch only: take exact imports and arguments from the repo's example for your model.
import litert_torch
model = build_reauthored_model(checkpoint_dir)          # from the generative API examples
assert_close(model, reference_hf_model, sample_prompts)  # verify before converting
quant_config = quant_recipes.full_int8_dynamic_recipe()  # one of the provided recipes
edge_model = (
    litert_torch.signature("prefill_64", model, prefill_inputs(64))
    .signature("prefill_256", model, prefill_inputs(256))
    .signature("decode", model, decode_inputs())
    .convert(quant_config=quant_config)
)
edge_model.export("model_q8.tflite")

The verification step is not optional. Re-authoring can introduce subtle mismatches in rotary position handling, normalization epsilon or attention masking that still produce fluent text but different answers. Compare logits on a fixed prompt set before and after re-authoring, and again after quantization, with a task-level eval on top.

The .tflite file alone cannot generate text: it lacks the tokenizer, special tokens, chat template and sampling defaults. Those are packed with it into a bundle, .litertlm for LiteRT-LM or .task for the older MediaPipe API. Ship and version the bundle as one unit so a tokenizer can never be paired with the wrong weights.

Running it with LiteRT-LM

LiteRT-LM exposes an engine and conversations. On Android and the JVM the Kotlin API looks like this, per Google's getting-started guide:

// build.gradle.kts: pin an exact version rather than latest.release
// implementation("com.google.ai.edge.litertlm:litertlm-android:<version>")

val engineConfig = EngineConfig(
    modelPath = "/data/local/tmp/model.litertlm",
    backend = Backend.GPU(),               // or Backend.CPU(), Backend.NPU()
    cacheDir = context.cacheDir.path       // optional: speeds up the second load
)
val engine = Engine(engineConfig)
engine.initialize()                        // slow: run off the main thread

engine.createConversation().use { conversation ->
    conversation.sendMessageAsync("Summarise this note: ...")
        .collect { chunk -> appendToUi(chunk.toString()) }
}

Google's guide notes that initialize() can take up to around 10 seconds, so it belongs on a background thread behind a loading state, ideally started before the user asks a question. The optional cache directory lets the runtime cache data between loads, which the guide says speeds up the second load. Keep one engine per process and close it when the feature is no longer in use; two engines double the memory.

Choosing a backend

BackendGood atWatch for
CPU (XNNPACK)Works everywhere; strong int8 and int4 kernels; predictableSlower prefill; competes with the UI thread for cores
GPUFast prefill; frees CPUFirst-load compilation time; fp16 precision; driver differences between devices
NPUBest performance per watt where supportedDevice-specific support and model preparation

Choose per device class, not once. Keep a CPU fallback for every model, test on the oldest device you support, and record which backend actually ran in your telemetry, because a silent fallback from GPU to CPU looks like a performance regression with no code change.

Failure modes

  • Partial delegation. One unsupported op splits every layer between GPU and CPU, and decode gets slower than CPU alone. Check the delegate's log of claimed ops.
  • Killed in the background. Weights plus a full-length cache exceed what the OS lets a background app keep. Shrink context, use int4 weights, release the engine on background.
  • Slow first answer. Initialization and GPU program compilation happen when the user taps send. Warm up early and set a cache directory.
  • Fluent but wrong after conversion. A re-authoring mismatch in positions or masking. Compare logits against the reference before shipping.
  • Context overflow. The conversation grows past the converted maximum length. Truncate or summarise history before you reach it.
  • Thermal throttling. Sustained generation heats the device and clocks drop. Benchmark over minutes, not single prompts.

Trade-offs

ChoiceGainCost
int4 weightsHalf the memory and roughly double decode speed of int8Some quality loss; needs evaluation
Longer max contextLonger conversationsCache memory allocated up front
Several prefill lengthsLess padding wasteLarger file, more signatures to test
Bundled model vs OS-provided modelYour choice of model and versionDownload size and update burden

What to do next

  1. Pick a model from the litert-torch examples that fits your memory budget using the calculation above.
  2. Convert with an int8 recipe first, verify logits against the reference, then try int4 and compare task quality.
  3. Package the model, tokenizer and template as one versioned .litertlm bundle.
  4. Integrate LiteRT-LM with initialization off the main thread and a cache directory.
  5. Measure time to first token and decode tokens per second separately on CPU and GPU, on your oldest supported device.
  6. Log the backend that actually ran, and keep a CPU fallback.
  7. Set the maximum context to what your feature needs, not the largest the model allows.
  8. If you still use the MediaPipe LLM Inference API, plan the move to LiteRT-LM.
Key takeaway: An LLM on LiteRT is a memory-mapped graph cut into prefill and decode signatures with an explicit, fixed-size KV cache; convert it with litert-torch and verify it, budget memory before choosing a model, run it through LiteRT-LM off the main thread, and measure prefill and decode separately on every backend you ship.