There are two ways to put a small language model on a user's device. You can ship one yourself, with its weights, runtime and update path, as covered in on-device LLM inference. Or you can use a model the platform already provides and maintains, and call it through an API. Google's Gemini Nano is the most widely deployed example of the second approach: it runs inside Android's AICore system service on supported phones, and inside Chrome on capable desktops.

Building on a system model changes the engineering problem. You no longer manage weights, but you no longer control them either. You must handle devices where the model does not exist yet, or never will; limits the platform imposes on your app; a model that can change underneath you; and a context window and capability level well below a cloud model. This article explains the architecture, then works through each of those problems with code, ending with a worked example and a checklist.

Advertisement

Why a platform owns the model

A small model quantised for a phone still occupies gigabytes. If every app bundled its own, a phone with five AI features would store five copies, load them into memory separately and update them through five release trains. A system model is stored once, kept warm by the platform, scheduled against the device's NPU or GPU by the operating system (see NPU acceleration for why that matters), and updated without app releases.

For developers the trade is simple to state. You gain zero download size, no runtime to maintain and hardware acceleration you did not have to build. You lose control of the model version, broad device coverage and the ability to fine-tune the base model yourself. Whether that trade suits a feature depends on whether the feature can degrade gracefully on devices without the model, which is the theme of the rest of this article. SLMs on mobile in 2026 compares this path with bundling in more detail.

The architecture on Android and in Chrome

On Android, Gemini Nano runs in AICore, a system service. Google's documentation says AICore manages the distribution of Gemini Nano and its updates, so apps do not download models or carry their memory cost, and that it isolates each request and does not keep a record of inputs or outputs after processing. Downloads go through Private Compute Services, a companion component that handles the network requests. Apps reach the model through the ML Kit GenAI APIs: task APIs for summarisation, proofreading, rewriting, image description and speech recognition, plus a general Prompt API.

In Chrome, the browser downloads Gemini Nano as a component the first time a site or extension needs it, then shares it across sites. Pages reach it through the built-in AI APIs: task APIs such as Summarizer, plus the general Prompt API. These run on desktop only, on Windows 10 or 11, macOS 13 or later, Linux and supported ChromeOS devices, and require at least 22 GB of free space on the volume holding the Chrome profile plus either a GPU with more than 4 GB of video memory or 16 GB of RAM and 4 CPU cores. Chrome on Android and iOS does not offer them.

AndroidChrome on desktopYour appML Kit GenAI APIsAICore system serviceisolates requests, enforces quotaGemini Nanoshared weightsTask APIssummarise, rewrite, ...Private Compute Servicesmodel download and updatesYour page or extensionPrompt, Summarizer, ...Built-in AI in the browseravailability, sessions, policyGemini Nano componentdownloaded once, shared by sitesGPU or CPUdevice hardware checksCloud modelfallback routeunavailableThe model belongs to the platform: you check availability, it handles weights, updates and scheduling
Two hosts for the same model family. In both, your code asks the platform for a session; the platform owns weights, updates, scheduling and policy, and your code owns the fallback.
Advertisement

Availability is a state machine, not a boolean

The first fact about a system model is that it may not be there. In Chrome, availability() returns one of four states. "unavailable" means the device or the requested options are not supported. "downloadable" means a download is needed first, which may be the language model, an expert model or fine-tuning data. "downloading" means one is in progress. "available" means a session can be created now. Android's ML Kit has an equivalent feature status check with download support.

Two rules follow. Pass the same options to the availability check as to create(), because supported languages and modalities change the answer. And while the model is not yet present, Chrome requires user activation, a recent click, tap or key press, before create() may start a session. A feature that silently calls it on page load will not trigger the download. Wrap the whole decision in one function:

// One function owns the decision; the UI never calls create() directly.
async function getNanoSession(options, { onProgress } = {}) {
  if (!("LanguageModel" in self)) return { route: "cloud", why: "api-missing" };

  const state = await LanguageModel.availability(options);   // same options as create()
  switch (state) {
    case "unavailable":
      return { route: "cloud", why: "unsupported" };
    case "downloadable":
    case "downloading":
      // create() needs a user gesture until the model is present. Only offer the
      // on-device path from a click handler; otherwise serve this request from the cloud.
      if (!navigator.userActivation?.isActive) return { route: "cloud", why: "needs-gesture" };
      // fall through: create() starts or joins the download
    case "available": {
      const session = await LanguageModel.create({
        ...options,
        monitor(m) {
          m.addEventListener("downloadprogress", (e) => onProgress?.(e.loaded));
        },
      });
      return { route: "device", session };
    }
    default:
      return { route: "cloud", why: "unknown-state:" + state };
  }
}

Returning a route rather than throwing keeps call sites simple, and the reason string goes to analytics so you can see what share of users actually run on device.

The Prompt API in practice

The Prompt API is stable for Chrome extensions since Chrome 138; for ordinary web pages, check the status table in Chrome's documentation before relying on it. Sessions keep a conversation history, so prime one with a system prompt and clone it per request. responseConstraint takes a JSON Schema and constrains the output to match it, which beats asking a small model nicely.

const options = {
  initialPrompts: [{ role: "system",
    content: "You tag personal notes. Reply with JSON only. Use at most 3 tags from the list." }],
  expectedInputs: [{ type: "text", languages: ["en"] }],
  expectedOutputs: [{ type: "text", languages: ["en"] }],
};

const schema = {
  type: "object",
  properties: {
    tags: { type: "array", maxItems: 3,
            items: { enum: ["work", "family", "health", "money", "travel", "ideas"] } },
    title: { type: "string", maxLength: 60 },
  },
  required: ["tags", "title"],
};

const r = await getNanoSession(options, { onProgress: showBar });
if (r.route === "device") {
  const base = r.session;                     // keep one primed session, clone per note
  const s = await base.clone();
  try {
    const json = await s.prompt(noteText.slice(0, 6000), { responseConstraint: schema });
    render(JSON.parse(json));
    if (s.contextUsage > 0.8 * s.contextWindow) console.warn("near context limit");
  } finally {
    s.destroy();                              // free memory; sessions are not free
  }
} else {
  render(await fallbackTag(noteText));        // cloud only if the user opted in
}

Sessions also expose contextUsage and contextWindow, and fire a contextoverflow event when older history is dropped to make room. Read the window size at run time instead of hard-coding it, truncate inputs to a measured budget, and destroy sessions when finished, because each one holds memory on the user's machine.

For common jobs, prefer a task API. The Summarizer API is stable on the web since Chrome 138 and needs no prompt at all; you choose a type, format and length and pass the text. You avoid owning a prompt that may behave differently after the next model update.

// Task API, stable on the web since Chrome 138. Run from a click handler:
// create() needs user activation while the model is still downloadable.
if ((await Summarizer.availability()) !== "unavailable") {
  const summarizer = await Summarizer.create({
    type: "key-points", format: "markdown", length: "short",
    sharedContext: "Personal notes written by the user",
  });
  const stream = summarizer.summarizeStreaming(longNote);
  for await (const chunk of stream) appendToPanel(chunk);
}

Living with platform limits on Android

Android imposes rules a cloud API does not. ML Kit's documentation states that GenAI inference is permitted only when the app is the top foreground application; calls from the background, including from a foreground service, fail with BACKGROUND_USE_BLOCKED. Quotas apply per app: BUSY for short-term overuse and PER_APP_BATTERY_USE_QUOTA_EXCEEDED for long-term use. The documentation recommends exponential backoff. Supported devices vary by feature, and the Prompt API documents several Gemini Nano versions with different device lists, so capability is a property of the device, not of your app.

// Kotlin sketch. Check the current ML Kit reference for exact exception and enum
// types in your SDK version; the error codes below are the documented ones.
suspend fun tagNote(note: String): Tags {
    if (!isAppInForeground()) return enqueueForLater(note)       // background is blocked
    repeat(3) { attempt ->
        try {
            return parseTags(nano.generateContent(buildPrompt(note)))
        } catch (e: GenAiException) {
            when (e.errorCode) {
                ErrorCode.BUSY -> delay(250L shl attempt)             // short-term limit
                ErrorCode.PER_APP_BATTERY_USE_QUOTA_EXCEEDED,         // long-term limit
                ErrorCode.BACKGROUND_USE_BLOCKED -> return fallbackTag(note)
                else -> return fallbackTag(note)   // cloud only with opt-in
            }
        }
    }
    return fallbackTag(note)
}

The design consequence is that on-device generation suits interactive, user-initiated work: summarise this note now, suggest a reply to this message. Batch jobs such as tagging a whole archive overnight belong in the cloud, or should run in small foreground slices while the user is active.

Hybrid fallback without two products

Every feature built on a system model needs a second path, and the danger is that the two paths drift into two products. Hold them together with one task contract: the same input limits, the same output schema, the same validator and the same evaluation set run against both engines. Users then get the same behaviour from either route, differing only in speed, privacy and occasionally quality.

You can build the router yourself, as above, or use a library. Firebase AI Logic offers hybrid inference for web apps: you create the model with an inference mode such as InferenceMode.PREFER_ON_DEVICE, and it uses Chrome's Prompt API when Gemini Nano is available, from Chrome 139, and falls back to a cloud Gemini model otherwise. Whichever you choose, decide per feature whether cloud fallback is even allowed. A feature sold as private, such as summarising medical notes, should fail closed on unsupported devices rather than quietly send text to a server.

Prompting a small model

A model sized for a phone is not a smaller copy of a frontier model with the same behaviour. It follows simple, explicit instructions well and long, nuanced ones poorly, and it has far less world knowledge. Prompts that work on device share a style: one task per call, a short system prompt, a closed set of allowed outputs, constrained JSON, and inputs trimmed to what matters. Split compound jobs, such as summarise, then tag, then title, into separate calls with separate checks; small models handle three easy calls better than one hard one.

Do not rely on the model for facts. Supply the facts in the prompt, or keep on-device work to transforming text the user already has: summarising, rewriting, extracting and classifying. You cannot change the system model, so your levers are prompt shape, input size and routing.

Worked example: a notes app on two platforms

A notes app adds two features: a short title and up to three tags for each new note, and a key-points summary for long notes. It ships on Android and as a web app. Privacy is a selling point, so notes should stay on the device when possible, but tagging is not sensitive enough to forbid cloud fallback for users who opt in.

On a capable laptop in Chrome, the first click on Suggest tags finds the model downloadable, starts the download from the click and shows progress; that request goes to the cloud if the user has opted in, otherwise the button waits. Later notes run on device through a cloned, primed session with the schema above, in a second or two for typical notes on that hardware. Summaries use the Summarizer API with streaming output. On a supported Android phone, the same features call ML Kit while the app is in the foreground; importing two thousand old notes queues the tagging and processes a few notes each time the user opens the app, rather than hitting the battery quota in one burst. On an unsupported phone, the opted-in user gets cloud tagging and everyone else gets manual tags.

One evaluation set of three hundred consented notes runs nightly against the cloud model and, on reference devices, against Gemini Nano on each platform, so a platform model update shows up in scores before it shows up in reviews.

Failure modes

  • Assuming availability: features built for the demo device that silently do nothing on most real ones.
  • Calling create() without a user gesture while the model is still downloadable, so the download never starts.
  • Background batch jobs on Android that fail with BACKGROUND_USE_BLOCKED, or exhaust the per-app quota.
  • Silent model changes: a platform update alters output style or quality and nobody re-runs the evaluation.
  • Context overflow: long inputs or long sessions push out the system prompt, and behaviour degrades mid-conversation.
  • Leaky fallback: a feature marketed as on-device quietly routes to the cloud on unsupported devices.

Trade-offs at a glance

ChoiceGainsCosts
System model (Gemini Nano)No download, platform updates, hardware accelerationLimited devices, no control of version, platform quotas
Bundled modelAny device you target, your own fine-tune, pinned versionApp size, memory, runtime and update work
Cloud modelHighest quality, every deviceLatency, per-call cost, data leaves the device
Task API over Prompt APITuned behaviour, no prompt to maintainOnly the tasks the platform offers

What to do next

  1. Pick one interactive, text-transforming feature, such as summarise or rewrite, as the first on-device candidate.
  2. Write a task contract: input limit, output schema, validator and a labelled evaluation set.
  3. Implement the availability state machine once, including the user-gesture rule and analytics on the route taken.
  4. Prefer a task API when one matches; otherwise use the Prompt API with a primed, cloned session and a response constraint.
  5. On Android, keep calls in the foreground and handle BUSY and quota errors with backoff and fallback.
  6. Decide per feature whether cloud fallback is allowed, and fail closed where privacy is the promise.
  7. Re-run the evaluation on reference devices whenever the platform model updates.
Key takeaway: Building on Gemini Nano means treating the model as a platform service: it is absent on many devices, governed by platform rules, and updated without you. Check availability as a four-state machine, respect user-gesture, foreground and quota limits, prefer task APIs, constrain outputs, keep one task contract across device and cloud routes, and evaluate continuously so platform updates never surprise your users.