Internationalising an agent is different from internationalising a web form. A form has a fixed set of strings that translators can work through. An agent produces most of its words at run time from a language model, calls tools that return dates, prices and quantities, and still shows a handful of fixed messages such as errors, refusals and consent notices that must be exact. Each of those three kinds of text needs a different mechanism, and treating them the same way is the root of most i18n bugs in agent systems.

This article shows how to do it with the Agent Development Kit for Java: resolving a locale once and keeping it in session state, steering the model's reply language through the instruction, serving fixed strings from resource bundles with correct plural rules, keeping tool contracts locale-neutral, and measuring language drift. ADK Java does not ship an i18n module of its own, so everything here is built from ADK's state, instruction templating and callbacks plus the JDK and ICU. Code targets the builder-style API described in the ADK documentation; check method names against the version you use. Related reading: prompts in ADK Java, callbacks and accessibility.

Three kinds of text

Three kinds of text in an ADK Java agent, and where each is localisedHTTP requestAccept-Language, profileLocale resolverBCP 47 tag, fallbackSession stateuser_locale, reply_languageRunnerrunAsync per turn1. Model textinstruction placeholder2. Fixed app stringsResourceBundle, ICU messages3. Tool dataISO dates, codes, numbersafterModel checklanguage drift metricFormatter layerjava.time, NumberFormatPer-locale evalssame cases, every localeModel text is steered, fixed strings are looked up, and tool data stays locale-neutral until display.
Locale is resolved once per session and read by three separate mechanisms. Only the first involves the model; the other two are deterministic code you can unit test.
Kind of textExampleMechanismTestable how
Model-generatedExplanations, summaries, follow-up questionsReply language in the instruction from session statePer-locale evaluation sets and a language-ID metric
Fixed application stringsRate-limit message, refusal text, consent notice, button labelsResourceBundle or ICU message files keyed by IDOrdinary unit tests and translation review
Tool dataOrder dates, prices, stock countsLocale-neutral values in tool contracts, formatted at the edgeContract tests on JSON, formatter tests per locale

The rule that falls out of the table: never ask the model to produce text that has to be exact. Legal notices, prices and error codes come from code, and the model is told to repeat or reference them, not to paraphrase them in another language.

Resolving the locale once

Decide the locale at the edge of your service, once, and store it as a BCP 47 language tag such as de-CH or pt-BR. Sources, in order of trust: an explicit user setting in their profile, the Accept-Language header, then a deployment default. Match against the locales you actually support, not every locale the JDK knows, so you never promise a language you have not evaluated.

import java.util.List;
import java.util.Locale;

public final class LocaleResolver {
  private static final List<Locale> SUPPORTED = List.of(
      Locale.forLanguageTag("en"), Locale.forLanguageTag("de"),
      Locale.forLanguageTag("fr"), Locale.forLanguageTag("ja"),
      Locale.forLanguageTag("pt-BR"));
  private static final Locale DEFAULT = Locale.forLanguageTag("en");

  public static Locale resolve(String profileTag, String acceptLanguage) {
    if (profileTag != null && !profileTag.isBlank()) {
      Locale p = Locale.forLanguageTag(profileTag);
      Locale hit = Locale.lookup(List.of(new Locale.LanguageRange(p.toLanguageTag())), SUPPORTED);
      if (hit != null) return hit;
    }
    if (acceptLanguage != null && !acceptLanguage.isBlank()) {
      try {
        Locale hit = Locale.lookup(Locale.LanguageRange.parse(acceptLanguage), SUPPORTED);
        if (hit != null) return hit;
      } catch (IllegalArgumentException malformed) {
        // fall through to default; log the header for diagnosis
      }
    }
    return DEFAULT;
  }
}

Locale.lookup applies the RFC 4647 lookup algorithm, so a request for pt-PT falls back to pt and then fails if only pt-BR is supported. That is usually what you want: Brazilian and European Portuguese differ enough that a silent substitution should be a product decision, made by adding an explicit mapping, not an accident of matching.

Locale in session state and the instruction

Put the resolved locale into session state when the session is created, together with a human-readable language name for the prompt. ADK replaces {key} placeholders in an agent's instruction with values from session state before each model call, so the instruction can stay a single English template while the reply language varies per user.

import com.google.adk.agents.LlmAgent;
import com.google.adk.runner.InMemoryRunner;
import com.google.adk.sessions.Session;
import java.util.Locale;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentMap;

LlmAgent supportAgent = LlmAgent.builder()
    .name("support_agent")
    .model("gemini-2.5-flash")
    .instruction("""
        You are the support assistant for Acme.
        Reply in {reply_language}. Keep product names, order IDs and codes exactly as given.
        When a tool returns a field named display_text, quote it verbatim; do not translate it.
        Dates and prices in tool results are already formatted for the user; do not reformat them.
        If the user writes in a different language, reply in that language and call
        set_reply_language with its BCP 47 tag.
        """)
    .tools(orderTools, languageTools)
    .build();

InMemoryRunner runner = new InMemoryRunner(supportAgent, "acme-support");

Locale locale = LocaleResolver.resolve(profile.languageTag(), request.header("Accept-Language"));
ConcurrentMap<String, Object> state = new ConcurrentHashMap<>();
state.put("user_locale", locale.toLanguageTag());                    // e.g. "de"
state.put("reply_language", locale.getDisplayLanguage(Locale.ENGLISH)); // e.g. "German"

Session session = runner.sessionService()
    .createSession("acme-support", userId, state, null)
    .blockingGet();

Two details matter. First, the language name is rendered in English with getDisplayLanguage(Locale.ENGLISH) because the instruction is English; mixing scripts inside the system prompt is a common cause of the model switching language halfway through. Second, every key an instruction references must exist in state before the first turn; a missing key is an error, not a blank. Keep the instruction in one language and evaluate it in all locales, rather than maintaining translated copies of the prompt that drift apart and multiply your evaluation work.

Fixed strings and plural rules

Everything the user must see exactly, and everything your support team needs to search logs for, should come from message files keyed by a stable ID. Since Java 9, .properties resource bundles are read as UTF-8, so non-Latin text no longer needs escaping. Simple messages are fine with java.text.MessageFormat, but plurals are not: English has two forms, Polish and Russian have several, Arabic has six categories, and Japanese has one. The JDK's ChoiceFormat cannot express those rules, so use ICU4J's MessageFormat, which implements CLDR plural rules.

# messages_en.properties
quota.exceeded=You have used your {count, plural, one {# request} other {# requests}} for today. Try again after {reset}.
# messages_pl.properties
quota.exceeded=Wykorzystano {count, plural, one {# zapytanie} few {# zapytania} many {# zapytań} other {# zapytania}} na dziś. Spróbuj ponownie po {reset}.
import com.ibm.icu.text.MessageFormat;
import java.util.Locale;
import java.util.Map;
import java.util.ResourceBundle;

public final class Messages {
  public static String get(Locale locale, String key, Map<String, Object> args) {
    ResourceBundle bundle = ResourceBundle.getBundle("messages", locale,
        ResourceBundle.Control.getNoFallbackControl(ResourceBundle.Control.FORMAT_PROPERTIES));
    MessageFormat mf = new MessageFormat(bundle.getString(key), locale);
    return mf.format(args);
  }
}

The no-fallback control matters on servers: by default a missing bundle falls back to the JVM's default locale before the base file, so a host configured for German would answer French users in German.

In an agent, fixed strings appear in two places. Code paths that never reach the model, such as a rate-limit rejection before the runner is called, use Messages.get directly. Short-circuits inside the agent, such as a beforeModel callback that blocks a request, should build their response from the same bundle, using the locale in callbackContext.state(), so that a policy refusal reads identically every time rather than being regenerated by the model in slightly different words.

Locale-neutral tools

Tool contracts are an API between the model and your code, and APIs should not be localised. Accept and return ISO 8601 dates, ISO 4217 currency codes, decimal amounts as numbers or strings without grouping separators, and enum values in English. The model is good at turning "nächsten Freitag" into 2026-10-09 if the tool schema says the parameter is an ISO date and the instruction gives today's date and the user's time zone; it is bad at guessing whether 1.234 means one thousand or one point two.

Where the tool result will be shown to the user, add a pre-formatted display_text field produced by deterministic code, and tell the model to quote it. That keeps prices and dates correct and consistent with the rest of your product.

import java.math.BigDecimal;
import java.text.NumberFormat;
import java.time.ZonedDateTime;
import java.time.format.DateTimeFormatter;
import java.time.format.FormatStyle;
import java.util.Currency;
import java.util.Locale;
import java.util.Map;

static Map<String, Object> orderResult(Order o, Locale locale) {
  NumberFormat money = NumberFormat.getCurrencyInstance(locale);
  money.setCurrency(Currency.getInstance(o.currencyCode()));   // e.g. "EUR"
  DateTimeFormatter date = DateTimeFormatter.ofLocalizedDate(FormatStyle.MEDIUM).withLocale(locale);
  ZonedDateTime eta = o.eta().atZone(o.customerZone());
  return Map.of(
      "order_id", o.id(),
      "total", o.total().toPlainString(),          // locale-neutral, for reasoning
      "currency", o.currencyCode(),
      "eta", eta.toLocalDate().toString(),          // ISO 8601
      "display_text", money.format(o.total()) + " · " + date.format(eta));
}

The currency comes from the order, not from the locale. A German-speaking customer in Switzerland paying in francs must see francs; deriving currency from language is a classic and expensive bug. The tool reads the locale from ToolContext state, so the model never has to pass it.

Measuring language drift

Models usually follow a reply-language instruction, but not always: long tool outputs in English, a pasted English error message, or a technical topic can pull the reply back to English. You cannot prevent this completely, so measure it. An afterModel callback sees each response before it is used; run a language-identification library over the text and record mismatches in state and metrics. Do not rewrite or translate the response in the callback, which would add latency and a second model's errors to every turn; use the metric to fix prompts and tools.

import com.google.adk.agents.Callbacks;
import java.util.Optional;

Callbacks.AfterModelCallbackSync languageDrift = (ctx, response) -> {
  String expected = (String) ctx.state().get("user_locale");
  String text = ResponseText.of(response);               // your helper: concatenated text parts
  if (text.length() > 40) {                                // too short to classify reliably
    String detected = LanguageId.detect(text);             // your chosen language-ID library
    if (!Locale.forLanguageTag(expected).getLanguage().equals(detected)) {
      metrics.counter("agent_language_mismatch", "expected", expected).increment();
    }
  }
  return Optional.empty();                                 // never alter the response here
};

Language switching by the user is a product decision. The instruction above follows the user's latest language and asks the model to call a small set_reply_language tool, which validates the tag against the supported list and updates user_locale and reply_language in state, so fixed strings and formatters switch too. Without that tool the model's prose switches but your error messages stay in the old language.

Failure modes

FailureCauseFix
Turkish users cannot match commands or IDstoLowerCase() without a locale maps I to dotless i in TurkishUse toLowerCase(Locale.ROOT) for identifiers; locale-aware case only for display
Tests break after a JDK upgradeCLDR data changed; for example, time formats in some locales now use a narrow no-break spaceAssert on parsed values or normalise whitespace; pin expectations per JDK
Same text fails equality checksComposed and decomposed Unicode forms of the same accented characterNormalise to NFC with java.text.Normalizer before comparing or hashing
Wrong currency shownCurrency derived from locale instead of the transactionCarry ISO 4217 code on the data
Replies drift to EnglishEnglish tool output and long contexts dominatedisplay_text fields, drift metric, per-locale evals
Token budget exceeded in some languagesNon-Latin scripts often use more tokens per wordMeasure tokens per locale and size context limits per locale
Right-to-left text renders garbledMixed Arabic or Hebrew with Latin IDs and no bidi isolationWrap embedded IDs in Unicode isolates or handle direction in the UI

Run your existing evaluation cases in every supported locale, with the same expected tool calls, and track pass rate per locale. A drop in one language is often a tool-description or date-parsing problem rather than a model limitation. The ADK Java evaluation overview covers building those suites.

Trade-offs

One English instruction with a reply-language placeholder is cheap to maintain and keeps evaluation comparable across languages; the cost is occasional drift and a slightly less natural tone in some languages. Fully translated instructions can read better and drift less for a single high-value market, but every prompt change then becomes a translation job and an evaluation run per language. Start with one template and add a localised variant only where measurements justify it.

Letting the model format dates and money removes code but makes correctness depend on sampling. Deterministic display_text adds a field to every tool and a formatter layer, and in exchange guarantees exact output. For anything with legal or financial weight, take the deterministic path.

What to do next

  1. Write down your supported locales and add a resolver that maps requests onto exactly that list.
  2. Store user_locale and reply_language in session state at creation; reference them in the instruction.
  3. Move every fixed user-facing string into message bundles with ICU plural syntax.
  4. Audit tool schemas: ISO dates, ISO 4217 codes, plain decimals, and a display_text field where users see results.
  5. Search the codebase for toLowerCase() and toUpperCase() without a locale argument.
  6. Add the language-drift metric and a per-locale run of your evaluation suite to CI.
  7. Add a set_reply_language tool and test a conversation that switches language midway.
Key takeaway: An agent speaks three kinds of text. Steer model text with a reply language held in session state, serve exact strings from ICU message bundles, and keep tool data locale-neutral until a deterministic formatter renders it. Resolve the locale once, let users switch it through a tool, measure drift, and run every evaluation in every supported locale.