An agent deployed to more than one country receives messages in more than one language, and the language a user writes in is not always the language of their account. A Spanish speaker uses an English-locale corporate account; a German engineer pastes an English stack trace and asks about it in German; a support queue in India receives Hindi, English and Hinglish in the same hour. Handling that well is a content problem: detecting the language of each turn, deciding how to answer, finding knowledge written in other languages, and proving quality in every language you claim to support.

This article builds that layer in ADK Java. It is the companion to ADK Java internationalization, which covers the product side: resolving a locale once, ICU message bundles, locale-neutral tools, a language-drift metric and a tool for switching reply language. Those are not repeated here. This article covers per-turn detection with a callback, mixed-language input, cross-lingual retrieval, native generation versus a translation pivot, token cost, and a per-language evaluation matrix.

Locale versus turn language

Keep two values apart. The locale is a product setting: it controls dates, currency, fixed strings and the default reply language, and it changes rarely. The turn language is an observation: the language the user actually wrote in this message. Most of the time they agree. When they disagree, a good default is to answer in the turn language, because the user just demonstrated it, while keeping formatting and fixed strings in the locale until the user switches explicitly.

Two exceptions matter. A message that is mostly pasted material (logs, code, an email in another language) says nothing about the user's language; and a very short message such as "ok" or "gracias" carries too little signal to override the session. Both cases fall back to the language already recorded for the session.

Architecture

Per-turn language: detect once per invocation, steer every model call, retrieve across languagesUser messageany language, mixedStrip non-prosecode, URLs, quoted logsLingua detectorrestricted set, thresholdSession stateturn_lang, lang_invwritebeforeModelCallbackSyncappends reply-language ruleLLM call (repeats)after each tool round tripRetrieval toolmultilingual embeddingsread once per callqueryLow confidencefall back to session languageReplyin turn language, cited sourcesEval matrixlanguage x task scoresDetection is cheap but not free; cache it per invocation id, not per model call.
Data flow for per-turn language handling. Red boxes are the safety valves: confidence fallback and measurement.

The flow has one writer and many readers. The detector runs once per invocation, on the user's own text, and writes the result to session state. A before-model callback reads it on every model call and appends a short instruction to the request. Tools such as retrieval read the same state to decide how to search. Evaluation runs offline over the same code path, in every supported language.

Detecting the language of a turn

Use a language-identification library rather than asking the model, which costs a call and is easily swayed by the instruction itself. Lingua (Maven com.github.pemistahl:lingua, version 1.2.2 at the time of writing) is a JVM library that is accurate on short text. Restrict it to the languages you support: fewer candidates means fewer confusions between close pairs such as Spanish and Portuguese, and less memory, because each language model is loaded separately. A minimum relative distance makes the detector return Language.UNKNOWN instead of guessing when the top two candidates are close.

import static com.github.pemistahl.lingua.api.Language.*;

import com.github.pemistahl.lingua.api.Language;
import com.github.pemistahl.lingua.api.LanguageDetector;
import com.github.pemistahl.lingua.api.LanguageDetectorBuilder;
import java.util.Locale;
import java.util.Optional;
import java.util.regex.Pattern;

public final class TurnLanguage {
  private static final LanguageDetector DETECTOR = LanguageDetectorBuilder
      .fromLanguages(ENGLISH, SPANISH, PORTUGUESE, GERMAN, FRENCH, HINDI)
      .withMinimumRelativeDistance(0.25)
      .withPreloadedLanguageModels()        // pay the load cost at startup, not on a user turn
      .build();

  private static final Pattern FENCED = Pattern.compile("```.*?```", Pattern.DOTALL);
  private static final Pattern NOISE = Pattern.compile(
      "https?://\\S+|\\S+@\\S+|\\b[\\w.]+Exception\\b|^\\s*(at|>) .*$", Pattern.MULTILINE);

  /** ISO 639-1 code of the user's own prose, or empty when the signal is too weak. */
  public static Optional<String> detect(String message) {
    String prose = NOISE.matcher(FENCED.matcher(message).replaceAll(" ")).replaceAll(" ").strip();
    if (prose.codePointCount(0, prose.length()) < 20) {
      return Optional.empty();             // "ok", "gracias": keep the session language
    }
    Language lang = DETECTOR.detectLanguageOf(prose);
    if (lang == Language.UNKNOWN) {
      return Optional.empty();
    }
    return Optional.of(lang.getIsoCode639_1().toString());   // "es", "de", ...
  }
}

The stripping step does most of the work on technical traffic. Without it, the German engineer's message with forty lines of English stack trace is classified as English. The 20-code-point floor and the 0.25 distance are starting points; tune both against labelled samples of your own traffic, because the right values depend on how short your users' messages are. Build the detector once as a static field: it is thread-safe and expensive to create.

Detect once, steer every model call

ADK runs the before-model callback on every model call in an invocation, and one user turn with two tool round trips makes three model calls. callbackContext.userContent() returns the same message each time, so detecting inside the callback without a guard runs the detector three times and can, in principle, produce different answers if thresholds sit on a boundary. Guard on the invocation id and store the result in state.

import com.google.adk.agents.Callbacks;
import com.google.adk.agents.LlmAgent;
import com.google.genai.types.Content;
import com.google.genai.types.Part;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.stream.Collectors;

static final Map<String, String> NAMES = Map.of(
    "en", "English", "es", "Spanish", "pt", "Portuguese",
    "de", "German", "fr", "French", "hi", "Hindi");

static final Callbacks.BeforeModelCallbackSync STEER_LANGUAGE = (ctx, request) -> {
  Map<String, Object> state = ctx.state();
  if (!ctx.invocationId().equals(state.get("lang_inv"))) {      // first model call of this turn
    String text = ctx.userContent().flatMap(Content::parts).orElse(List.of()).stream()
        .map(p -> p.text().orElse(""))
        .collect(Collectors.joining(" "));
    String previous = (String) state.getOrDefault("turn_lang", state.getOrDefault("locale_lang", "en"));
    state.put("turn_lang", TurnLanguage.detect(text).orElse(previous));
    state.put("lang_inv", ctx.invocationId());
  }
  String lang = NAMES.getOrDefault((String) state.get("turn_lang"), "English");
  request.appendInstructions(List.of(
      "Write your reply in " + lang + ". Keep code, identifiers, product names and quoted "
      + "error messages unchanged. If you quote a source written in another language, "
      + "quote it in the original and give a short " + lang + " summary."));
  return Optional.empty();                                      // continue to the model
};

LlmAgent agent = LlmAgent.builder()
    .name("support_agent")
    .model("gemini-2.5-flash")
    .instruction("You are a support agent for Acme storage products. ...")
    .beforeModelCallbackSync(STEER_LANGUAGE)
    .build();

Appending per call is correct here: the request builder is new for each model call, so the instruction does not accumulate. Writes to ctx.state() go through the callback context's event actions, and the next turn's fallback depends on them persisting; add a two-turn test that reads turn_lang back from the session service to confirm it on your ADK version. The locale_lang key is the default seeded from the user's locale at session creation, as described in the i18n article. Callback mechanics in general are covered in ADK Java callbacks.

Code-switching and mixed scripts

Code-switching, mixing two languages inside one sentence, is normal speech for hundreds of millions of people, and single-label detection handles it poorly. Romanised Hindi written in Latin script is a common example: a detector trained on Devanagari Hindi has little to go on and often returns English or UNKNOWN. Do not fight this with a lower threshold. Decide a policy instead: when detection is unknown, keep the session language, and let the model mirror the user's style through an instruction such as "if the user mixes languages, you may reply in the same mix". Then measure satisfaction for those conversations separately rather than letting them disappear into the English average.

Scripts are a cheaper signal than languages. Counting code points by Character.UnicodeScript tells you instantly that a message is Devanagari, Cyrillic or Han, which is enough to route between very different scripts even when the language within a script is ambiguous.

Cross-lingual retrieval

The hardest multilingual problem is usually knowledge, not chat: the user asks in Spanish and the answer is in a German manual. There are two designs. Translate-then-search translates the query into the corpus language and uses a monolingual index; it needs one extra model call per query and one index per corpus language. Cross-lingual search embeds queries and documents with a multilingual embedding model, which places text with the same meaning near each other regardless of language, so one index serves every query language. On Google Cloud, text-multilingual-embedding-002 and gemini-embedding-001 are documented as multilingual; check the current model list and quotas before choosing.

Either way, store each chunk's language as metadata at indexing time, using the same detector. It lets you filter (prefer chunks in the user's language when they exist), monitor coverage (how many Hindi questions are answered from English sources), and tell the model which language each passage is in. Evaluate retrieval per language pair: cross-lingual recall is typically lower than monolingual recall, and the gap varies by pair. The retrieval tool itself follows the pattern in RAG in ADK Java.

Native generation or a translation pivot

ApproachHow it worksStrengthsWeaknesses
Native generationModel reasons and replies directly in the user's languageOne call, natural phrasing, keeps nuanceQuality varies by language; harder to review
Translation pivotTranslate input to English, run the agent, translate the reply backOne English prompt and eval suite; predictable toolsTwo extra calls, latency, translation errors in both directions, loses tone
HybridNative for conversation; translate only retrieved passages or tool outputsBest answer quality when sources are in one languageMore moving parts to evaluate

Current large models generate well in widely spoken languages, so native generation is the usual default, with the pivot reserved for languages where your evaluation shows native quality is not good enough. Decide per language from measurements, not once for all of them.

Token cost differs by language

Tokenizers split text from different languages and scripts into very different numbers of tokens for the same meaning, so the same conversation can cost noticeably more, and fill the context window sooner, in one language than in another. Do not estimate this from published averages; measure it on your own traffic with the model's token-counting endpoint and record input and output tokens per turn tagged with turn_lang. The pattern for counting is in token counting across models. Use the result to size history truncation and summarisation thresholds per language, and to price plans honestly.

Worked example

A user with an English locale writes: "¿Cómo amplío el volumen RAID sin perder datos? Me sale esto:" followed by a fenced block containing an English error. The stripper removes the block; the remaining prose is well over 20 code points, and Lingua returns Spanish with a clear margin. The callback stores turn_lang = es and appends the Spanish reply rule. The model calls the retrieval tool with the Spanish question; the multilingual index returns two German manual sections and one English knowledge-base article. The model answers in Spanish, quotes the English error unchanged, and cites the German section with a one-line Spanish summary. The second model call, after the tool result, skips detection because lang_inv matches. The user's next message, "gracias", is below the floor, so the session keeps Spanish rather than reverting to the English locale.

A per-language evaluation matrix

Claiming a language means measuring it. Build an evaluation matrix with supported languages as rows and task types (answer from documents, tool use, refusal policy, multi-turn follow-up) as columns, and fill each cell with the same scoring you use for English. Do not create the non-English cases by machine-translating the English set and stopping there: translated prompts are unnaturally clean and miss the slang, code-switching and regional variants that real users send. Seed each language with real anonymised messages and have a fluent reviewer check a sample of graded outputs, because an LLM judge can also be weaker in some languages. Add detector accuracy as its own row, measured on labelled real messages, since every downstream decision depends on it. Running these suites in CI is covered in ADK Java evaluation in CI.

Failure modes

FailureCauseFix
Replies flip to English after a pasteDetector saw logs or code, not proseStrip fenced blocks, URLs and stack frames before detection
Short replies reset the languageTwo-word messages classified at low confidenceLength floor and minimum relative distance; fall back to session
Spanish and Portuguese confusedClose languages, all 75 candidates enabledRestrict candidates to supported languages; add labelled tests for the pair
Latency spike on first messageLanguage models loaded lazilyPreload models at startup; build the detector once
Different language inside one turnDetection re-run on each model callGuard on invocation id; store in state
Good chat, bad answers in one languageCross-lingual retrieval recall lower for that pairPer-pair retrieval eval; translate queries for weak pairs

Trade-offs

A library detector adds a dependency and memory for each language model in exchange for no extra model calls and stable results. Following the turn language serves bilingual users well but can surprise someone who writes one sentence in another language; following the locale is predictable but ignores what the user just showed you. Cross-lingual embeddings make one index serve all languages at some cost in recall; translation keeps recall high at the cost of a call per query. The safe pattern is to make each choice per language, backed by the evaluation matrix.

What to do next

  1. List your supported languages, and add Lingua restricted to exactly those, built once with preloaded models.
  2. Strip code, URLs and stack frames before detection, and set a length floor tuned on your traffic.
  3. Add the before-model callback with the invocation-id guard, and seed the session language from the locale.
  4. Tag every indexed chunk with its language, and measure retrieval recall for each query-corpus language pair.
  5. Record tokens per turn by language, and set truncation thresholds per language.
  6. Build the language-by-task evaluation matrix from real messages, with fluent review of a sample.
  7. Decide native or pivot per language from the matrix, and revisit the choice when you change models.
Key takeaway: Treat the language of each message as an observation separate from the user's locale. Detect it once per invocation from the user's own prose with a restricted, thresholded detector, fall back to the session language when the signal is weak, steer every model call from a callback, search across languages with language-tagged chunks, and prove each supported language with its own evaluation row.