An agent deployed to more than one country receives messages in more than one language, and the language a user writes in is not always the language of their account. A Spanish speaker uses an English-locale corporate account; a German engineer pastes an English stack trace and asks about it in German; a support queue in India receives Hindi, English and Hinglish in the same hour. Handling that well is a content problem: detecting the language of each turn, deciding how to answer, finding knowledge written in other languages, and proving quality in every language you claim to support.
This article builds that layer in ADK Java. It is the companion to ADK Java internationalization, which covers the product side: resolving a locale once, ICU message bundles, locale-neutral tools, a language-drift metric and a tool for switching reply language. Those are not repeated here. This article covers per-turn detection with a callback, mixed-language input, cross-lingual retrieval, native generation versus a translation pivot, token cost, and a per-language evaluation matrix.
Locale versus turn language
Keep two values apart. The locale is a product setting: it controls dates, currency, fixed strings and the default reply language, and it changes rarely. The turn language is an observation: the language the user actually wrote in this message. Most of the time they agree. When they disagree, a good default is to answer in the turn language, because the user just demonstrated it, while keeping formatting and fixed strings in the locale until the user switches explicitly.
Two exceptions matter. A message that is mostly pasted material (logs, code, an email in another language) says nothing about the user's language; and a very short message such as "ok" or "gracias" carries too little signal to override the session. Both cases fall back to the language already recorded for the session.
Architecture
The flow has one writer and many readers. The detector runs once per invocation, on the user's own text, and writes the result to session state. A before-model callback reads it on every model call and appends a short instruction to the request. Tools such as retrieval read the same state to decide how to search. Evaluation runs offline over the same code path, in every supported language.
Detecting the language of a turn
Use a language-identification library rather than asking the model, which costs a call and is easily swayed by the instruction itself. Lingua (Maven com.github.pemistahl:lingua, version 1.2.2 at the time of writing) is a JVM library that is accurate on short text. Restrict it to the languages you support: fewer candidates means fewer confusions between close pairs such as Spanish and Portuguese, and less memory, because each language model is loaded separately. A minimum relative distance makes the detector return Language.UNKNOWN instead of guessing when the top two candidates are close.
import static com.github.pemistahl.lingua.api.Language.*;
import com.github.pemistahl.lingua.api.Language;
import com.github.pemistahl.lingua.api.LanguageDetector;
import com.github.pemistahl.lingua.api.LanguageDetectorBuilder;
import java.util.Locale;
import java.util.Optional;
import java.util.regex.Pattern;
public final class TurnLanguage {
private static final LanguageDetector DETECTOR = LanguageDetectorBuilder
.fromLanguages(ENGLISH, SPANISH, PORTUGUESE, GERMAN, FRENCH, HINDI)
.withMinimumRelativeDistance(0.25)
.withPreloadedLanguageModels() // pay the load cost at startup, not on a user turn
.build();
private static final Pattern FENCED = Pattern.compile("```.*?```", Pattern.DOTALL);
private static final Pattern NOISE = Pattern.compile(
"https?://\\S+|\\S+@\\S+|\\b[\\w.]+Exception\\b|^\\s*(at|>) .*$", Pattern.MULTILINE);
/** ISO 639-1 code of the user's own prose, or empty when the signal is too weak. */
public static Optional<String> detect(String message) {
String prose = NOISE.matcher(FENCED.matcher(message).replaceAll(" ")).replaceAll(" ").strip();
if (prose.codePointCount(0, prose.length()) < 20) {
return Optional.empty(); // "ok", "gracias": keep the session language
}
Language lang = DETECTOR.detectLanguageOf(prose);
if (lang == Language.UNKNOWN) {
return Optional.empty();
}
return Optional.of(lang.getIsoCode639_1().toString()); // "es", "de", ...
}
}The stripping step does most of the work on technical traffic. Without it, the German engineer's message with forty lines of English stack trace is classified as English. The 20-code-point floor and the 0.25 distance are starting points; tune both against labelled samples of your own traffic, because the right values depend on how short your users' messages are. Build the detector once as a static field: it is thread-safe and expensive to create.
Detect once, steer every model call
ADK runs the before-model callback on every model call in an invocation, and one user turn with two tool round trips makes three model calls. callbackContext.userContent() returns the same message each time, so detecting inside the callback without a guard runs the detector three times and can, in principle, produce different answers if thresholds sit on a boundary. Guard on the invocation id and store the result in state.
import com.google.adk.agents.Callbacks;
import com.google.adk.agents.LlmAgent;
import com.google.genai.types.Content;
import com.google.genai.types.Part;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.stream.Collectors;
static final Map<String, String> NAMES = Map.of(
"en", "English", "es", "Spanish", "pt", "Portuguese",
"de", "German", "fr", "French", "hi", "Hindi");
static final Callbacks.BeforeModelCallbackSync STEER_LANGUAGE = (ctx, request) -> {
Map<String, Object> state = ctx.state();
if (!ctx.invocationId().equals(state.get("lang_inv"))) { // first model call of this turn
String text = ctx.userContent().flatMap(Content::parts).orElse(List.of()).stream()
.map(p -> p.text().orElse(""))
.collect(Collectors.joining(" "));
String previous = (String) state.getOrDefault("turn_lang", state.getOrDefault("locale_lang", "en"));
state.put("turn_lang", TurnLanguage.detect(text).orElse(previous));
state.put("lang_inv", ctx.invocationId());
}
String lang = NAMES.getOrDefault((String) state.get("turn_lang"), "English");
request.appendInstructions(List.of(
"Write your reply in " + lang + ". Keep code, identifiers, product names and quoted "
+ "error messages unchanged. If you quote a source written in another language, "
+ "quote it in the original and give a short " + lang + " summary."));
return Optional.empty(); // continue to the model
};
LlmAgent agent = LlmAgent.builder()
.name("support_agent")
.model("gemini-2.5-flash")
.instruction("You are a support agent for Acme storage products. ...")
.beforeModelCallbackSync(STEER_LANGUAGE)
.build();Appending per call is correct here: the request builder is new for each model call, so the instruction does not accumulate. Writes to ctx.state() go through the callback context's event actions, and the next turn's fallback depends on them persisting; add a two-turn test that reads turn_lang back from the session service to confirm it on your ADK version. The locale_lang key is the default seeded from the user's locale at session creation, as described in the i18n article. Callback mechanics in general are covered in ADK Java callbacks.
Code-switching and mixed scripts
Code-switching, mixing two languages inside one sentence, is normal speech for hundreds of millions of people, and single-label detection handles it poorly. Romanised Hindi written in Latin script is a common example: a detector trained on Devanagari Hindi has little to go on and often returns English or UNKNOWN. Do not fight this with a lower threshold. Decide a policy instead: when detection is unknown, keep the session language, and let the model mirror the user's style through an instruction such as "if the user mixes languages, you may reply in the same mix". Then measure satisfaction for those conversations separately rather than letting them disappear into the English average.
Scripts are a cheaper signal than languages. Counting code points by Character.UnicodeScript tells you instantly that a message is Devanagari, Cyrillic or Han, which is enough to route between very different scripts even when the language within a script is ambiguous.
Cross-lingual retrieval
The hardest multilingual problem is usually knowledge, not chat: the user asks in Spanish and the answer is in a German manual. There are two designs. Translate-then-search translates the query into the corpus language and uses a monolingual index; it needs one extra model call per query and one index per corpus language. Cross-lingual search embeds queries and documents with a multilingual embedding model, which places text with the same meaning near each other regardless of language, so one index serves every query language. On Google Cloud, text-multilingual-embedding-002 and gemini-embedding-001 are documented as multilingual; check the current model list and quotas before choosing.
Either way, store each chunk's language as metadata at indexing time, using the same detector. It lets you filter (prefer chunks in the user's language when they exist), monitor coverage (how many Hindi questions are answered from English sources), and tell the model which language each passage is in. Evaluate retrieval per language pair: cross-lingual recall is typically lower than monolingual recall, and the gap varies by pair. The retrieval tool itself follows the pattern in RAG in ADK Java.
Native generation or a translation pivot
| Approach | How it works | Strengths | Weaknesses |
|---|---|---|---|
| Native generation | Model reasons and replies directly in the user's language | One call, natural phrasing, keeps nuance | Quality varies by language; harder to review |
| Translation pivot | Translate input to English, run the agent, translate the reply back | One English prompt and eval suite; predictable tools | Two extra calls, latency, translation errors in both directions, loses tone |
| Hybrid | Native for conversation; translate only retrieved passages or tool outputs | Best answer quality when sources are in one language | More moving parts to evaluate |
Current large models generate well in widely spoken languages, so native generation is the usual default, with the pivot reserved for languages where your evaluation shows native quality is not good enough. Decide per language from measurements, not once for all of them.
Token cost differs by language
Tokenizers split text from different languages and scripts into very different numbers of tokens for the same meaning, so the same conversation can cost noticeably more, and fill the context window sooner, in one language than in another. Do not estimate this from published averages; measure it on your own traffic with the model's token-counting endpoint and record input and output tokens per turn tagged with turn_lang. The pattern for counting is in token counting across models. Use the result to size history truncation and summarisation thresholds per language, and to price plans honestly.
Worked example
A user with an English locale writes: "¿Cómo amplío el volumen RAID sin perder datos? Me sale esto:" followed by a fenced block containing an English error. The stripper removes the block; the remaining prose is well over 20 code points, and Lingua returns Spanish with a clear margin. The callback stores turn_lang = es and appends the Spanish reply rule. The model calls the retrieval tool with the Spanish question; the multilingual index returns two German manual sections and one English knowledge-base article. The model answers in Spanish, quotes the English error unchanged, and cites the German section with a one-line Spanish summary. The second model call, after the tool result, skips detection because lang_inv matches. The user's next message, "gracias", is below the floor, so the session keeps Spanish rather than reverting to the English locale.
A per-language evaluation matrix
Claiming a language means measuring it. Build an evaluation matrix with supported languages as rows and task types (answer from documents, tool use, refusal policy, multi-turn follow-up) as columns, and fill each cell with the same scoring you use for English. Do not create the non-English cases by machine-translating the English set and stopping there: translated prompts are unnaturally clean and miss the slang, code-switching and regional variants that real users send. Seed each language with real anonymised messages and have a fluent reviewer check a sample of graded outputs, because an LLM judge can also be weaker in some languages. Add detector accuracy as its own row, measured on labelled real messages, since every downstream decision depends on it. Running these suites in CI is covered in ADK Java evaluation in CI.
Failure modes
| Failure | Cause | Fix |
|---|---|---|
| Replies flip to English after a paste | Detector saw logs or code, not prose | Strip fenced blocks, URLs and stack frames before detection |
| Short replies reset the language | Two-word messages classified at low confidence | Length floor and minimum relative distance; fall back to session |
| Spanish and Portuguese confused | Close languages, all 75 candidates enabled | Restrict candidates to supported languages; add labelled tests for the pair |
| Latency spike on first message | Language models loaded lazily | Preload models at startup; build the detector once |
| Different language inside one turn | Detection re-run on each model call | Guard on invocation id; store in state |
| Good chat, bad answers in one language | Cross-lingual retrieval recall lower for that pair | Per-pair retrieval eval; translate queries for weak pairs |
Trade-offs
A library detector adds a dependency and memory for each language model in exchange for no extra model calls and stable results. Following the turn language serves bilingual users well but can surprise someone who writes one sentence in another language; following the locale is predictable but ignores what the user just showed you. Cross-lingual embeddings make one index serve all languages at some cost in recall; translation keeps recall high at the cost of a call per query. The safe pattern is to make each choice per language, backed by the evaluation matrix.
What to do next
- List your supported languages, and add Lingua restricted to exactly those, built once with preloaded models.
- Strip code, URLs and stack frames before detection, and set a length floor tuned on your traffic.
- Add the before-model callback with the invocation-id guard, and seed the session language from the locale.
- Tag every indexed chunk with its language, and measure retrieval recall for each query-corpus language pair.
- Record tokens per turn by language, and set truncation thresholds per language.
- Build the language-by-task evaluation matrix from real messages, with fluent review of a sample.
- Decide native or pivot per language from the matrix, and revisit the choice when you change models.