A prompt that works well in English often degrades in other languages, and not always in obvious ways. The model may answer in English when the user wrote in German, drift into Spanish when asked for Portuguese, switch from a formal to a casual register halfway through, translate a product name that legal says must never be translated, or format 03/04 as a date that means different days in different countries. None of these show up in an English-only test set.

Multilingual prompting is the engineering that prevents this. Its core idea is a division of labour: code decides which language and locale a response must use, the model writes in that language, and code checks the result. This article explains why models behave differently across languages, then builds a template, a validator and an evaluation plan around one worked example carried through German, Japanese and Brazilian Portuguese.

Advertisement

Why languages behave differently

Large language models learn from text, and the training mix is uneven: English dominates, a group of widely published languages follows, and many languages have far less data. Quality, instruction-following and safety behaviour all track that imbalance, so a model can follow a complex instruction in English and follow the same instruction less reliably in a lower-resource language. Tokenizers add a cost dimension. The same sentence can take noticeably more tokens in some scripts than in English, which raises latency and price and uses more of the context window; the mechanics are in multilingual tokenization.

Two practical consequences follow. First, you cannot assume an English evaluation predicts quality elsewhere; every language you ship needs its own test set. Second, prompt choices that are neutral in English, such as where instructions sit, which examples you include and how you phrase the output format, can shift behaviour in other languages, so they should be decided deliberately rather than inherited.

Instruction language versus output language

Keep two questions separate: what language the instructions are written in, and what language the answer must be in. A common, maintainable choice is one instruction language, often English, for every locale, plus an explicit statement of the output language. One template is easier to review, version and test than a dozen translations that drift apart. Fully native prompts can help for some languages and tasks, but they multiply maintenance, so adopt them per language only when an evaluation shows a gain.

Whatever you choose, never rely on the model to infer the output language from the user's message. Inference fails on short inputs, on mixed-language inputs (a Japanese question containing an English error message) and when retrieved context is in a different language from the user. State the language by name and by tag, for example German (de), because names are less ambiguous than bare codes, and state the register too: formal or informal address, polite forms, regional variety. Brazilian and European Portuguese differ enough that users notice.

Advertisement

The request pipeline

The diagram shows where each decision lives. The target language comes from an explicit user or account setting when there is one, and from language detection on the input only as a fallback. The template is filled with the language pin, register, glossary and few-shot examples written in that language. After the call, a validator checks the schema, detects the language of the free-text fields and checks protected terms. Dates, numbers and currency are formatted in code from structured values, never by the model.

User inputany languageDetect languagefastText lid.176Resolve targetuser setting, then detectionAssemble promptpin, glossary, few-shotsModel callJSON, values localisedValidateschema, language, glossaryRepair oncerestate the target languagefailsretryFormat in codedates, numbers, currencypassesResponseplus per-language metricsThe model writes the language; code decides which language and checks that it got it.
Code resolves the target language and checks the result; the model only writes it.

A template that pins language and protects terms

The template below keeps the JSON keys and enum values in English so the parsing code is identical for every locale, and asks for the free-text values in the target language. Few-shot examples come from a per-language pool, because examples in the target language anchor both the language and the register far more strongly than an instruction alone; few-shot prompting covers how to choose them. The glossary lists terms that must survive untouched, such as product names, feature names and legal phrases.

LANGS = {
    "de": dict(name="German", register="Use the formal 'Sie' form."),
    "ja": dict(name="Japanese", register="Use polite desu/masu style."),
    "pt-BR": dict(name="Brazilian Portuguese", register="Use 'voce', not European forms."),
}
GLOSSARY = ["Cassia Cloud", "Workspace", "SSO"]  # never translate these

SYSTEM = """You summarise customer support tickets for agents.
Return JSON with keys "summary", "sentiment", "next_step".
Keys and the sentiment enum (positive|neutral|negative) stay in English.
Write the values of "summary" and "next_step" in {name} ({tag}). {register}
Keep these terms exactly as written: {glossary}.
Do not format dates or amounts; copy them as they appear in the ticket."""

def build_messages(ticket: str, tag: str, examples: dict[str, list]) -> list[dict]:
    lang = LANGS[tag]
    system = SYSTEM.format(name=lang["name"], tag=tag, register=lang["register"],
                           glossary=", ".join(GLOSSARY))
    msgs = [{"role": "system", "content": system}]
    for user, assistant in examples.get(tag, []):  # few-shots in the target language
        msgs += [{"role": "user", "content": user}, {"role": "assistant", "content": assistant}]
    msgs.append({"role": "user", "content": f"<ticket>\n{ticket}\n</ticket>"})
    return msgs

The instruction to copy dates and amounts verbatim is deliberate. If the model reformats 2026-03-04 into a local style, it may also silently reinterpret it. Keep these values structured in your data and render them in the presentation layer with a locale library such as ICU, Babel or the platform's Intl API, which are tested against real locale data. The model is not.

Validate the output language

Instructions reduce language drift but do not eliminate it, so check every response. Language identification is a solved, cheap problem for text of reasonable length. The fastText lid.176 model identifies 176 languages and runs in microseconds per string on a CPU. Two details matter: its predict method processes one line at a time, so strip newlines first, and it returns base languages, so it cannot tell Brazilian from European Portuguese. Register and regional variety need an evaluation, not a detector.

import json
import fasttext  # pip install fasttext; model file lid.176.bin from fasttext.cc

LID = fasttext.load_model("lid.176.bin")
BASE = {"de": "de", "ja": "ja", "pt-BR": "pt"}  # lid.176 predicts base languages only

def detect(text: str) -> tuple[str, float]:
    labels, probs = LID.predict(text.replace("\n", " "), k=1)  # one line at a time
    return labels[0].removeprefix("__label__"), float(probs[0])

def validate(raw: str, tag: str, min_conf: float = 0.6) -> list[str]:
    errors = []
    try:
        out = json.loads(raw)
    except json.JSONDecodeError:
        return ["not JSON"]
    if out.get("sentiment") not in {"positive", "neutral", "negative"}:
        errors.append("sentiment enum")
    for field in ("summary", "next_step"):
        value = out.get(field, "")
        if len(value) >= 20:  # very short strings give unreliable detections
            lang, conf = detect(value)
            if lang != BASE[tag] or conf < min_conf:
                errors.append(f"{field}: detected {lang} ({conf:.2f})")
    for term in GLOSSARY:
        if term.lower() in raw.lower() and term not in raw:
            errors.append(f"glossary term altered: {term}")
    return errors

On failure, retry once with a repair message that quotes the error and restates the target language, then fall back: return the answer with a flag, or route it to a human queue. Log every failure with the language tag, because the failure rate per language is the metric that tells you where the prompt is weak. Structured output validation in general is covered in structured output.

Worked example: one ticket, three languages

A support tool summarises incoming tickets for agents who read them in their own language. The same ticket, a customer who cannot sign in through SSO to a Cassia Cloud Workspace after a password reset on 03/04, is routed to agents in Germany, Japan and Brazil. The first version used a single English prompt with the line "answer in the user's language" and no validation. Testing it on a set of translated tickets showed four distinct failures.

LanguageObserved failureFix
GermanSummary switched from Sie to du halfway throughRegister stated in the pin; German few-shots written in formal style
JapaneseSSO and Workspace rendered in katakanaGlossary in the system prompt; validator flags altered terms
Brazilian PortugueseOccasional European Portuguese vocabulary; once, SpanishPin names the variety; validator catches Spanish; human review samples for variety
All03/04 rewritten as 3 April and as 4 MarchModel copies dates verbatim; the UI renders the structured date

The fixed version resolves the language from the agent's profile, not from the ticket, because the reader is the agent. That one change matters more than any wording: a Japanese agent reading a ticket written in English still wants a Japanese summary, and "the user's language" was ambiguous between the customer and the agent. The lesson generalises. Decide whose language the output is for, write it into code, and test the cases where the input language and the target language differ.

Retrieval across languages

Retrieval-augmented systems add a second axis: the documents may be in a different language from the question. There are two broad designs. Multilingual embedding models place text from different languages in one vector space, so a Japanese question can retrieve an English document directly. The alternative is to translate the query into the corpus language before retrieval, which works with an existing monolingual index at the cost of an extra call and translation errors in short queries. Either way, the model then answers in the target language from source text in another, which is a translation task hidden inside an answering task. Tell it so: quote sources in their original language, answer in the target language, and keep citations pointing at the original passages.

Evaluating per language

Build a test set for each shipped language rather than translating your English set once and forgetting it. Machine-translated test inputs are a reasonable start, but they are cleaner than real user text, which mixes languages, uses slang and contains typos, so add real samples as soon as you have them. Track at least four metrics per language: language-match rate from the validator, schema-valid rate, glossary preservation and a task-quality score. LLM-as-judge grading can scale the quality score, but the judge also varies by language, so calibrate it against native-speaker ratings for each language before trusting it. The general method is in prompt evaluations.

Gate releases per language. A prompt change that improves English and German but regresses Japanese should not ship to Japanese users, and with per-language metrics that decision becomes a configuration choice rather than a rollback.

Failure modes and trade-offs

  • Language drift in long outputs. Responses start correctly and slide into English after a quoted English passage. Restate the pin after long context and validate each free-text field.
  • Detection on short strings. A two-word answer gives an unreliable language guess. Skip detection below a length threshold and rely on the pin.
  • Translated keys. The model localises JSON keys or enum values and the parser fails only for some languages. Keep keys in English and say so explicitly.
  • Token cost surprises. The same feature costs more and hits context limits sooner in some scripts. Budget tokens per language, not globally.
  • Safety asymmetry. Refusal and moderation behaviour can differ between languages. Include safety cases in every language's test set.
  • Native prompts versus one template. Native prompts may help a weak language but multiply maintenance. Use them per language, justified by evaluation.

What to do next

  1. Decide, in code, whose language each output is for, and pass the target language and register explicitly into every prompt.
  2. Keep JSON keys and enums in one language and render dates, numbers and currency with a locale library.
  3. Write a glossary of protected terms and add a validator check for each.
  4. Build target-language few-shot pools for every shipped language.
  5. Add fastText language detection to the response path with one repair retry and logging by language.
  6. Create a per-language test set, track language-match, schema, glossary and quality per language, and gate releases per language.
Key takeaway: Multilingual prompting works when code owns the decisions and the model only writes. Resolve the target language and register explicitly, pin them by name in the prompt, keep machine-readable fields in one language, protect product terms with a glossary, format locale-sensitive values outside the model, validate the output language on every response, and evaluate and release each language separately.