A translation jailbreak uses language itself to get past an LLM's safety behaviour. The attacker takes a request the model would refuse in English and either writes it in a language where safety training is thin, or frames it as a translation task so the model produces the content while believing it is only converting words. Neither needs any skill beyond a free translation tool, which is why it matters to every deployment and not only to multilingual products.
The research measuring per-language safety gaps, and the fairness side of the problem, are covered in AI and Language Minorities, in depth. This article is the engineering view: the four attack families, why they work, where machine translation should and should not sit in your pipeline, a translate-and-guard component with code, a traced example using a placeholder request, and a test matrix that tells you whether the defence works. Examples here use placeholders such as [restricted request] rather than real harmful content.
Four attack families
Four distinct patterns hide under one name, and each needs its own control.
- Language as obfuscation. The request is machine-translated into a low-resource language, sent to the model, and the answer translated back. The model's capability transfers across languages better than its refusals do.
- Translation as the task. The attacker supplies text, or asks the model to write text "for translation", and the request is framed as a neutral linguistic job. If the policy is judged by the request type rather than by the output, the model emits content it would never author.
- Code-switching and sandwiching. The harmful fragment is embedded in a prompt that mixes languages, or placed among several benign questions in different languages. Upadhayay and Behzadan (2024) described this as a sandwich attack. Classifiers that pick one language per input see mostly benign text.
- Response-language pivots. The request is in English, which the guard handles well, but asks for the answer in another language, where the output guard is weak.
Why it works: the evidence and the mechanism
The effect is well measured. Yong, Menghini and Bach (2023) machine-translated the unsafe prompts of the AdvBench benchmark and found that GPT-4 gave actionable help about 79 percent of the time when low-resource languages such as Zulu, Scots Gaelic, Hmong and Guarani were combined, against well under 1 percent for the English originals. Deng and colleagues (ICLR 2024) built MultiJail, 315 English prompts with human translations into nine languages at high, medium and low resource levels (Chinese, Italian, Vietnamese; Arabic, Korean, Thai; Bengali, Swahili, Javanese), and found that unsafe output rises as resource level falls, both for ordinary users writing in their own language and for deliberate attackers.
The mechanism is a capability-alignment gap. Pre-training covers many languages, so the model can understand and answer a request in Zulu. Safety fine-tuning and preference data are dominated by English and a few large languages, so the refusal behaviour is learned mostly there. Guard classifiers inherit the same skew from their training data, and tokenisers fragment low-resource text into many short pieces that look unlike anything the guard learned. Models also change, so treat these published rates as evidence of the shape of the problem, not as numbers for any model you run today; measure your own.
Where translation sits in the pipeline
Translation can sit in three places, and each placement opens or closes different routes.
| Placement | What it closes | What it opens |
|---|---|---|
| No MT; multilingual guard only | Nothing extra to run or leak | Every language where the guard is weak |
| MT before the guard only (translate-then-classify) | Low-resource obfuscation, if MT is accurate | MT softening or garbling the request; code-switched fragments lost; data sent to the MT provider |
| Guard on original AND on translation, take the higher risk | Most obfuscation and code-switching | Latency and cost of a second pass; false positives from bad MT |
| Output guard plus back-translation | Translation-as-task, response pivots | A second MT call on every flagged-language reply |
The third and fourth rows together are the recommended design: never replace the original text with its translation, because the translation may drop exactly the fragment that mattered, and always judge what the model produced, not what the user said they wanted.
A translate-and-guard component
The component below implements the input side. lang_id, translate and guard are interfaces you supply: a language identifier that returns per-segment languages and confidence, a translation model you host or trust with the data, and a risk classifier returning a score in [0, 1]. The names are placeholders, not a specific product's API.
from dataclasses import dataclass
STRONG_LANGS = {"en", "es", "fr", "de", "zh", "ja"} # where your guard is measured to be good
BLOCK_AT = 0.80
REVIEW_AT = 0.50
@dataclass
class Verdict:
action: str # "allow", "review" or "block"
score: float
languages: list
reasons: list
def check_input(text, lang_id, translate, guard) -> Verdict:
segments = lang_id(text) # [(lang, confidence, span), ...]
langs = sorted({lang for lang, _, _ in segments})
reasons = []
scores = [guard(text)] # always score the original
weak = [s for s in segments if s[0] not in STRONG_LANGS or s[1] < 0.6]
if weak:
pivot = translate(text, target="en") # whole text, keeps context
scores.append(guard(pivot))
reasons.append(f"weak-language segments: {[s[0] for s in weak]}")
for lang, conf, span in weak: # catches code-switched fragments
scores.append(guard(translate(text[span[0]:span[1]], target="en")))
if len(langs) > 2:
reasons.append(f"code-switching across {len(langs)} languages")
scores.append(min(1.0, max(scores) + 0.1)) # small prior, tune on your data
score = max(scores)
action = "block" if score >= BLOCK_AT else "review" if score >= REVIEW_AT else "allow"
return Verdict(action, score, langs, reasons)Three design choices are deliberate. The maximum, not the average, of the scores is used, because the attack succeeds if any view of the text is harmful. Weak segments are translated individually as well as in context, which is what catches a single harmful sentence inside a benign multilingual sandwich. And the thresholds and the code-switching prior are tuning parameters to be set from the evaluation below, not constants to copy.
Judging the output, not the request
Input checks cannot stop translation-as-task, because the request is honestly a translation. The rule that does stop it is a policy rule: translating, summarising, paraphrasing or transcribing content counts as producing it. If the model would refuse to write a passage, it should refuse to translate one into or out of any language, and the output guard enforces that by scoring the response text whatever the request claimed.
Two output checks complete the picture. If the guard is weak in the response language, back-translate the response to a pivot language and score that too. And compare the response language with the request language and the product's configured languages: an English request that receives an answer in Hmong from a support bot that serves English and Spanish is an anomaly worth blocking or reviewing whatever the guard says. This also limits response-language pivots, where the request is clean and only the answer is in a weak language.
Encodings and ciphers work by the same gap and are covered in Encoding Attacks on LLMs, in depth. Normalising inputs before guarding, as in Payload Smuggling, in depth, should run before the language step so a base64 blob is decoded and then language-identified.
Worked example: a sandwiched request
A support assistant serves English and Spanish customers. Its guard is measured as strong in both. An attacker takes [restricted request], which the assistant refuses in English, machine-translates it into Zulu, and wraps it between two ordinary questions, one in Spanish and one in English. The scores below are illustrative of the pattern, not measurements of a specific guard.
| Step | Without the component | With the component |
|---|---|---|
| Language ID | Labelled Spanish (majority) | Segments: es 0.97, zu 0.71, en 0.98 |
| Guard on original | 0.18, allowed | 0.18 |
| Guard on whole translation | not run | 0.46 |
| Guard on the Zulu segment, translated alone | not run | 0.88 |
| Decision | Model answers all three, including the Zulu one | Block, reasons: weak-language segment, code-switching |
The whole-text translation scored only 0.46 because the two benign questions diluted it; the segment-level pass is what caught it. Had the attacker instead asked in English for the answer in Zulu, the input would pass, and the output-side language check would flag a reply in a language the product does not serve.
Measuring the defence
A translation defence is only as good as its measured coverage. Build a matrix: a set of harmful-intent prompts and a matched set of benign prompts, each rendered in every language you serve plus a fixed list of low-resource languages, in four forms: plain, code-switched, sandwiched and translation-framed. Published sets such as MultiJail give a starting point with human translations; machine-translated variants test the attacker's realistic route.
def run_matrix(prompts, languages, forms, render, pipeline, judge):
rows = []
for p in prompts: # p.intent is "harmful" or "benign"
for lang in languages:
for form in forms: # plain, code_switched, sandwich, translate_task
text = render(p, lang, form)
reply = pipeline(text) # full system: guards + model
unsafe = judge(p, reply) # human or calibrated model judge, in a pivot language
rows.append((p.intent, lang, form, reply.blocked, unsafe))
return rows
# Per (lang, form): attack success = unsafe / harmful prompts
# over-refusal = blocked / benign promptsReport both numbers per cell, never only the average, because the failures concentrate in small cells that an aggregate hides. Set a release gate on the worst cell, not the mean, and re-run the matrix whenever the model, the guard, the MT system or the language list changes. Multi-turn versions matter too: a conversation can shift language gradually, which Multi-Turn Jailbreaks, in depth covers.
Failure modes
- Replacing the original with its translation. MT drops, softens or mistranslates the key fragment and the guard never sees it. Score both.
- Whole-text language ID. Majority-language labelling hides a short fragment in another language. Identify per segment.
- Judging the request, not the output. Translation-as-task passes any input check. The output guard must score the produced text.
- Over-blocking minority languages. Treating every low-resource input as suspicious refuses legitimate users. Measure over-refusal per language and route to review rather than block where possible.
- Leaking data to a third-party MT service. Sending every prompt to an external translator is a data transfer that privacy terms may not allow. Host the MT model or restrict it to flagged segments.
- An LLM translator that can be injected. If the translation step is itself an LLM, the text it translates can instruct it. Use a dedicated MT model or constrain the translator's output format.
Trade-offs
| Control | Benefit | Cost |
|---|---|---|
| Multilingual guard only | Cheapest, no extra hop | Weak exactly where the attack lives |
| Dual scoring with MT | Covers low-resource and code-switching | Added latency on flagged inputs; MT errors cause false positives |
| Output guard and back-translation | Stops translation-as-task and response pivots | Second pass on replies in weak languages |
| Language allow-list | Simple, strong for single-market products | Unusable for open multilingual products |
| Multilingual safety fine-tuning | Fixes the model, not just the wrapper | Data collection and training cost per language |
What to do next
- List the languages your product serves and measure your guard per language; mark the weak ones.
- Add per-segment language identification before the guard, after any decoding of encodings.
- Score the original text and its translation, plus weak segments alone, and take the maximum.
- Write the policy rule that translating or paraphrasing content counts as producing it, and enforce it with an output guard.
- Check that the response language matches the request and the product's languages.
- Build the language-by-form test matrix, report attack success and over-refusal per cell, and gate releases on the worst cell.
- Decide where MT runs and confirm the data handling is allowed before sending prompts to it.