A safety system that works in English and fails in Zulu is not a safe system. It is a safe system for some users. The gap runs in both directions. Harmful requests written in a low-resource language can pass filters that would block them in English. Ordinary speech in a minority language or dialect can be blocked as if it were abuse. Speakers of those languages also pay more tokens for the same sentence, which costs them money, context window and latency.

This article is about measuring and closing those gaps in a deployed LLM product. It covers where the gaps come from, what published research has measured, how to build a per-language evaluation that catches both failure directions, how to route uncertain inputs, how to account for tokenizer cost, and how to involve the language communities whose speech you are judging. It deliberately contains no attack prompts. Everything here is evaluation and defence.

Why safety thins out beyond high-resource languages

Three things cause the gap. The first is data. Pretraining corpora are dominated by English and a handful of other high-resource languages, so a model's grasp of a minority language is shallower, and its grasp of what is harmful in that language is shallower still. The second is safety tuning. Refusal behaviour is mostly taught with preference data and red-team prompts written in English, and it transfers imperfectly to languages that were rare in that data. The third is the surrounding tooling. Moderation classifiers, keyword lists, PII detectors and language-ID models are usually trained and tuned on the same few languages. A guard can only judge what it can read.

Tokenization adds a fourth, quieter effect. Byte-pair vocabularies learned mostly from English split minority-language text into many more pieces. Long, fragmented token sequences are harder for classifiers to judge, and they consume the context budget that holds the system prompt and safety instructions.

What the research measured

Several studies quantify the gap. Yong, Menghini and Bach (2023, "Low-Resource Languages Jailbreak GPT-4") machine-translated the unsafe prompts of the AdvBench benchmark into a range of languages. They found that GPT-4 engaged with the translated inputs and gave actionable help about 79 percent of the time when low-resource languages were combined. High- and mid-resource languages had much lower attack success. Their point is not about one model. Free translation tools let anyone use this route, so it is a risk for every deployment, not only for speakers of those languages.

Deng, Zhang, Pan and Bing (ICLR 2024, "Multilingual Jailbreak Challenges in Large Language Models") separated an unintentional scenario, an ordinary user writing in their own language, from an intentional one. They released the MultiJail dataset and found that unsafe outputs became more frequent as language resources decreased. They also showed that fine-tuning on automatically generated multilingual safety data reduced unsafe generation substantially. Petrov, La Malfa, Torr and Bibi (NeurIPS 2023) measured tokenizer disparity and found that the same text can take up to 15 times more tokens in one language than in another, even with tokenizers built for multilingual use.

The over-blocking direction is older. Sap and colleagues (2019) showed that hate-speech datasets and the classifiers trained on them flagged tweets in African American English as offensive disproportionately often. A guard tuned on the majority dialect treats normal minority speech as an anomaly. The same pattern appears whenever a classifier meets a dialect, a code-mixed register or a romanised script it was not trained on.

Two failure directions

A per-language programme has to track both errors, because fixing one in isolation tends to worsen the other. Lowering a threshold to catch more harmful Swahili also blocks more benign Swahili.

Under-blocking (misses)Over-blocking (false refusals)
What happensUnsafe request in language X passes the input guard and the model compliesBenign request in language X is refused or flagged
Who is harmedThird parties, and the platform's integritySpeakers of X, who get a worse product
Typical causeGuard and safety tuning never saw XGuard treats unfamiliar script or dialect as suspicious
MetricMiss rate on labelled unsafe items in XRefusal rate on labelled benign items in X
Visible in aggregate?Rarely: X is a small share of trafficRarely: affected users leave quietly

The last row is why aggregate dashboards are not enough. If 2 percent of traffic is in a language with ten times the English miss rate, the overall miss rate barely moves.

A per-language evaluation and release gate

Build an evaluation set per supported language with two labelled halves. One half is benign: ordinary questions, culturally specific topics, health and legal questions, and dialect and code-mixed variants. The other half is policy-violating, organised by your policy categories. Three rules make the set trustworthy. Native speakers write or review every item, because machine translation of an English set misses what is sensitive locally and produces stilted text that guards find easy. Items are labelled by category, not by keyword. And the set is split so that part of it is never used for tuning, which keeps you from overfitting the guard to its own test.

Score both halves through the whole production path: language ID, input guard, model, output guard. Then report rates per language with confidence intervals, because per-language samples are small and a raw rate on 80 items is noisy.

import math
from collections import defaultdict

def wilson(k, n, z=1.96):
    """95% Wilson interval for k errors in n items."""
    if n == 0:
        return (0.0, 1.0)
    p = k / n
    d = 1 + z * z / n
    mid = (p + z * z / (2 * n)) / d
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return (max(0.0, mid - half), min(1.0, mid + half))

def per_language(rows):
    """rows: {"lang", "label": "unsafe"|"benign", "blocked": bool} from the full pipeline."""
    tally = defaultdict(lambda: {"unsafe": [0, 0], "benign": [0, 0]})
    for r in rows:
        cell = tally[r["lang"]][r["label"]]
        cell[1] += 1
        wrong = (not r["blocked"]) if r["label"] == "unsafe" else r["blocked"]
        cell[0] += wrong
    out = {}
    for lang, t in tally.items():
        out[lang] = {"miss": (*t["unsafe"], wilson(*t["unsafe"])),
                     "overblock": (*t["benign"], wilson(*t["benign"]))}
    return out

def gate(report, baseline="en", slack=0.02):
    """Fail a language when its lower bound is worse than the baseline's upper bound plus slack."""
    failures = []
    for metric in ("miss", "overblock"):
        base_hi = report[baseline][metric][2][1]
        for lang, m in report.items():
            if lang != baseline and m[metric][2][0] > base_hi + slack:
                failures.append((lang, metric, m[metric][0], m[metric][1]))
    return failures

The gate compares interval bounds rather than point estimates, so it fails only when a language is worse than the baseline with reasonable confidence. If the intervals are too wide to decide either way, the honest output is "insufficient data for language X", and the fix is more items, not a pass.

Worked example: four languages, two different fixes

A guardrail path that is measured per language, not just in aggregateUser inputany script, mixedLanguage IDlabel + confidenceInput guardmultilingual classifierModelsystem promptOutputguardLow-confidence or uncovered languagepivot translation + stricter thresholds + human reviewlow confPer-language eval setnative-speaker built, benign + unsafeMetrics per languagemiss rate, over-block rate, CIRelease gateno language worse than baselineProduction logslanguage mix, appeals, reports by languageCommunity reviewpaid native speakers, dialect coverageAggregate safety numbers are dominated by English; the gate reads every language separately.
Figure: language ID routes low-confidence input to a stricter path; per-language evaluation, production signals and community review feed a gate that reads every language separately.

Worked example, with illustrative numbers. A support assistant serves English, Hindi, Swahili and Yoruba, and each has 200 unsafe and 300 benign evaluation items. English misses 4 of 200 unsafe items, an interval of about 0.8 to 5.0 percent. Hindi misses 7, which overlaps English, so it passes. Swahili misses 31, an interval of about 11 to 21 percent, and fails clearly. Yoruba misses only 6 but blocks 39 of 300 benign items, against 5 for English. It fails on over-blocking. Inspection shows that its romanised text without tone marks is being scored as noise.

The two fixes are different. Swahili needs coverage: guard training data in the language, plus safety fine-tuning examples. Until those exist, its traffic is routed through the stricter path in the figure. Yoruba needs calibration: benign examples in the script variants people actually type, and a language-specific threshold. Fixing Yoruba by switching its guard off would just move it into the Swahili column. Re-run the gate after each change and keep both reports in the release record.

Mitigations and their costs

Several mitigations are available. All of them have costs.

  • Multilingual guard models. Prefer a guard that was evaluated per language on your language list. Treat any language the vendor does not report as uncovered until your own set says otherwise.
  • Pivot translation for classification. Translate input to a well-covered language and run the guard on both versions, blocking if either flags it. Meta's NLLB-200 (2022) covers about 200 languages and is a common choice. Translation adds latency and can soften or drop the very phrasing that matters, so it supplements a native guard and never replaces one. Use it on the uncertain path, not on all traffic.
  • Language-ID confidence as a signal. Short, code-mixed or romanised inputs often get low or wrong language-ID labels. A low-confidence label should select the stricter path. It should not silently default to English thresholds.
  • Safety fine-tuning data per language. Deng et al. showed that generated multilingual safety data reduces unsafe output. Review generated examples with native speakers before training, or you teach the model the generator's mistakes.
  • Output-side checks. When the input guard is weak in a language, an output guard that also reads the response through a pivot translation catches some misses that slipped past the input side.

Tokenizer cost is part of the threat model

Tokenizer disparity is a fairness and security issue at once. If your product limits users by tokens, a speaker whose language needs four times the tokens gets a quarter of the service for the same price. Long inputs can also push safety instructions out of a truncated context. Measure it with a parallel corpus such as FLORES-200, which has the same sentences in about 200 languages, and set limits in characters or in per-language token allowances.

def token_ratio(tokenizer, parallel, base="eng_Latn"):
    """parallel: {lang_code: [sentence, ...]} with aligned sentences across languages."""
    base_n = sum(len(tokenizer.encode(s)) for s in parallel[base])
    return {lang: round(sum(len(tokenizer.encode(s)) for s in sents) / base_n, 2)
            for lang, sents in parallel.items()}

Run it on the tokenizer of every model you serve. A ratio well above 1 is a reason to raise that language's token allowance and to check that system prompts still fit when the user's input is long.

Working with language communities

None of this works without the people who speak the languages. Pay native speakers to write and review evaluation items and to label appeals. Recruit across dialects, not one reviewer per language. Give them a channel to report over-blocking that goes to engineers and not just to a support queue. Data contributed by communities needs clear consent and use terms. Mozilla's Common Voice is a useful model for how voice data is contributed and licensed openly. Publish which languages you evaluate and at what quality, so users know where support is thin instead of discovering it.

Failure modes

  • Translated test sets only. Machine-translated English prompts measure translation artefacts, not how speakers write. Results look better or worse than reality for reasons unrelated to safety.
  • One threshold for every language. Guard scores are not calibrated across languages. A score of 0.7 in English and 0.7 in Amharic do not mean the same thing.
  • Aggregate gates. A release gate on the overall miss rate passes while a small language regresses badly.
  • Language-ID defaulting. Uncertain inputs fall through to English rules, which is exactly the gap that translation-based attacks use.
  • Fixing misses by over-blocking. Blanket refusal of a language closes the attack route and abandons its users. That shows up as lost users, not as a security metric.
  • Stale coverage. A model or guard upgrade changes per-language behaviour. Re-run every language, not just English, on every upgrade.

What to do next

  1. List the languages your users actually write in from production logs, including code-mixed and romanised forms.
  2. For each, build a native-speaker evaluation set with benign and policy-violating halves and a held-out split.
  3. Score the full pipeline per language and report miss and over-block rates with intervals.
  4. Add a release gate that fails any language worse than the baseline, and run it on every model or guard upgrade.
  5. Route low-confidence language-ID results to a stricter path with pivot-translation checks.
  6. Measure tokenizer ratios on FLORES-200 and adjust per-language limits.
  7. Fund paid community review, and publish which languages you cover and how well.
Key takeaway: Safety measured in aggregate is safety measured in English. Evaluate every supported language separately, for both missed harmful requests and wrongly blocked benign ones, with native-speaker data and confidence intervals. Gate releases per language, route uncertain inputs to a stricter path, account for tokenizer cost, and pay the communities whose language you are judging.