Border and immigration agencies use AI at nearly every step a traveller or applicant passes through: screening advance passenger data, matching faces at automated gates, checking documents, triaging visa and asylum files, and translating interviews. These systems decide who is questioned, delayed or refused, and they act on people who often cannot see the system, cannot easily contest it, and may be in a vulnerable position. That combination makes border AI one of the hardest settings for secure and fair AI engineering.

This article is written for engineers, security teams and governance staff building or assessing such systems. It covers where AI sits in the pipeline, the base-rate arithmetic that dominates screening, how biometric matching errors behave, the specific risks of language models in casework, a threat model, the EU AI Act provisions that apply, and the records and monitoring that make decisions accountable. It is not a news summary; it is a design guide.

Where AI sits in the border pipeline

The main uses, roughly in the order a traveller meets them:

  • Pre-arrival screening. Advance passenger information and booking data are checked against watchlists and risk rules, sometimes with statistical models, before the traveller boards.
  • Travel authorisation. Visa-exempt travellers apply online and the application is checked automatically against databases and risk indicators. The EU's ETIAS is one example; it is scheduled for the last quarter of 2026, so check the official ETIAS site for the confirmed date.
  • Biometric verification at the border. Automated gates compare a live face image with the chip photo in the passport (1:1). The EU Entry/Exit System, which began operating on 12 October 2025 with a six-month progressive rollout, records facial images and fingerprints of third-country nationals on short stays.
  • Identification. A face or fingerprint is searched against a database of many people (1:N), for watchlists or to establish identity.
  • Document checks and casework. Models authenticate documents, triage applications, summarise files and translate, increasingly with language models.
  • Surveillance. Sensors, towers and maritime tracking detect crossings and vessels.

Architecture: models refer, humans decide

A defensible architecture separates three things that are often blurred: the model's output (a score, a match, a summary), the referral (a decision to look more closely), and the decision (entry, refusal, grant), which a named, accountable human makes. Models refer; people decide; every step is written to a record that review and redress can use.

A border risk pipeline: models refer, humans decide, every step leaves a recordAdvance datapassenger and travel dataApplicationsvisa, permit, asylum filesBiometricsface, fingerprintsRisk enginerules + models, versionedBiometric matcher1:1 verify, 1:N searchReferral queuereasons attachedOfficer decisionhuman, accountableDecision recordinputs, versions, reasonsReview and redressappeal, correctionscore + reasonsmatch + scoreloggedNo automatedrefusal pathMonitoring reads the decision records: referral rates and outcomes per group, drift, overrides.
Inputs flow through versioned engines into a referral queue with reasons attached. Only an officer's decision affects the person, and the decision record feeds monitoring and appeals.

Two properties matter most. First, there is no automated refusal path: a high score sends a case to a human with the reasons that produced it, never straight to a denial. Second, the officer sees reasons rather than just a number, because a bare score invites automation bias, the documented tendency to defer to a machine. Recording when officers override the referral, and why, gives you the best signal for whether the model helps.

The base-rate problem

Screening looks for rare events, and rare events make accurate-sounding models produce mostly false alarms. Suppose 1 in 1,000 travellers is a genuine case of interest, and the model catches 90 percent of them while wrongly flagging 2 percent of everyone else. That sounds strong. Run it on 100,000 travellers:

def referral_stats(travellers, prevalence, tpr, fpr):
    """What a screening model's error rates mean at the gate."""
    true_cases = travellers * prevalence
    tp = true_cases * tpr                       # real cases referred
    fp = (travellers - true_cases) * fpr        # innocent travellers referred
    return {"referred": tp + fp, "true_positives": tp,
            "false_positives": fp, "ppv": tp / (tp + fp)}

referral_stats(100_000, 0.001, 0.90, 0.02)
# referred 2,088  true positives 90  false positives 1,998  ppv 4.3%
referral_stats(100_000, 0.001, 0.90, 0.005)
# false positives 500  ppv 15.3%

The model refers 2,088 people, of whom 90 are real cases: a positive predictive value of 4.3 percent. Ninety-six of every hundred people pulled aside did nothing. Cutting the false positive rate to 0.5 percent raises the predictive value only to 15.3 percent. Three consequences follow. Officers must be told what a referral means, so they do not treat it as evidence. Capacity, not accuracy, often sets the threshold, so document that trade-off explicitly. And if the model's false positive rate differs between groups, the burden of those 1,998 wrongful stops falls unevenly: a group with three times the false positive rate is referred three times as often for the same innocence.

Biometric matching: thresholds, demographics and attacks

Biometric matchers output a similarity score; a threshold turns it into match or no match. Two error rates trade against each other: the false match rate (two different people accepted as the same) and the false non-match rate (the same person rejected). Verification at a gate uses 1:1 comparison and fails safe: a non-match sends the traveller to an officer. Identification against a large gallery is riskier, because every search is many comparisons, so even a tiny per-comparison false match rate produces false candidates as the gallery grows.

Error rates are not uniform across people. The US National Institute of Standards and Technology's 2019 demographic study of face recognition (NISTIR 8280) found false positive rates that varied across demographic groups by large factors for many algorithms, with the size of the effect depending heavily on the algorithm. The engineering response is to measure your own system per group on representative data, set thresholds with those numbers in view, and re-measure after every model or camera change.

Two attack classes need defences. Face morphing blends two people's faces into one passport photo that can match both; defences include live capture of the photo at enrolment by the issuing authority and morphing attack detection at issuance and at the border. Presentation attacks, such as printed photos or masks held to a camera, are addressed by presentation attack detection evaluated against ISO/IEC 30107-3. Treat both as requirements with measured error rates, not as checkbox features.

Language models in casework

Language models are entering casework: summarising long asylum files, drafting decision letters, translating interviews and searching country information. Each use carries a specific risk. A summary that invents or drops a detail can change a credibility assessment, and credibility is often decisive in asylum cases. A translation that flattens a hedge or an idiom can make a consistent account look contradictory. Generated country information can be fluent and wrong.

There is also a classic security problem: applicant-submitted documents are untrusted input. A file that contains text addressed to the model, hidden in an attachment or written in a way staff do not read, can try to steer a summary or classification. Treat document content as data, never as instructions; keep the model without tools that change case state; and require every claim in a summary to cite a page and passage in the source so a caseworker can verify it. Use language models to help staff find and read material faster, not to assess credibility or recommend outcomes, and keep qualified interpreters for interviews, with machine translation at most as an aid they check.

Threat model

ThreatExampleControl
Data poisoningManipulated historical decisions or labels bias the risk modelProvenance for training data, label audits, holdout comparisons per release
Probing the rulesRepeated applications to learn which answers avoid referralRate and pattern monitoring; random secondary checks independent of the score
Insider misuseStaff querying biometric or watchlist data without a casePurpose-bound access, per-query logging, regular access review
Function creepA border system reused for unrelated policingLegal basis per purpose, technical separation, documented approvals
Injection via documentsHidden text steering an LLM summaryDocuments as data, cited summaries, no write tools for the model
Biometric spoofingMorphed photos, presentation attacksLive enrolment capture, morphing and presentation attack detection
Data breachExfiltration of biometric templatesEncryption, template protection, minimal retention, segmented storage

Biometric data is a high-value, permanent target: a leaked password can be changed, a face cannot. Minimise what is stored, set retention limits in the system rather than in policy documents, and log every access in a form auditors can query.

The EU AI Act in this domain

In the EU, the AI Act treats this domain as high-risk. Annex III point 7 lists four uses in migration, asylum and border control management, when used by or on behalf of public authorities: polygraphs and similar tools; systems that assess a security, irregular-migration or health risk posed by a person who intends to enter or has entered a Member State; systems that assist the examination of asylum, visa and residence applications and related complaints; and systems that detect, recognise or identify people, with an exception for verifying travel documents. Remote biometric identification and emotion recognition also fall under the biometrics point of Annex III.

Two points are often misstated. The Article 5 prohibition on inferring emotions covers workplaces and education institutions; at the border, emotion recognition is not prohibited but is high-risk, with all the obligations that brings. And for remote biometric identification the Act generally requires that a match be separately verified by at least two people before action, but allows Union or national law to disapply that for border, migration and asylum use where it is considered disproportionate. High-risk systems in this area are registered in a non-public section of the EU database, and public-body deployers must carry out a fundamental rights impact assessment.

On timing, Regulation (EU) 2026/1744, the AI Omnibus, entered into force on 27 July 2026 and moved the obligations for Annex III high-risk systems to 2 December 2027. Check the published text for your exact case; the EU AI Act guide covers roles and the classification procedure in detail. Outside the EU, data protection law, administrative law on reasons and appeals, and human rights obligations apply regardless of AI-specific rules.

Decision records and monitoring

Accountability is an engineering artefact. Every referral should produce a record that lets an auditor reconstruct what the system saw, which versions ran, what it said and what the human did. Store field names rather than personal values where possible, so records can be shared with oversight bodies without new disclosures:

@dataclass(frozen=True)
class ReferralRecord:
    case_id: str
    timestamp_utc: str
    system: str                     # "entry-risk"
    model_version: str              # exact build, and rules version below
    rules_version: str
    input_fields_used: list[str]    # names, not values, so the record is shareable
    score: float
    threshold: float
    reasons: list[str]              # human-readable factors shown to the officer
    officer_id: str | None = None
    officer_decision: str | None = None     # "cleared", "further_checks", ...
    officer_rationale: str | None = None    # required when deviating from the referral
    affected_person_notified: bool = False

def group_referral_rates(records, group_of):
    """Referral rate per group for the monitoring dashboard; group_of uses protected data under strict access."""
    counts = defaultdict(lambda: [0, 0])
    for r in records:
        g = group_of(r.case_id)
        counts[g][1] += 1
        counts[g][0] += r.score >= r.threshold
    return {g: referred / total for g, (referred, total) in counts.items()}

Feed these records into monitoring: referral rates and outcomes per group, override rates per officer and site, and drift in input distributions. The per-group calculation needs protected characteristics, which are themselves sensitive, so compute it in a restricted environment and publish only aggregates. A referral rate that differs sharply between groups with similar outcome rates is the signal to investigate thresholds, features and training data; the fairness guide explains which metrics conflict and why. Make sure affected people learn that a system was involved and how to challenge the decision, and that corrections to wrong data actually propagate to every system that copied it.

Failure modes

  • Score as verdict. Officers treat a referral as evidence of wrongdoing; train them on the base rates and show reasons.
  • Silent threshold changes. Moving a threshold for queue capacity changes who is stopped; version and log thresholds like code.
  • Unmeasured demographic gaps. Vendor accuracy figures from other populations are not your error rates.
  • Stale watchlists and data errors. A wrong record keeps generating referrals until someone corrects it everywhere it was copied.
  • Unverified LLM summaries. An invented detail reaches a decision because nobody checked the citation.
  • No redress path. People cannot contest what they cannot see; notification and appeal are part of the system.

What to do next

  1. Map every AI component in your border or casework pipeline against the four Annex III point 7 uses and the biometrics point, and record the role you play.
  2. Remove any automated refusal path; make every model output a referral with human-readable reasons.
  3. Compute the base-rate numbers for each screening model with your own prevalence estimates and show them to the officers who use it.
  4. Measure biometric error rates per demographic group on representative data, and repeat after every model or hardware change.
  5. Treat applicant documents as untrusted input to any language model and require cited summaries; see human-in-the-loop approval design.
  6. Implement referral records and tamper-evident audit logging, then build per-group monitoring in a restricted environment.
  7. Plan the fundamental rights impact assessment and the 2 December 2027 obligations now, and document notification and appeal routes.
Key takeaway: Border AI acts on people who rarely see it and can rarely contest it, so the design has to supply the accountability. Let models refer and humans decide, show officers reasons and base rates, measure biometric errors per group, treat applicant documents as untrusted input to language models, record every referral with versions and outcomes, and build notification and appeal into the system. In the EU these uses are high-risk under Annex III, with obligations from 2 December 2027.