Aviation is adopting machine learning in places that range from harmless to safety-critical: predictive maintenance on engine data, assistants that search maintenance manuals, decision support for air traffic flow, camera-based runway and traffic detection, and detect-and-avoid for uncrewed aircraft. The industry already has some of the most mature safety processes in engineering. What it does not yet have is settled practice for systems whose behaviour is learned from data, and whose inputs an adversary may control.

This article looks at AI in aviation from a security engineer's seat. It covers the assurance frameworks a model must fit into, where attackers can reach an ML function, how runtime monitoring keeps a learned component honest, why ground-side LLM tools are the nearest-term risk, and what an organisation should do now. It is not legal or certification advice. Standards in this area are still moving, and where a document is in draft the text says so.

The assurance landscape

Conventional airborne software is assured under DO-178C (ED-12C in Europe), whose core idea is traceability: every requirement traces to code and tests, and structural coverage shows no untested logic. System development follows ARP4754 and safety assessment follows ARP4761. Airworthiness security, meaning protection against intentional unauthorised electronic interaction, has its own process standard, DO-326A (ED-202A), with methods in companion documents. Neural networks strain the first of these badly. Their behaviour lives in millions of weights trained from data, not in requirements written line by line, so statement coverage of the inference code says almost nothing about correctness.

Regulators have responded with learning-specific guidance. EASA's AI Roadmap classifies applications by level: Level 1 assists a human, Level 2 is human-AI teaming in which the system may take decisions under human oversight, and Level 3 is more autonomous. Its AI Concept Paper Issue 02, published in March 2024, gives guidance for Level 1 and 2 machine learning applications built around learning assurance, explainability and an ethics-based assessment. Learning assurance treats the data as a development artefact: requirements on data completeness and representativeness, independence of training and test sets, and verification of the trained model against its operational design domain. A joint EUROCAE WG-114 and SAE G-34 standard, ED-324 / ARP6983, is being developed to turn this into a certification process standard; at the time of writing treat it as in development and check its status before citing it. The FAA has also published a roadmap for AI safety assurance. None of this replaces DO-326A: a learned component still needs a security risk assessment.

The attack surface

Start a threat model by drawing where data flows into the model, both during development and in service. Four surfaces matter.

Training data. Labelled imagery, simulator output, flight-data recordings and maintenance histories. Poisoning a small fraction can implant a backdoor, for example a detector that misses aircraft when a particular pattern is present, or skew a maintenance model toward deferring a class of findings. Aviation data often passes through many contractors, which widens this surface; the general mechanics are in the data poisoning article.

Model supply chain. Pretrained backbones, training frameworks, compilers that turn a network into code for an embedded accelerator, and the deployment package itself. A swapped weight file or a compromised conversion step yields a model that passed tests and is not the model that flies. See running an AI supply chain programme.

Sensor inputs in service. A camera can be shown an adversarial pattern; laser illumination of cockpits is already a real hazard. Radio inputs are worse. ADS-B broadcasts are neither authenticated nor encrypted, so fake traffic can be injected by anyone with a transmitter, and GNSS jamming and spoofing have become routine in some regions, prompting safety bulletins from EASA and others. An ML function that consumes these signals inherits their weaknesses, and may hide them behind a confident output.

Ground-side text. LLM assistants read manuals, service bulletins, work orders, pilot reports and email. Any of those documents can carry instructions aimed at the model rather than the human.

Where an attacker can touch an aviation ML functiontraining datalabels, sim, flight logsmodel supply chainweights, toolchainsensor inputscamera, GNSS, ADS-BML componentfrozen, signed, in an ODDpoisoningtamperingspoofing, patchesruntime monitorODD + cross-checkscrew / operatordecides, can overrideground LLM toolsMRO, ops, dispatchdocuments, ticketsprompt injectionRed: attacker-reachable inputs. The monitor and the human are the last line, so they must not share the model's blind spots.
Attack surface of an aviation ML function. The monitor and human are the final barrier only if they are independent of the model's inputs.

Runtime assurance: a worked example

The most useful architectural pattern is runtime assurance, often called a simplex architecture. The learned component is allowed to act or advise only while a small, conventionally verified monitor confirms that its inputs are inside the operational design domain (ODD) and that its output agrees with independent evidence. When the monitor objects, the system falls back to a verified path or tells the crew the function is unavailable. ASTM F3269 describes run-time assurance for aircraft systems that contain complex functions, and it maps naturally onto DO-178C because the part that needs the highest assurance is the simple arbiter, not the network.

Worked example: a camera-based runway detector gives the crew a visual cue on approach. Its ODD says daylight, visibility above a set value, a runway in the onboard database, and an approach angle within a band. The independent reference is the expected runway position computed from navigation sensors and the database. The arbiter below enables the cue only when everything agrees, and it cross-checks GNSS against inertial data so that a spoofed position cannot validate a fooled detector.

Runtime assurance: the ML path advises only while independent checks agreecamera framerunway in viewML runway detectorpose estimate + scorereference pathGNSS/ILS + runway databasenav sensorsposition, attitudearbiterODD, OOD, agreementadvisory onML cue shownadvisory offflagged, loggedThe arbiter is small, deterministic and conventionally verified; it never needs to understand the network.Every disagreement is logged with the input, which is the field evidence for drift or attack.
Runtime assurance around a learned runway detector. The arbiter is the certified part; the network only advises.
from dataclasses import dataclass


@dataclass
class Frame:
    visibility_m: float
    sun_elevation_deg: float
    runway_in_db: bool
    glide_angle_deg: float


@dataclass
class Detection:
    lateral_offset_m: float   # detector's estimate of offset from centreline
    score: float              # detector confidence, not a probability
    ood_score: float          # distance from training distribution, from a separate model


@dataclass
class Reference:
    lateral_offset_m: float   # from navigation solution + runway database
    gnss_ins_divergence_m: float


ODD = dict(min_vis=1500.0, min_sun=5.0, glide=(2.0, 4.0))
LIMITS = dict(max_disagree_m=8.0, max_ood=0.7, min_score=0.6, max_gnss_ins=30.0)


def arbitrate(f: Frame, d: Detection, r: Reference):
    reasons = []
    if f.visibility_m < ODD["min_vis"] or f.sun_elevation_deg < ODD["min_sun"]:
        reasons.append("outside ODD: conditions")
    if not f.runway_in_db or not (ODD["glide"][0] <= f.glide_angle_deg <= ODD["glide"][1]):
        reasons.append("outside ODD: geometry")
    if d.ood_score > LIMITS["max_ood"] or d.score < LIMITS["min_score"]:
        reasons.append("input unfamiliar or low score")
    if r.gnss_ins_divergence_m > LIMITS["max_gnss_ins"]:
        reasons.append("navigation reference untrusted")   # possible spoofing
    elif abs(d.lateral_offset_m - r.lateral_offset_m) > LIMITS["max_disagree_m"]:
        reasons.append("detector disagrees with reference")
    return (len(reasons) == 0), reasons   # caller logs reasons with the frame

Note what the arbiter does not do: it does not trust the detector's confidence on its own, because adversarial inputs typically produce confident wrong answers. It uses an independent out-of-distribution signal, and it refuses to use navigation data as a reference when GNSS and inertial data disagree. The thresholds are illustrative; real ones come from the safety assessment and flight-test data.

Lifecycle controls

Security controls across the lifecycle follow from the surfaces.

  • Data provenance. Record source, collection conditions, labeller and transformations for every training item, and hash datasets so a release names the exact data it learned from. Keep a held-out test set controlled by a different team from the training data.
  • Poisoning checks. Audit label agreement, look for clusters of training items that dominate a class's decision boundary, and test for triggers by stamping candidate patterns onto clean validation images and watching for flipped outputs.
  • Frozen models. Airborne models should not learn in service. Every change is a new configuration item with its own verification, which also removes online poisoning.
  • Signed artefacts. Sign weights and the compiled inference binary, verify signatures at load, and reproduce the build so the flying artefact provably comes from the reviewed data and code.
  • Robustness evidence. Test against weather, lighting, sensor degradation and adversarial perturbations inside the ODD. Formal verification of small networks over bounded input regions is maturing and fits learning assurance well, but it does not yet scale to large vision backbones.
  • Field monitoring. Log every monitor rejection with its input and analyse them centrally. A cluster of rejections at one airport is drift, a new building, or an attack, and you cannot tell which without the data.

Ground-side LLM tools

The nearest-term exposure is not in the cockpit. It is in maintenance, repair and overhaul (MRO), operations control and dispatch, where LLM assistants are being deployed to search manuals, draft work orders and summarise defects. These tools are not airborne software, but their output can lead a technician to a wrong torque value or a skipped inspection. Four rules make them defensible.

  1. Approved data is authoritative, the model is not. Every answer must cite a specific task in the approved maintenance data at a named revision, and the technician works from that document. The assistant is a search aid, never the source.
  2. Retrieve only from controlled sources. Index approved manuals and bulletins from the document control system. Free text from work orders, email and pilot reports is untrusted input and may carry prompt injection; mark it as such and never let it change instructions.
  3. No write actions without a human. Drafting a work order is fine; signing one off, deferring a defect or changing an inventory record requires a person.
  4. Verify citations mechanically. Check that each cited task number exists in the cited revision and that numeric values in the answer appear in the cited text.
import re


def check_answer(answer, citations, approved_index):
    """approved_index maps (doc_id, revision, task) -> approved task text."""
    problems = []
    if not citations:
        problems.append("no citation to approved data")
    cited_text = ""
    for cit in citations:
        key = (cit["doc"], cit["rev"], cit["task"])
        if key not in approved_index:
            problems.append(f"cited task not in approved revision: {key}")
        else:
            cited_text += approved_index[key]
    for value in re.findall(r"[0-9]+(?:[.][0-9]+)?", answer):
        if value not in cited_text:
            problems.append(f"number {value} not found in cited text")
    return problems   # any problem: show sources only, suppress the summary

The general patterns for isolating untrusted text from instructions are in the prompt isolation article and apply directly here.

Failure modes

  • Trusting confidence. A fooled detector is usually confident; gating on its own score gives no protection against adversarial inputs.
  • Monitor shares the input. If the reference path consumes the same spoofed GNSS as the model, the cross-check validates the attack.
  • ODD written, not enforced. The design domain lives in a document and no code checks it at runtime.
  • Unreproducible model. Nobody can rebuild the flying weights from recorded data and code, so a poisoning investigation has nothing to compare against.
  • Assistant treated as authority. A technician copies a value from a chat answer instead of the manual revision it claims to quote.
  • Automation bias. Crews and controllers stop checking an aid that is usually right. Human-factors evaluation is part of security, not separate from it.

Trade-offs

ChoiceGainCost
Advisory only (Level 1)Human check on every outputAutomation bias, workload
Runtime assurance monitorCertifiable arbiter, contains failuresLower availability of the ML function
Narrow ODDSmaller test space, stronger evidenceFunction unavailable more often
Frozen modelNo online poisoning, stable evidenceSlow to adapt to new conditions
Ground LLM with citation checksFast search with an audit trailRefuses more, needs document control integration

What to do next

  1. Inventory every ML and LLM use in your organisation, airborne and ground, and assign each an EASA-style level and a security owner.
  2. For each, draw the data flows and run a DO-326A-style security risk assessment that includes poisoning, supply chain, spoofed inputs and prompt injection.
  3. Write each ODD as executable checks, and put an independent runtime monitor around any learned function that influences flight.
  4. Make training data and model builds reproducible and signed, with provenance per item.
  5. Restrict ground assistants to approved, revision-controlled sources and add mechanical citation checks.
  6. Track the EASA concept paper and ED-324 / ARP6983 as they mature. For a regulated comparison, read AI medical device security.
Key takeaway: A learned component in aviation is only as trustworthy as its data, its build and its inputs, and attackers can reach all three. Fit it into DO-178C, DO-326A and EASA learning assurance, enforce its design domain at runtime with an independent monitor, keep models frozen, signed and reproducible, and make ground LLM tools cite approved data rather than replace it.