An autonomous vehicle is a robot that runs neural networks at highway speed. Its perception models read the world through cameras, LiDAR and radar. A planner turns that picture into a trajectory, and in a growing number of research and production stacks a vision-language model (VLM) now explains scenes or proposes actions. Every one of those learned components can be fooled, and the attacker does not need to break into the car. A sticker, a projector or a laser aimed at a sensor reaches the model through the same channel as the road itself.

This page treats AI in autonomous vehicles as a security engineering problem: the stack and its attack surface, what published attacks actually demonstrated, why learned components fail this way, and a layered defence built around a runtime plausibility monitor in code. It stays at the level of threat models and defences, not attack recipes.

The stack and its two loops

A typical modular stack has five stages. Sensing produces camera frames, LiDAR point clouds, radar returns and GNSS/IMU readings. Perception runs deep networks that detect objects, segment drivable space and find lane lines. Localization matches sensor data against a high-definition map. Fusion and tracking merges detections across sensors and time into a list of tracked objects with positions and velocities. Prediction and planning forecasts what those objects will do and chooses a trajectory, which control executes. End-to-end stacks collapse several of these stages into one network, which removes inspectable interfaces that a defender could monitor.

Around the online loop sits an offline loop: fleet data is collected, labelled, used to train new models, validated and shipped back to vehicles over the air (OTA). That loop is a software supply chain with all the usual risks, plus one specific to ML: whoever influences the training data influences the model.

A learned driving stack and where attackers can reach itCamerastickers, projection, glareLiDARspoofed / removed pointsRadarjamming, ghost targetsGNSS + IMUposition spoofingPerception DNNsdetect, segment, lanesLocalizationmap match, odometryFusion + trackingplausibility monitorDriving VLMadvisory onlyPrediction + planningsafety envelopeControlsteer, brake, throttleproposalFleet data + labelspoisoning riskTraining + releasemodel supply chainOTA updatesigned (Uptane-style)Physical-world inputs need no network access: anyone who can stand by the road is in the threat model.The offline loop (data, training, OTA) is a classic supply chain and needs classic controls.
The online loop (top) and the offline loop (bottom). Attacks on sensors and scenes enter on the left; data and model attacks enter through the offline loop.

Three standards frame the engineering. ISO 26262 covers functional safety, meaning hazards caused by faults. ISO 21448 (SOTIF, safety of the intended functionality) covers hazards that arise even when nothing is broken, such as a perception model that is simply wrong about an unusual scene. ISO/SAE 21434 covers cybersecurity engineering. UNECE Regulations 155 and 156 make a cybersecurity management system and a software update management system conditions of type approval in the markets that apply them, including the EU. Adversarial attacks on ML touch all three.

Threat model: start from attacker capability

Start with what an attacker can touch, because that determines which defences can help. Each row below is a different capability, not a different technique.

Attacker capabilityEntry pointExample effectPrimary defence layer
Place objects in the sceneCamera, LiDARMisread sign, phantom or hidden obstacleCross-modal plausibility, map priors
Emit signals at a sensorLiDAR, radar, GNSS, camera glareInjected points, ghost targets, position driftSensor integrity checks, redundancy
Put text in view of a VLMDriving VLMTypographic instruction changes the proposalVLM is advisory; output validated
Influence fleet data or labelsTraining pipelineBackdoor or degraded classData provenance, holdout audits
Tamper with models or updatesRelease, OTAMalicious weights on the fleetSigning, Uptane-style OTA, attestation
Send V2X or remote messagesConnectivityFalse hazard or map dataMessage authentication, plausibility

Rank threats by physical realizability and consequence: a printed sign anyone can tape to a pole is cheaper and more repeatable than a laser tracking a moving car.

What the published attacks showed

The published record is research against specific systems and versions: evidence of attack classes, not statements about any current product.

  • Physical adversarial stickers. Eykholt et al. (CVPR 2018) showed that black-and-white stickers placed on a stop sign could make road-sign classifiers read it as a speed-limit sign across many distances and angles, in lab and drive-by tests. The attack targeted classifiers, not a deployed car.
  • Lane perception. Tencent Keen Security Lab's 2019 research on Tesla Autopilot reported that small marks on the road surface could influence lane recognition and steer the car toward the adjacent lane in their tests. Tesla disputed the real-world relevance.
  • LiDAR spoofing. Cao et al. (ACM CCS 2019) injected spoofed points that made Baidu Apollo's LiDAR-based perception report a fake obstacle in front of the vehicle, and showed the spoofed pattern had to be optimised to survive the model's preprocessing.
  • Phantoms. Nassi et al. (ACM CCS 2020) projected or briefly displayed images of road signs and pedestrians and caused driver-assistance systems to react to them.
  • Fusion is not automatically safe. Cao et al. (IEEE S&P 2021) built a single 3D-printed object that both camera and LiDAR perception in a multi-sensor fusion pipeline failed to detect, challenging the assumption that redundancy alone defeats physical attacks.
  • Typographic attacks on driving VLMs. Research in 2024 (arXiv 2405.14169) placed text in traffic scenes and showed it misled the reasoning of open vision-language models such as LLaVA, Qwen-VL, VILA and Imp, and that such attacks transfer between models better than gradient-based perturbations, because models are trained to align text they see with text they read.

Why learned components fail this way

Three properties explain why these attacks work, and each points to a defence.

Learned features are not the features humans use. A classifier maps pixels to classes through a function whose gradient can be large in directions people do not notice. An optimiser that follows that gradient finds small, structured changes that move the output a long way. Physical attacks add one step, expectation over transformation (Athalye et al., 2018): optimise the perturbation averaged over random viewpoints, distances, lighting and print errors so it survives the real world. Defence implication: do not trust a single model's confidence; a high score is exactly what the attacker optimised.

Fusion encodes trust assumptions. A pipeline that accepts an object if any sensor sees it is vulnerable to phantoms; one that requires every sensor is vulnerable to hiding. Defence implication: make the fusion policy explicit, per object class and per situation, and monitor disagreements.

VLMs read instructions from pixels. To a VLM, a sign saying "ignore the red light" is the same kind of token stream as its prompt. This is indirect prompt injection with the road as the document. Defence implication: scene text is untrusted data, and a VLM must never hold actuation authority without a validator that does not read scene text.

A layered defence

No single layer is sufficient, so the architecture stacks independent ones and asks each to fail differently.

  1. Sensor integrity. Check physical consistency: GNSS position against wheel odometry and IMU, LiDAR returns against expected intensity and timing patterns, camera exposure against glare. Some LiDARs randomise pulse timing to make spoofing harder; ask suppliers what theirs do.
  2. Perception robustness. Adversarial training and data augmentation raise the cost of attacks but do not eliminate them. Certified-robustness methods give guarantees only for small perturbation budgets, far below a sticker.
  3. Runtime plausibility monitoring. Cross-check modalities, time and map. An object that appears from nowhere, has no LiDAR support, or contradicts the map gets a lower trust score.
  4. A planner safety envelope. A rule-based layer, for example a formal model of safe following distance such as Mobileye's Responsibility-Sensitive Safety (RSS), bounds what any learned proposal can do. When trust drops, the system degrades: slower speed, more margin, ultimately a minimal-risk manoeuvre.
  5. VLM containment. The VLM proposes from a closed action vocabulary; a validator that reads only structured perception and rules accepts or rejects it.
  6. Supply chain. Track data provenance, audit training sets for anomalous clusters, sign models and use a compromise-resilient OTA design such as Uptane, which separates signing roles so one stolen key cannot push arbitrary updates.

A runtime plausibility monitor in code

The core of layer three is small. The monitor below scores each tracked object on four pieces of independent evidence and maps the fused score to an action. It is a teaching sketch, not a certified component, but every check corresponds to a published attack class.

from dataclasses import dataclass, field

@dataclass
class Track:
    track_id: int
    cls: str                  # "pedestrian", "vehicle", "stop_sign", ...
    range_m: float
    camera_conf: float        # 0..1 from the camera detector
    lidar_points: int         # points inside the projected 3D box
    radar_hit: bool
    age_frames: int           # how long this track has existed
    on_drivable_map: bool     # is it where the HD map says this class can be?
    history: list = field(default_factory=list)

EXPECTED_POINTS = {"pedestrian": 25, "vehicle": 80}   # at 20 m; tune per sensor

def expected_points(cls, range_m):
    base = EXPECTED_POINTS.get(cls)
    if base is None:
        return None
    return base * (20.0 / max(range_m, 1.0)) ** 2   # point density falls ~1/r^2

def plausibility(t: Track) -> tuple[float, list[str]]:
    reasons, score = [], 1.0
    exp = expected_points(t.cls, t.range_m)
    if exp is not None and t.lidar_points < 0.2 * exp and not t.radar_hit:
        score *= 0.4; reasons.append("no range-sensor support")
    if t.age_frames < 3 and t.range_m < 30:
        score *= 0.7; reasons.append("appeared suddenly at close range")
    if not t.on_drivable_map and t.cls in ("vehicle",):
        score *= 0.6; reasons.append("contradicts map")
    if len(t.history) >= 2:
        jump = abs(t.history[-1] - t.history[-2])
        if jump > 3.0:                                  # metres per frame at 10 Hz = 108 km/h
            score *= 0.5; reasons.append("physically implausible jump")
    return score, reasons

def decide(t: Track, ego_speed_mps: float) -> str:
    score, reasons = plausibility(t)
    if score >= 0.6:
        return "accept"
    # Low trust is not "ignore": slow down and gather more evidence.
    if ego_speed_mps > 8 and t.range_m < 40:
        return "degrade: reduce speed, widen margin, alert"
    return "hold: keep tracking, do not plan around it yet"

Note what the code does not do: it never deletes a detection. Low trust changes the response from plan-around-it to slow-down-and-look, because a monitor that silently drops objects converts a phantom attack into a hiding attack. The same principle governs the VLM path:

ALLOWED = {"keep_lane", "slow_down", "stop", "change_lane_left", "change_lane_right"}

def accept_vlm_proposal(proposal: dict, perception, rules) -> str:
    action = proposal.get("action")
    if action not in ALLOWED:
        return "keep_lane"                       # unknown output: ignore, log
    # The validator reads structured perception and traffic rules, never scene text.
    if not rules.permits(action, perception):
        return rules.safest_permitted(perception)
    return action

Worked example: a projected pedestrian at night

Night, urban road, ego vehicle at 50 km/h (about 14 m/s). A projector on a parked van throws the image of a pedestrian onto the road 25 m ahead. The camera detector reports a pedestrian at 0.91 confidence. LiDAR finds 2 points inside the projected box where roughly 16 are expected at 25 m. Radar reports nothing. The track is one frame old, and the map says the location is roadway, which is plausible for a crossing pedestrian, so it adds no penalty.

The monitor multiplies 0.4 for missing range support and 0.7 for sudden close appearance: score 0.28. At 14 m/s and 25 m the decision is to degrade: ease off, widen margin, alert the operator. Over the next 300 ms the track gets more frames. A real pedestrian in dark clothing would gain LiDAR points as range closes; the phantom never does. A frame-accurate projector that moves the image would also trigger the jump check.

Compare the two naive policies. Trust-the-camera brakes hard for a phantom, inviting a rear-end collision, which is the effect the phantom research targeted. Require-LiDAR ignores a real pedestrian whose dark clothing returns few points. The graded response is the trade-off made explicit, and its thresholds are a safety-case decision, not a tuning detail.

Testing the defences

Test the defences the way the attacks were built. In simulation (CARLA is a common open simulator), place patches, phantom objects and spoofed point clusters across a grid of distances, angles, speeds and lighting. Measure attack success with and without the monitor, its time to flag, and its false-alarm rate per thousand clean kilometres. Promote successful cases to hardware-in-the-loop and test tracks. For the VLM, keep a red-team set of scene-text injections and count how many unsafe proposals the validator catches. Feed the results into the SOTIF analysis.

Failure modes

  • Confidence as trust. Using detector confidence as the trust signal rewards exactly what an adversarial optimiser maximises.
  • Monitors that delete. Dropping low-plausibility objects turns every phantom defence into a hiding attack.
  • Correlated redundancy. Two cameras running the same network share the same weakness; diversity must be in modality or model, not just hardware count.
  • VLM with authority. Letting a model that reads scene text emit actuation commands, even "just for edge cases", reopens the injection channel.
  • Untracked data lineage. Without provenance for fleet clips and labels, a poisoning incident cannot be scoped or rolled back.
  • Lab-only evaluation. Robustness measured on digital perturbations says little about printed, weathered, moving artefacts.

Trade-offs

ChoiceGainCost
Strict cross-modal agreementFewer phantom reactionsMisses real objects one sensor handles poorly
Graded degradationSafe response under uncertaintyComfort and throughput loss; more disengagements
Adversarial trainingRaises attack costTraining compute; some clean accuracy loss
VLM advisory onlyCloses text-injection path to actuatorsLess benefit from VLM reasoning

What to do next

  1. Draw your stack as in the diagram and run STRIDE over every interface, including the offline data loop.
  2. List which sensors must support each object class before the planner commits to it, and write it down as policy rather than leaving it implicit in fusion code.
  3. Implement a plausibility monitor that grades instead of deletes, and log its scores on every clean drive to set a false-alarm baseline.
  4. Build a simulation suite of patch, phantom and spoofing scenarios, and track attack success rate per release next to accuracy.
  5. If a VLM is in the loop, give it a closed action vocabulary and a validator that never reads scene text.
  6. Audit model signing and OTA key separation against your cybersecurity management system obligations.
  7. Keep learning: STRIDE threat modelling for AI systems, universal adversarial attacks, an AI supply-chain security programme and running an AI red-team programme.
Key takeaway: An autonomous vehicle's AI can be attacked through the road itself: stickers, projections, spoofed sensor signals and, for driving VLMs, text in the scene. Defend in independent layers: sensor integrity, cross-modal plausibility monitoring that grades trust instead of deleting objects, a rule-based safety envelope, a VLM with no actuation authority, and a signed, provenance-tracked data and update pipeline. Measure attack success per release, not only accuracy.