AI pollution detection fuses several data sources into a judgement: is the air, water or stack emission at this place and time abnormal, and who is likely responsible? The inputs are regulatory-grade reference monitors, dense networks of low-cost sensors, continuous monitors on industrial stacks, satellite retrievals and weather data. The outputs are public alerts, operator notifications, inspection priorities and, increasingly, LLM-written summaries for regulators and the public.

That makes it a security problem, not only a modelling problem, for one structural reason: some of the parties being measured have a financial incentive for the measurement to be wrong. An emitter may own the stack monitor, can sometimes reach a nearby low-cost sensor, and can learn when a satellite passes overhead. This article treats integrity as the core property, walks through the threats at each layer with a tested integrity-check example, and ends with the governance controls that keep a detector from becoming either a rubber stamp or a source of false accusations. It deliberately avoids quoting legal limits, which vary by jurisdiction and pollutant; look those up for your own regulator.

The system and its trust boundaries

An AI pollution detection system and its trust boundariesReference monitorsregulatory gradeLow-cost sensorsdense, indicativeStack monitorsoperator ownedSatellite, weatherthird party feedsIngest + integritysignatures, QC flagsDetection modelsanomaly, attributionAlertspublic, operatorsLLM assistantreports, Q and AHuman reviewbefore enforcementDashed trust boundary: every arrow into ingestcrosses from devices or parties you do not control.
Four classes of input cross a trust boundary into ingest. Integrity checks sit at that boundary, before any model sees the data.

Each input class has a different trust level. Reference monitors are calibrated and audited but sparse. Low-cost optical particle sensors are cheap enough to deploy by the hundred but drift and respond to humidity, so they are indicative rather than authoritative. Stack monitors are precise but owned and maintained by the operator whose emissions they measure. Satellite and weather feeds come from third parties and arrive on their own schedules, with gaps under cloud.

Detection models sit downstream: anomaly detectors that flag readings far from what weather and history predict, source-attribution models that combine wind and plume physics to point at likely emitters, and forecasting models that predict exceedances. An LLM assistant may then summarise events and answer questions over the data. The principle that organises the rest of this article: no model output can be more trustworthy than the integrity of the data entering it, so checks belong at the ingest boundary, and every output must carry the provenance of the readings behind it.

Threat model: who wants the measurement wrong

ActorGoalTypical methodPrimary control
Regulated emitterHide or shift emissionsEmit when sensors are offline or the satellite is not overhead; tamper with an owned monitorIndependent sensors, unannounced checks, coverage-gap alerts
Local vandal or neighbourDisable or discredit a sensorCover the inlet, move the device, spoof readingsPeer cross-checks, tamper switches, location attestation
Malicious data contributorPoison a crowdsourced networkFeed fabricated readings or labelsDevice keys, contributor reputation, robust aggregation
InsiderSuppress or fabricate alertsEdit thresholds, delete data, change model versionsSigned, append-only logs; two-person changes
Activist or competitorCreate false alarmsPlace a source near a sensor inletPeer agreement, plume plausibility, human review
Nobody (nature)None, but looks like an attackCalibration drift, humidity, power loss, firmware bugsCo-location with reference monitors, drift models

The last row matters most in practice. Benign faults vastly outnumber attacks, and a system that treats every anomaly as an attack drowns its reviewers. The design goal is a pipeline that flags the reading, explains why, and lets the same evidence support "broken sensor", "tampering" or "real event" without prejudging which.

Sensor-layer attacks and their benign twins

Sensor-layer attacks are cheap and physical. Blocking or bagging an inlet pins a particulate reading low. Moving a sensor a few hundred metres upwind of a source changes everything it reports while its metadata says nothing changed. Unplugging it during a planned release creates a gap that a naive dashboard renders as "no data" rather than as suspicious. A sensor's firmware or cloud account can be compromised and made to report a plausible-looking fabricated series.

Timing attacks need no tampering at all. Satellite instruments observe a location at predictable overpass times and cannot see through cloud, so an emitter can schedule releases for when the view is absent. Ground networks with known maintenance windows have the same exposure. The defence is to treat coverage as data: record when each location was observable by which instrument, and flag emission patterns that correlate with gaps.

Benign failures mimic all of these. Low-cost optical particulate sensors are known to over-read at high relative humidity, because particles take up water and grow; drift accumulates over months; a failing fan produces a slow decline that resembles a blocked inlet. This is why co-locating a sample of low-cost sensors with reference monitors and fitting per-device correction models is not optional. Without it, you cannot tell a tampered sensor from a tired one.

Worked example: integrity checks at ingest

The most effective cheap control is to judge each sensor against its neighbours and against itself, before any model consumes its readings. The function below applies range, step, flatline and robust peer-agreement checks to an hourly PM2.5 series. It flags rather than deletes, so reviewers can see what the detector excluded.

import random, statistics

def integrity_flags(series, neighbours, max_step=80.0, flat_n=12, z_lim=4.0, noise_floor=3.0):
    """series: hourly PM2.5 readings (ug/m3) from one sensor.
    neighbours: lists of the same length from nearby sensors.
    Returns (hour, flag) pairs; downstream models see only unflagged hours."""
    flags = []
    for t, v in enumerate(series):
        if v is None:
            flags.append((t, "missing")); continue
        if v < 0 or v > 1000:
            flags.append((t, "out_of_range")); continue
        if t > 0 and series[t - 1] is not None and abs(v - series[t - 1]) > max_step:
            flags.append((t, "step"))
        if t >= flat_n and len(set(series[t - flat_n:t + 1])) == 1:
            flags.append((t, "flatline"))
        peers = [n[t] for n in neighbours if n[t] is not None]
        if len(peers) >= 3:
            med = statistics.median(peers)
            mad = max(statistics.median(abs(p - med) for p in peers), noise_floor)
            if abs(v - med) / (1.4826 * mad) > z_lim:
                flags.append((t, "peer_disagree"))
    return flags

rng = random.Random(3)
base = [25 + 10 * ((h % 24) in range(7, 10)) + rng.gauss(0, 3) for h in range(72)]
peers = [[b + rng.gauss(0, 3) for b in base] for _ in range(5)]
target = [b + rng.gauss(0, 3) for b in base]
for h in range(30, 48):          # tampering: target pinned at 12.0 during a real event
    target[h] = 12.0
    for p in peers:              # the event, seen by every peer
        p[h] += 60
target[60] = 400.0               # a single spike: a source right next to the inlet

Run on that synthetic 72-hour scenario, the checks produced three kinds of flag. peer_disagree fired on hours 30 to 47, the whole tampering window, plus hour 60. flatline fired only from hour 42, twelve hours into the tampering, because a constant series needs a window before it is suspicious. step fired at hours 60 and 61, the jump to 400 and the fall back. The lesson is that peer agreement is the fast signal and self-consistency checks are slower confirmation.

The noise_floor parameter is there because the first version of this code did not have one. With five peers, the median absolute deviation is sometimes tiny by chance, the robust z-score explodes, and that run flagged six extra clean hours outside both events. Set the floor from the instrument's measured noise at co-location, not from the data being judged. Real deployments also need wind-aware peers: a sensor downwind of a source should disagree with upwind neighbours, so choose the comparison set by meteorology, not only by distance.

Model-layer threats: poisoning, evasion, provenance

Above the sensors, the models face their own attacks. Data poisoning targets any model trained on crowdsourced readings or analyst labels: a contributor who consistently reports clean air near one facility teaches the anomaly model that the area is clean. The general mechanics are covered in the data poisoning guide; the domain-specific defence is to weight training data by device trust, cap any single contributor's influence, and keep a held-out set from reference monitors that the training pipeline never sees, so a poisoned model shows up as a regression on trusted data.

Evasion targets the attribution model. An emitter who knows the model relies on wind direction can release when wind carries the plume away from the sensor network, or blend releases into periods when a neighbouring facility is already emitting. The model is not wrong, it is blind. Report attribution with uncertainty and with the coverage that supported it; an attribution resting on one sensor and a wind estimate is a lead for an inspector, not a finding.

Provenance ties the layers together. Each reading should carry a device identity, a timestamp and a signature from a key held in the device, and each model output should record the model version and the input set it used. Then any alert can be traced back to specific readings from specific devices, and a challenge to that alert can be answered with evidence. The dataset supply chain article covers signed manifests and provenance records in depth; for low-cost devices without secure hardware, at minimum sign at the first gateway you control and record that the device itself was unauthenticated.

LLM assistants on top of detection

LLM assistants are being added on top of these systems to draft incident summaries, answer public questions and help inspectors prioritise. They add three risks.

First, fabricated numbers. An LLM asked "what were PM2.5 levels near the plant last Tuesday" will produce a confident figure whether or not it retrieved one. Every number in an output must come from a tool call against the measurement store, carry its integrity flags, and be rendered by code rather than retyped by the model. Reject drafts containing numbers that do not appear in the tool results.

Second, prompt injection through free text. Complaint forms, inspection notes and operator submissions are untrusted input. An operator's self-report that contains instructions to describe the facility as compliant is an injection attempt aimed at the summariser. Treat these fields as data, quote them rather than obey them, and never give the assistant tools that change thresholds, delete flags or close cases.

Third, tone that outruns the evidence. A summary that says a facility "caused" an event when the attribution model gave a probability with sparse coverage creates legal and reputational harm. Constrain the template: the assistant reports what was measured, what was flagged and the stated uncertainty, and leaves conclusions to the human reviewer. The same pattern appears in the agriculture article, where grounded advisors must not invent field data.

From detection to action

The governance question is how a detection becomes an action. A false negative lets a polluter continue; a false positive accuses someone publicly on bad data. Both errors are expensive, so the pipeline should separate tiers explicitly: indicative alerts from low-cost sensors can trigger public health advice and inspection, while enforcement relies on reference-grade or audited measurements reviewed by a person.

Operationally, that means a few non-negotiable controls. Thresholds, model versions and sensor exclusions change only through reviewed, logged changes, because quietly raising a threshold is the cheapest insider attack. Raw readings are stored append-only, with flags added as annotations rather than edits. Coverage dashboards show where and when the network was blind. And the detector's performance is measured continuously against reference monitors, with drift alarms, so degradation is noticed before someone exploits it.

Failure modes

  • Silent gaps. Missing data is rendered as normal, so an offline sensor during a release reads as a clean day.
  • Tampering mistaken for drift. A pinned-low sensor is corrected by a drift model instead of flagged, because nobody compared it with peers.
  • Drift mistaken for tampering. A humid week triggers tamper alerts across a network of optical sensors and burns reviewer trust.
  • Overconfident attribution. A single sensor plus a wind estimate is reported as identifying a source.
  • Poisoned retraining. Crowdsourced labels shift the baseline near one facility, and nobody notices because evaluation uses the same crowdsourced data.
  • LLM-invented figures. A public summary quotes a concentration that appears in no measurement.
  • Editable history. Readings or flags can be overwritten, so no alert survives a challenge.

Trade-offs

DecisionOption AOption BGuidance
Sensor density vs gradeMany low-cost sensorsFew reference monitorsUse both: density for detection, reference for calibration and enforcement
Flag vs drop suspect dataFlag and keepDrop at ingestFlag; dropping hides tampering evidence
Alert sensitivitySensitive, more false alarmsConservative, more missesSensitive for inspection triggers, conservative for public accusations
SigningOn deviceAt first trusted gatewayDevice when hardware allows; record which applies
LLM roleDrafts and answersDecides and actsDraft only; numbers from tools; humans decide

What to do next

  1. Draw your system's trust boundaries and list, for each input, who owns the device and who benefits if it is wrong.
  2. Co-locate a sample of low-cost sensors with a reference monitor and fit per-device correction and noise models before trusting their alerts.
  3. Run the integrity-check code on a week of your own data, tune noise_floor from co-location noise, and review every flag by hand once.
  4. Publish coverage: record when each location was observable by ground and satellite, and alert on emission patterns that align with gaps.
  5. Make raw readings append-only and put thresholds, exclusions and model versions behind reviewed changes.
  6. If an LLM drafts reports, enforce that every number comes from a tool result and that free-text inputs are quoted, never followed.
Key takeaway: Pollution detection is measured by parties who may want it wrong, so integrity comes before accuracy. Check every reading against peers and itself at ingest, flag rather than delete, calibrate low-cost sensors against reference monitors so drift and tampering can be told apart, sign readings and record model provenance, keep LLMs to drafting with numbers from tools, and require human review on reference-grade evidence before any enforcement.