AI pollution detection fuses several data sources into a judgement: is the air, water or stack emission at this place and time abnormal, and who is likely responsible? The inputs are regulatory-grade reference monitors, dense networks of low-cost sensors, continuous monitors on industrial stacks, satellite retrievals and weather data. The outputs are public alerts, operator notifications, inspection priorities and, increasingly, LLM-written summaries for regulators and the public.
That makes it a security problem, not only a modelling problem, for one structural reason: some of the parties being measured have a financial incentive for the measurement to be wrong. An emitter may own the stack monitor, can sometimes reach a nearby low-cost sensor, and can learn when a satellite passes overhead. This article treats integrity as the core property, walks through the threats at each layer with a tested integrity-check example, and ends with the governance controls that keep a detector from becoming either a rubber stamp or a source of false accusations. It deliberately avoids quoting legal limits, which vary by jurisdiction and pollutant; look those up for your own regulator.
The system and its trust boundaries
Each input class has a different trust level. Reference monitors are calibrated and audited but sparse. Low-cost optical particle sensors are cheap enough to deploy by the hundred but drift and respond to humidity, so they are indicative rather than authoritative. Stack monitors are precise but owned and maintained by the operator whose emissions they measure. Satellite and weather feeds come from third parties and arrive on their own schedules, with gaps under cloud.
Detection models sit downstream: anomaly detectors that flag readings far from what weather and history predict, source-attribution models that combine wind and plume physics to point at likely emitters, and forecasting models that predict exceedances. An LLM assistant may then summarise events and answer questions over the data. The principle that organises the rest of this article: no model output can be more trustworthy than the integrity of the data entering it, so checks belong at the ingest boundary, and every output must carry the provenance of the readings behind it.
Threat model: who wants the measurement wrong
| Actor | Goal | Typical method | Primary control |
|---|---|---|---|
| Regulated emitter | Hide or shift emissions | Emit when sensors are offline or the satellite is not overhead; tamper with an owned monitor | Independent sensors, unannounced checks, coverage-gap alerts |
| Local vandal or neighbour | Disable or discredit a sensor | Cover the inlet, move the device, spoof readings | Peer cross-checks, tamper switches, location attestation |
| Malicious data contributor | Poison a crowdsourced network | Feed fabricated readings or labels | Device keys, contributor reputation, robust aggregation |
| Insider | Suppress or fabricate alerts | Edit thresholds, delete data, change model versions | Signed, append-only logs; two-person changes |
| Activist or competitor | Create false alarms | Place a source near a sensor inlet | Peer agreement, plume plausibility, human review |
| Nobody (nature) | None, but looks like an attack | Calibration drift, humidity, power loss, firmware bugs | Co-location with reference monitors, drift models |
The last row matters most in practice. Benign faults vastly outnumber attacks, and a system that treats every anomaly as an attack drowns its reviewers. The design goal is a pipeline that flags the reading, explains why, and lets the same evidence support "broken sensor", "tampering" or "real event" without prejudging which.
Sensor-layer attacks and their benign twins
Sensor-layer attacks are cheap and physical. Blocking or bagging an inlet pins a particulate reading low. Moving a sensor a few hundred metres upwind of a source changes everything it reports while its metadata says nothing changed. Unplugging it during a planned release creates a gap that a naive dashboard renders as "no data" rather than as suspicious. A sensor's firmware or cloud account can be compromised and made to report a plausible-looking fabricated series.
Timing attacks need no tampering at all. Satellite instruments observe a location at predictable overpass times and cannot see through cloud, so an emitter can schedule releases for when the view is absent. Ground networks with known maintenance windows have the same exposure. The defence is to treat coverage as data: record when each location was observable by which instrument, and flag emission patterns that correlate with gaps.
Benign failures mimic all of these. Low-cost optical particulate sensors are known to over-read at high relative humidity, because particles take up water and grow; drift accumulates over months; a failing fan produces a slow decline that resembles a blocked inlet. This is why co-locating a sample of low-cost sensors with reference monitors and fitting per-device correction models is not optional. Without it, you cannot tell a tampered sensor from a tired one.
Worked example: integrity checks at ingest
The most effective cheap control is to judge each sensor against its neighbours and against itself, before any model consumes its readings. The function below applies range, step, flatline and robust peer-agreement checks to an hourly PM2.5 series. It flags rather than deletes, so reviewers can see what the detector excluded.
import random, statistics
def integrity_flags(series, neighbours, max_step=80.0, flat_n=12, z_lim=4.0, noise_floor=3.0):
"""series: hourly PM2.5 readings (ug/m3) from one sensor.
neighbours: lists of the same length from nearby sensors.
Returns (hour, flag) pairs; downstream models see only unflagged hours."""
flags = []
for t, v in enumerate(series):
if v is None:
flags.append((t, "missing")); continue
if v < 0 or v > 1000:
flags.append((t, "out_of_range")); continue
if t > 0 and series[t - 1] is not None and abs(v - series[t - 1]) > max_step:
flags.append((t, "step"))
if t >= flat_n and len(set(series[t - flat_n:t + 1])) == 1:
flags.append((t, "flatline"))
peers = [n[t] for n in neighbours if n[t] is not None]
if len(peers) >= 3:
med = statistics.median(peers)
mad = max(statistics.median(abs(p - med) for p in peers), noise_floor)
if abs(v - med) / (1.4826 * mad) > z_lim:
flags.append((t, "peer_disagree"))
return flags
rng = random.Random(3)
base = [25 + 10 * ((h % 24) in range(7, 10)) + rng.gauss(0, 3) for h in range(72)]
peers = [[b + rng.gauss(0, 3) for b in base] for _ in range(5)]
target = [b + rng.gauss(0, 3) for b in base]
for h in range(30, 48): # tampering: target pinned at 12.0 during a real event
target[h] = 12.0
for p in peers: # the event, seen by every peer
p[h] += 60
target[60] = 400.0 # a single spike: a source right next to the inletRun on that synthetic 72-hour scenario, the checks produced three kinds of flag. peer_disagree fired on hours 30 to 47, the whole tampering window, plus hour 60. flatline fired only from hour 42, twelve hours into the tampering, because a constant series needs a window before it is suspicious. step fired at hours 60 and 61, the jump to 400 and the fall back. The lesson is that peer agreement is the fast signal and self-consistency checks are slower confirmation.
The noise_floor parameter is there because the first version of this code did not have one. With five peers, the median absolute deviation is sometimes tiny by chance, the robust z-score explodes, and that run flagged six extra clean hours outside both events. Set the floor from the instrument's measured noise at co-location, not from the data being judged. Real deployments also need wind-aware peers: a sensor downwind of a source should disagree with upwind neighbours, so choose the comparison set by meteorology, not only by distance.
Model-layer threats: poisoning, evasion, provenance
Above the sensors, the models face their own attacks. Data poisoning targets any model trained on crowdsourced readings or analyst labels: a contributor who consistently reports clean air near one facility teaches the anomaly model that the area is clean. The general mechanics are covered in the data poisoning guide; the domain-specific defence is to weight training data by device trust, cap any single contributor's influence, and keep a held-out set from reference monitors that the training pipeline never sees, so a poisoned model shows up as a regression on trusted data.
Evasion targets the attribution model. An emitter who knows the model relies on wind direction can release when wind carries the plume away from the sensor network, or blend releases into periods when a neighbouring facility is already emitting. The model is not wrong, it is blind. Report attribution with uncertainty and with the coverage that supported it; an attribution resting on one sensor and a wind estimate is a lead for an inspector, not a finding.
Provenance ties the layers together. Each reading should carry a device identity, a timestamp and a signature from a key held in the device, and each model output should record the model version and the input set it used. Then any alert can be traced back to specific readings from specific devices, and a challenge to that alert can be answered with evidence. The dataset supply chain article covers signed manifests and provenance records in depth; for low-cost devices without secure hardware, at minimum sign at the first gateway you control and record that the device itself was unauthenticated.
LLM assistants on top of detection
LLM assistants are being added on top of these systems to draft incident summaries, answer public questions and help inspectors prioritise. They add three risks.
First, fabricated numbers. An LLM asked "what were PM2.5 levels near the plant last Tuesday" will produce a confident figure whether or not it retrieved one. Every number in an output must come from a tool call against the measurement store, carry its integrity flags, and be rendered by code rather than retyped by the model. Reject drafts containing numbers that do not appear in the tool results.
Second, prompt injection through free text. Complaint forms, inspection notes and operator submissions are untrusted input. An operator's self-report that contains instructions to describe the facility as compliant is an injection attempt aimed at the summariser. Treat these fields as data, quote them rather than obey them, and never give the assistant tools that change thresholds, delete flags or close cases.
Third, tone that outruns the evidence. A summary that says a facility "caused" an event when the attribution model gave a probability with sparse coverage creates legal and reputational harm. Constrain the template: the assistant reports what was measured, what was flagged and the stated uncertainty, and leaves conclusions to the human reviewer. The same pattern appears in the agriculture article, where grounded advisors must not invent field data.
From detection to action
The governance question is how a detection becomes an action. A false negative lets a polluter continue; a false positive accuses someone publicly on bad data. Both errors are expensive, so the pipeline should separate tiers explicitly: indicative alerts from low-cost sensors can trigger public health advice and inspection, while enforcement relies on reference-grade or audited measurements reviewed by a person.
Operationally, that means a few non-negotiable controls. Thresholds, model versions and sensor exclusions change only through reviewed, logged changes, because quietly raising a threshold is the cheapest insider attack. Raw readings are stored append-only, with flags added as annotations rather than edits. Coverage dashboards show where and when the network was blind. And the detector's performance is measured continuously against reference monitors, with drift alarms, so degradation is noticed before someone exploits it.
Failure modes
- Silent gaps. Missing data is rendered as normal, so an offline sensor during a release reads as a clean day.
- Tampering mistaken for drift. A pinned-low sensor is corrected by a drift model instead of flagged, because nobody compared it with peers.
- Drift mistaken for tampering. A humid week triggers tamper alerts across a network of optical sensors and burns reviewer trust.
- Overconfident attribution. A single sensor plus a wind estimate is reported as identifying a source.
- Poisoned retraining. Crowdsourced labels shift the baseline near one facility, and nobody notices because evaluation uses the same crowdsourced data.
- LLM-invented figures. A public summary quotes a concentration that appears in no measurement.
- Editable history. Readings or flags can be overwritten, so no alert survives a challenge.
Trade-offs
| Decision | Option A | Option B | Guidance |
|---|---|---|---|
| Sensor density vs grade | Many low-cost sensors | Few reference monitors | Use both: density for detection, reference for calibration and enforcement |
| Flag vs drop suspect data | Flag and keep | Drop at ingest | Flag; dropping hides tampering evidence |
| Alert sensitivity | Sensitive, more false alarms | Conservative, more misses | Sensitive for inspection triggers, conservative for public accusations |
| Signing | On device | At first trusted gateway | Device when hardware allows; record which applies |
| LLM role | Drafts and answers | Decides and acts | Draft only; numbers from tools; humans decide |
What to do next
- Draw your system's trust boundaries and list, for each input, who owns the device and who benefits if it is wrong.
- Co-locate a sample of low-cost sensors with a reference monitor and fit per-device correction and noise models before trusting their alerts.
- Run the integrity-check code on a week of your own data, tune
noise_floorfrom co-location noise, and review every flag by hand once. - Publish coverage: record when each location was observable by ground and satellite, and alert on emission patterns that align with gaps.
- Make raw readings append-only and put thresholds, exclusions and model versions behind reviewed changes.
- If an LLM drafts reports, enforce that every number comes from a tool result and that free-text inputs are quoted, never followed.