Forestry has quietly become one of the most AI-dependent compliance domains. Buyers of timber, pulp and wood products must show that their supply did not come from recently cleared land; carbon projects are paid according to how much forest they say they kept standing; fire agencies triage satellite and camera alerts faster than any human team could. In each case a model turns pixels and paperwork into a decision with money attached, and in each case some party gains from a wrong answer.

This article treats forestry AI as a security problem: it maps the evidence chain from supplier documents and satellite scenes to a decision, names who attacks each link, and shows concrete guards, ending with one shipment traced through the pipeline. The companion pieces on environmental monitoring and agriculture cover sensor tampering and field machinery; here the focus is remote sensing, geolocation and document trust.

Where AI sits in forestry and who wants it wrong

Four AI workloads dominate. Deforestation-free due diligence checks that each plot a commodity came from was not cleared or degraded after a cut-off date. The EU Deforestation Regulation (EUDR) made this concrete for wood, among other commodities: operators must file geolocation for every plot, the cut-off date is 31 December 2020, and plots larger than four hectares must be described by a polygon rather than a single point. The application date has moved twice; after Regulation (EU) 2025/2650 it is 30 December 2026 for large operators and 30 June 2027 for micro and small ones. Treat those details as things to re-check against the consolidated text before you build on them.

Carbon measurement, reporting and verification (MRV) estimates biomass from satellite, lidar and field plots and credits carbon against a baseline. Fire detection fuses satellite hotspots with tower cameras. And document assistants, increasingly LLM-based, read permits, licences and supplier declarations at scale.

The common thread is that the party supplying the input is often the party the decision is about. A supplier wants a clean verdict; a project developer wants a high baseline; an illegal logger wants no alert. That makes forestry closer to fraud detection than to ordinary analytics.

The evidence chain and its threat model

The design rule that matters is the split between components that interpret and the one that decides. Models and the LLM produce evidence with stated confidence; a versioned, deterministic rule engine turns it into release, hold or reject, and anything uncertain goes to a human.

Forestry AI evidence chain: every arrow crosses a trust boundarySupplier inputsplots, permits, volumesSatellite archiveoptical + radar scenesField and crowd labelsplots, annotationsGeometry gatevalidate, dedupe, areaChange detectionloss since 2020-12-31Biomass modelcarbon per hectareLLM extractorfields + cited spansRule enginedeterministic decisionHuman reviewholds, sign-offdocsfieldsholdRed: parties who benefit from a wrong answer. The LLM may extract and summarise; only the rule engine decides.
The forestry evidence chain. Inputs from interested parties (red) pass through gates that validate them before any model sees them, and only the rule engine issues a decision.
AssetAttackerGoalPrimary control
Plot geolocationSupplier, traderLaunder wood from cleared landGeometry gate and duplicate detection
Change-detection outputIllegal logger, supplierClearing goes unseenMulti-sensor fusion, abstention
Biomass and baselineProject developerOver-credit carbonUncertainty deductions, independent plots
Training labelsCrowd contributor, insiderBias the modelLabel provenance, audit sampling
Supplier documentsSupplierSteer the LLM verdictExtraction only, cited spans, rules decide

Geolocation laundering and the geometry gate

Geolocation is the cheapest place to cheat. Dilution: a polygon much larger than the harvest area makes a cleared corner a small fraction of it. Borrowing: wood from a cleared site is declared against a neighbour's intact plot, sometimes one polygon reused by several suppliers. Imprecision: truncated coordinates, or a point where a polygon is required, which makes any satellite check meaningless.

Geolocation laundering: the declared polygon versus where the wood came fromdeclared plot: intact forestcleared 2023actual harvest: cleared 2022Attack 1: a large polygon dilutes a cleared cornerAttack 2: a clean neighbour's polygon is reusedDefences: share-of-area thresholds, yield-per-hectare plausibility, cross-supplier duplicate detection
Two laundering patterns. A plot-level verdict must look at the share and location of loss inside the polygon, and at whether the same geometry appears under other suppliers.

A geometry gate runs before any imagery is fetched. It rejects malformed shapes, checks precision, compares declared volume with what that area could plausibly yield, and hashes normalised geometry so the same plot under two suppliers is flagged. The sketch below uses shapely and pyproj for area in square metres.

from shapely.geometry import shape
from shapely.validation import explain_validity
from pyproj import Geod
import hashlib

GEOD = Geod(ellps="WGS84")
MAX_YIELD_M3_PER_HA = {"temperate": 400, "tropical": 250}   # set from your own data

def gate_plot(feature, declared_ha, declared_m3, biome, seen_hashes):
    issues = []
    geom = shape(feature["geometry"])
    if not geom.is_valid:
        issues.append("invalid geometry: " + explain_validity(geom))
    if geom.geom_type == "Point":
        area_ha = declared_ha
        if declared_ha > 4:
            issues.append("point given for a plot over 4 ha")
    else:
        area_ha = abs(GEOD.geometry_area_perimeter(geom)[0]) / 10_000
    if area_ha and declared_m3 / area_ha > MAX_YIELD_M3_PER_HA[biome]:
        issues.append(f"implausible yield {declared_m3/area_ha:.0f} m3/ha")
    h = hashlib.sha256(geom.normalize().wkb).hexdigest()
    if h in seen_hashes:
        issues.append("geometry already declared by supplier " + seen_hashes[h])
    return area_ha, h, issues

Check coordinate precision on the raw GeoJSON text, since parsing to floats hides truncation. Exact-hash matching catches lazy reuse but not a polygon nudged by a few metres, so also flag declarations from different suppliers whose polygons overlap heavily. The yield ceiling is deliberately crude: it catches a two-hectare plot backing two thousand cubic metres. Check points against a buffer too; a point in intact forest beside a clearing is the borrowing attack in miniature.

Change detection that can say not observed

Change detection compares the plot's condition now with its condition at the cut-off date. Public alert products show the pattern: optical alerts such as the University of Maryland's GLAD system and radar alerts such as Wageningen's RADD, built on Sentinel-1, flag tree cover loss at roughly ten to thirty metre resolution. Each has blind spots, and attackers learn them.

Cloud is the obvious one: optical sensors can be blind for weeks in the tropics, so clearing timed to the wet season may be missed. Radar sees through cloud but is noisy over steep terrain. Selective logging of individual trees often falls below loss-product resolution and shows only as degradation, which the regulation treats separately for wood. And any model trained on labels can be steered through them: if crowd or contractor labels feed retraining, a contributor who marks cleared patches as intact moves the boundary. That is the classic data poisoning attack with a geographic twist, because the poison can be targeted at one concession.

The defence is to make absence of evidence explicit. Each plot gets one of three states per period: observed clear, observed loss, or not observed. A plot with observation gaps longer than a set window is unverified, and the rule engine holds it. Fuse optical and radar to fill gaps, version the model and scene list behind every verdict, and keep a field-verified holdout, chosen by you, that every retrained model must still pass.

Carbon estimates and the measured party

Carbon projects that claim avoided deforestation are paid for the difference between a baseline, the loss that would have happened without the project, and the loss that was observed. Both terms are model outputs, and the developer chooses much of the baseline method. A 2023 study in Science by West and colleagues compared many such projects with synthetic controls and concluded that most had overstated their baselines. Whatever the exact share, the incentive is clear: the measured party benefits from a pessimistic counterfactual and an optimistic biomass map.

Biomass maps regress carbon from spectral, radar and lidar inputs, with large plot-level errors that grow in dense forest where optical signals saturate. Baseline models predict loss from roads, slope and past clearing, and their skill depends on a reference region someone chose. Guard both the same way:

  • Credit the conservative bound of the uncertainty interval, not the mean.
  • Fix and hash the reference region and model version before the monitoring period, so a baseline cannot be re-fitted after the outcome is known.
  • Validate against field plots measured by a party with no stake in the result.
  • Re-run every calculation from stored inputs; a credit that cannot be reproduced is not issued.

An LLM that reads supplier documents

Due-diligence teams now use LLMs to read the documents suppliers send: harvest permits, concession maps, transport papers, declarations of legality. This is where the newest attack surface sits, because the documents are written by the party being assessed. A PDF can carry hidden text such as an instruction to report all plots as verified, or simply phrased claims designed to be summarised as fact. This is indirect prompt injection with a commercial motive.

The defence is architectural. The LLM extracts typed fields, each with the exact span of source text it came from, and nothing else. It never emits a verdict, and its output is checked: a permit number must match the issuing registry's format, a date must parse, a cited span must actually appear in the document text. The decision is computed by rules over extracted fields and the geometry and change-detection results.

from dataclasses import dataclass

@dataclass
class Field:
    name: str
    value: str
    span: str          # verbatim text from the document

def verify_fields(fields, doc_text, validators):
    ok, rejected = {}, []
    for f in fields:
        if f.span not in doc_text:
            rejected.append((f.name, "span not in document"))
        elif not validators[f.name](f.value):
            rejected.append((f.name, "failed format check"))
        else:
            ok[f.name] = f.value
    return ok, rejected

def decide(plot_states, fields, rejected):
    # The LLM cannot reach this function except through verified fields.
    if rejected or "permit_id" not in fields:
        return "HOLD", "document evidence incomplete"
    if any(s == "loss" for s in plot_states.values()):
        return "REJECT", "loss after cut-off on declared plot"
    if any(s == "unobserved" for s in plot_states.values()):
        return "HOLD", "plot not observable for part of the period"
    return "RELEASE", "all plots observed clear; documents verified"

The extractor has no tools, so an injected instruction has nothing to call, and the summary a reviewer reads is generated from verified fields and rule outcomes, never from raw documents. The grounding check pattern applies: every summary sentence must trace to a field or a rule result.

Fire detection and alert budgets

Fire detection is rarely attacked, but it fails in security-shaped ways. Hotspot and smoke classifiers raise false alarms from glint, industrial heat and fog, and a system that pages for every one trains operators to ignore it. Fuse independent sources before paging, treat alert volume as a budget, log dismissals with reasons, and alert when a camera goes silent, since silence is what an intruder or a failed link looks like.

Worked example: one shipment, three plots

A sawmill imports one shipment declared against three plots. Here is what the pipeline does with it.

PlotDeclaredGeometry gateChange detectionState
P1polygon, 12 ha, 900 m3valid; 75 m3/ha plausibleobserved clear in every quarterclear
P2point, 3 ha, 400 m3point allowed; 133 m3/ha plausibleradar fills a 10-week cloud gap; clearclear
P3polygon, 40 ha, 1,100 m3overlaps 85 percent with a polygon from another supplier1.8 ha loss in 2023 in the south-east cornerloss

The LLM extracts a permit number, an issue date and a declared species from four documents. One declaration contains, in white text, the line that all plots were independently verified as deforestation-free. The extractor has no field for that claim, so it never enters the decision. The permit number passes the registry format check and its span is found verbatim.

The rule engine sees P3 in the loss state and returns REJECT for the shipment, with two reasons a reviewer can check: the loss polygon with dates and scene identifiers, and the overlap with the other supplier's declaration. Note what a naive design would have done. Loss of 1.8 ha in a 40 ha polygon is 4.5 percent of the area; a share-of-area threshold of 5 percent would have passed it, which is exactly the dilution attack. The rule is any confirmed loss after the cut-off, with the share reported only as context. The hidden instruction is logged as a supplier-integrity signal that earns that supplier's other declarations a closer look.

Failure modes

FailureHow it showsGuard
Unobserved treated as clearClearing timed to cloud season passesThree-state verdicts; gaps mean hold
Share-of-area thresholdLarge polygons hide small clearingsAny confirmed loss rejects; share is context
Unpinned model versionsSame plot gets different verdicts after retrainVersion and scene list stored per verdict
Supplier-chosen validation plotsAccuracy looks high, field reality differsYour own randomly sampled field holdout
LLM summary as evidencePersuasive supplier prose reaches reviewersSummaries generated from verified fields only
Mean carbon creditedOver-crediting within the model errorCredit the conservative bound

Operations and trade-offs

Run the pipeline as an evidence system. Every verdict stores its inputs by hash: geometry, scene identifiers, model versions, extracted fields and spans, and rule-set version. That record answers both auditors and disputing suppliers. Keep the rule engine small and change it through review, like code.

The trade-offs are real. Holding unobserved plots delays shipments in cloudy regions; the answer is radar coverage, not a looser rule. Strict geometry checks reject smallholders whose plots someone else mapped badly; offer a field-verified correction path rather than waiving the check. Prefer fewer, typed LLM fields to rich summaries. And because measured parties adapt, red-team each season: try dilution, borrowing and injection against your own pipeline before suppliers do.

What to do next

  1. Draw your evidence chain and mark which inputs come from parties who gain from the verdict.
  2. Add a geometry gate: validity, precision, size rule, yield plausibility, hash and overlap checks.
  3. Change verdicts to three states and make not-observed a hold, with radar to fill optical gaps.
  4. Pin model versions, scene lists and rule sets per verdict, and store them by hash.
  5. Build a field-verified holdout that you choose, and gate every retrain on it.
  6. Restrict the LLM to typed extraction with verbatim spans, and decide only in rules.
  7. Credit carbon at the conservative bound and make every calculation reproducible.
  8. Re-check the current EUDR text and dates before relying on any threshold quoted here.
Key takeaway: Forestry AI decides things that the people supplying its inputs care about, so build it like a fraud system: validate geometry before imagery, treat unobserved as unverified, credit carbon at the conservative bound, let the LLM extract but never decide, and store every verdict with the hashed evidence that produced it.