Machine-learning models now sit inside operational weather prediction and are spreading into climate risk work. ECMWF took its Artificial Intelligence Forecasting System (AIFS) into operations on 25 February 2025, alongside its physics-based model. DeepMind's GenCast produces 15-day probabilistic forecasts at 0.25 degree resolution. Open checkpoints for several of these models can be downloaded and run by anyone with a GPU, and generative downscalers turn coarse fields into kilometre-scale maps that feed flood models, insurance pricing and disclosure reports.

That shift changes the security question. A physics model has a long chain of human review around its code; an ML emulator is a file of weights, a data pipeline and an evaluation report. Each can be wrong, stale or tampered with, and the failure mode is not a crash but a plausible map. This article walks the pipeline, names the integrity risks at each stage, and gives the controls, with code, that a team consuming or building ML climate products should run.

The pipeline and its trust boundaries

An ML forecast does not start from observations directly. Observations are combined with a model background through data assimilation into an analysis, the best estimate of the current atmosphere. The emulator takes one or two analysis states and steps forward, typically six or twelve hours per step, feeding its own output back as input. Ensembles run many perturbed or sampled trajectories. Downscalers then add local detail, impact models convert weather into flood depth or crop stress, and increasingly an LLM assistant summarises the result for people.

Training uses reanalysis: a consistent multi-decade reconstruction such as ERA5. The trust boundaries are where data or weights cross from someone else's control into yours: the reanalysis archive, the real-time analysis feed, the downloaded checkpoint, third-party downscaled products, and any document an assistant reads.

An ML weather and climate pipeline and its integrity checkpointsObservationsstations, satellitesAssimilationanalysis, reanalysisML emulatorGraphCast-style, AIFSEnsemblespread, probabilitiesinit stateDownscalerkm-scale samplesImpact modelsflood, heat, cropLLM assistantreads products, answersDecisionsalerts, pricing, disclosureCheckpoints1 dataset manifest + plausibility 2 signed checkpoint, weights-only load 3 leakage-free evaluation4 spread-skill and rank-histogram monitor 5 downscaled output labelled as samples6 every number an assistant states must come from a tool result with dataset and version
Each arrow is a place where data or weights change hands; the six checkpoints are the controls this article builds.

Threat model

StageWhat goes wrongHow it shows upControl
Input feedStale, preliminary or substituted fieldsForecast drifts from peers on day 1Manifest, checksums, plausibility checks
Training dataRevised or mixed dataset versions; poisoned additionsBias in a region or seasonPinned snapshots, data lineage
CheckpointTampered or swapped weights; unsafe deserialisationCode execution, or subtle skill lossSignatures, weights-only loading
EvaluationTest years overlap training; cherry-picked scoresPaper skill not reproduced in operationsYear-split audit, independent scorecard
EnsembleToo little spread, worst on extremesOverconfident probabilitiesSpread-skill ratio, rank histograms
DownscalingInvented fine-scale detailSharp maps with false precisionMultiple samples, labelling, validation
AssistantPrompt injection, invented numbersConfident wrong answersTool-grounded numbers, citation check

Input and training data integrity

Start with the data the model reads. ERA5 has a preliminary stream, ERA5T, released a few days behind real time and potentially revised when the final product replaces it. A training set assembled at two different dates can silently mix the two. Real-time initial conditions arrive as GRIB files, often through mirrors and caches that are not the original publisher. Neither problem is exotic; both produce a model that is slightly wrong in ways no one will notice without a check.

The control is a manifest that pins exactly what was read, plus physical plausibility tests that catch corrupted or substituted fields before they reach the model:

import hashlib, json
import numpy as np
import xarray as xr

BOUNDS = {                       # loose physical bounds; tighten per level and season
    "2m_temperature": (180.0, 340.0),          # kelvin
    "specific_humidity": (0.0, 0.04),          # kg/kg, never negative
    "mean_sea_level_pressure": (85000.0, 110000.0),
}

def sha256(path, chunk=1 << 20):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        while block := f.read(chunk):
            h.update(block)
    return h.hexdigest()

def check_field(ds, var):
    lo, hi = BOUNDS[var]
    a = ds[var].values
    problems = []
    if np.isnan(a).any():
        problems.append("nan")
    if a.min() < lo or a.max() > hi:
        problems.append(f"out of range [{a.min():.4g}, {a.max():.4g}]")
    if np.nanstd(a) == 0:
        problems.append("constant field")    # a classic sign of a fill-value substitution
    return problems

def build_manifest(paths, source, stream):
    entries = []
    for p in paths:
        ds = xr.open_dataset(p)
        issues = {v: check_field(ds, v) for v in BOUNDS if v in ds}
        entries.append({"file": str(p), "sha256": sha256(p), "stream": stream,
                        "time": [str(ds.time.values.min()), str(ds.time.values.max())],
                        "issues": {k: v for k, v in issues.items() if v}})
    bad = [e for e in entries if e["issues"]]
    if bad:
        raise RuntimeError(f"{len(bad)} files failed plausibility: {bad[:3]}")
    return {"source": source, "stream": stream, "files": entries}

Record the stream (final or preliminary) in the manifest and refuse to train on a mix. For real-time runs, compare the day-zero state with a second independent analysis; a large disagreement in one region is a reason to hold the run, not to publish it. The general case of poisoned additions to training data is covered in data poisoning attacks.

Checkpoint provenance

Open weather and climate checkpoints are distributed through code repositories, cloud buckets and model hubs. Two separate risks travel with them. The first is code execution: a PyTorch checkpoint saved with pickle can run arbitrary code when loaded. Since PyTorch 2.6, torch.load defaults to weights_only=True; keep it that way, prefer safetensors where the publisher offers it, and never flip the flag to make a download work. The second is integrity: a swapped checkpoint that loads cleanly and forecasts slightly worse, or worse only in one basin, is far harder to notice than malware.

Pin the exact files by digest, verify a publisher signature where one exists (the mechanics are in model signing), and run a short regression forecast on fixed historical initial conditions after every update, comparing against stored reference output:

def regression_gate(model, cases, reference, tol):
    # cases: fixed historical initial states; reference: outputs from the approved checkpoint
    worst = {}
    for case_id, init in cases.items():
        out = model.rollout(init, steps=8)                 # two days at 6 h steps
        for var, ref in reference[case_id].items():
            rmse = float(np.sqrt(np.mean((out[var] - ref) ** 2)))
            worst[var] = max(worst.get(var, 0.0), rmse)
    failed = {v: e for v, e in worst.items() if e > tol[v]}
    if failed:
        raise RuntimeError(f"checkpoint diverges from approved reference: {failed}")
    return worst

Bind the checkpoint digest, training data manifest and evaluation report together in a model card so that a reader can tell which weights a published skill score belongs to; see model cards.

Evaluation you can trust

Weather data is strongly autocorrelated in time, so a random train-test split leaks: tomorrow is very like today. Credible work splits by year and tests on years the model never saw. When you adopt a model, read which years it trained on and evaluate only on later ones; when you fine-tune it, keep that boundary. A vendor scorecard that does not state its years is not evidence.

Probabilistic forecasts need probabilistic checks. The spread-skill ratio compares ensemble spread with the error of the ensemble mean; a value well below 1 means the ensemble is overconfident, and ML ensembles that look sharp on average can be most underdispersive on the extremes that matter. Monitor it continuously, by region and by event type:

def spread_skill(ens, obs):
    # ens: (members, cases) forecasts at one lead time; obs: (cases,) verifying analysis
    m = ens.shape[0]
    mean = ens.mean(axis=0)
    rmse = np.sqrt(np.mean((mean - obs) ** 2))
    spread = np.sqrt(np.mean(ens.var(axis=0, ddof=1)))
    return float(np.sqrt((m + 1) / m) * spread / rmse)   # about 1.0 when calibrated

def rank_histogram(ens, obs):
    ranks = (ens < obs[None, :]).sum(axis=0)              # 0..members
    return np.bincount(ranks, minlength=ens.shape[0] + 1)

ratio = spread_skill(ens_t2m_day5, obs_t2m_day5)
if ratio < 0.8:
    alert("t2m day-5 ensemble underdispersive", ratio=ratio)

A U-shaped rank histogram means observations keep falling outside the ensemble, the signature of overconfidence. Compute these on the top few percent of events separately; average scores hide the tails.

Downscaling and out-of-distribution climates

Generative downscalers, for example the diffusion-based models NVIDIA ships with its Earth-2 platform, produce realistic kilometre-scale fields from coarse input. Realistic is the risk. Each output is one sample from a learned distribution, and the fine structure, such as which valley gets the heaviest rain, may not be constrained by any observation. A single sample rendered as a flood map looks like precision that does not exist.

Controls: always generate several samples and publish the spread, label outputs as samples in metadata and on every chart, validate against high-resolution station data where it exists, and forbid downstream systems from treating one sample as a deterministic answer. Climate-scenario work adds another trap: an emulator trained on the historical climate is being asked about a warmer one it has never seen, so its behaviour outside the training range needs explicit out-of-distribution tests, not an assumption of physics it does not encode.

Assistants over climate data

Assistants that answer questions over forecasts and climate reports inherit two risks. Retrieved documents can carry injected instructions, and models produce fluent numbers that appear nowhere in the data. The defence that works is structural: numbers come only from tool calls against versioned datasets, and a checker rejects any answer containing a number that no tool returned.

import re

NUM = re.compile(r"-?\d+(?:\.\d+)?")

def grounded(answer, tool_results, question="", rel_tol=0.01):
    allowed = [float(x) for r in tool_results for x in NUM.findall(r["text"])]
    allowed += [float(x) for x in NUM.findall(question)]   # years, lead times the user asked about
    for tok in NUM.findall(answer):
        v = float(tok)
        if not any(abs(v - a) <= rel_tol * max(abs(a), 1.0) for a in allowed):
            return False, tok
    if not all(r.get("dataset") and r.get("version") for r in tool_results):
        return False, "missing dataset or version"
    return True, None

Years and lead times in the question will appear as numbers too, so pass the user question into the allowed set, as the function does. The same approach is used for marine data in AI in ocean science.

Worked example: a flood pricing pipeline

An insurer prices flood cover using a pipeline that downloads an open emulator, runs ensembles from real-time analyses, downscales with a generative model and feeds a flood model. Three things happen in one month. A mirror serves a checkpoint that differs from the publisher's digest. The training refresh mixes preliminary and final reanalysis for the last quarter. An analyst asks the assistant for the 1-in-100 rainfall for a postcode and gets a crisp figure.

With controls in place: the digest pin rejects the mirrored checkpoint before load; the manifest builder refuses the mixed-stream training set; the regression gate would have caught a quietly degraded checkpoint had one got through; the spread-skill monitor flags that day-5 rainfall spread is 0.6 of its error in the basin of interest, so pricing widens its uncertainty load; and the grounding check blocks the assistant answer because its figure came from a single downscaled sample rather than the ensemble statistic. None of these needed a security team to notice anything. They are gates that fail closed.

Failure modes

  • Mixed dataset streams. Preliminary and final data in one training set. Record the stream per file.
  • Unpinned checkpoints. A tag or branch name instead of a digest lets weights change underneath you.
  • Leaky evaluation. Random splits on autocorrelated data inflate skill. Split by year.
  • Average-only metrics. Mean scores hide underdispersion on extremes. Score the tails separately.
  • One sample as truth. Downscaled detail treated as deterministic. Publish spread.
  • Ungrounded assistant numbers. Fluent figures with no source. Enforce the grounding check.

Trade-offs

Every control costs something. Plausibility checks with tight bounds produce false alarms in genuine record events, exactly when you least want a held forecast, so keep hard bounds loose and route borderline cases to a person. Regression gates slow checkpoint updates. Multiple downscaling samples multiply GPU cost. ECMWF reports AIFS forecasts use about 1,000 times less energy than its physics model, which is what makes extra samples affordable, but that is the publisher's figure for its system. The balance is to make the cheap gates mandatory and keep humans for the borderline calls; for comparable trade-offs in another field, see AI in agriculture.

What to do next

  1. Inventory every ML weather or climate product you consume and note who controls its data and weights.
  2. Build a dataset manifest with digests and stream labels, and refuse mixed-stream training sets.
  3. Pin checkpoints by digest, load with weights-only deserialisation, and verify signatures where offered.
  4. Store reference forecasts for fixed cases and run the regression gate on every model update.
  5. Audit evaluation year splits and add spread-skill and rank-histogram monitoring by region and tail.
  6. Label downscaled outputs as samples and publish spread alongside every map.
  7. Put a numeric grounding check in front of any assistant that answers climate questions.
Key takeaway: In ML climate and weather pipelines the dangerous failure is a plausible wrong map. Pin and check the data, pin and verify the weights, evaluate on unseen years and on the tails, treat downscaled detail as samples, and make every number an assistant states trace back to a versioned dataset.