Security teams plan for a threat landscape that will exist when their controls ship, not the one that exists today. For most threats that gap is small. For threats that depend on what AI systems can do, it is not: an evaluation that showed a model could not complete a multi-step intrusion task last year says little about the model your vendor ships next quarter. AI forecasting is the discipline of making explicit, scored predictions about that capability curve, and for a security programme its job is narrow and practical: decide when to re-run evaluations, when to tighten agent permissions, and how much budget to reserve for defences you do not need yet.

This article explains the evidence forecasters use, why each kind misleads in its own way, how to fit a trend without fooling yourself, how to score forecasts so the good forecasters rise, and how to wire probabilities to tripwires that change controls. It includes runnable Python for a log-linear trend fit with a bootstrap interval and for Brier and log scoring, plus a worked example with deliberately illustrative numbers. It is not a prediction of when any particular capability arrives; it is a method for making and using such predictions responsibly.

Four kinds of evidence

Forecasts rest on four classes of evidence. They fail differently, which is the main reason to use more than one.

Input trends. The resources that go into models are measured more reliably than the capabilities that come out. Epoch AI's 2024 analysis estimated that training compute for frontier models has grown by roughly four to five times per year over the last decade, and the same group tracks hardware cost, dataset size and algorithmic efficiency. Input trends are steady and well documented, but they only become a capability forecast through a second model that maps compute to ability, and that mapping is where most uncertainty lives.

Capability trends. These measure outputs directly. The best-known recent example is METR's March 2025 paper on task time horizons: it timed skilled humans on software and research tasks, measured which tasks agents could complete, and expressed each model's ability as the human task length it completes with 50% success. That horizon had doubled roughly every seven months since 2019. Capability trends are closer to what a security team cares about, but they depend on the task suite, the success threshold and the agent scaffolding, and they are noisier.

Benchmarks. Individual benchmarks move fast and then saturate. A benchmark near its ceiling stops carrying information, and contamination of public test sets inflates scores. Benchmarks are useful for spotting a jump, not for long extrapolation.

Structured judgement. Expert surveys and forecasting platforms such as Metaculus pool human judgement. The AI Impacts survey of 2,778 published AI researchers in October 2023 put a 50% aggregate probability on machines outperforming humans at every task by 2047, thirteen years earlier than the same survey a year before. That shift is itself the lesson: survey answers move with news, and framing changes them substantially. Platforms with scored track records do better than one-off surveys because forecasters are rewarded for accuracy rather than for opinion.

EvidenceStrengthTypical failure
Input trends (compute, cost)Long, well-measured seriesNeeds a mapping to capability
Capability trends (time horizons)Measures what agents can doTask-suite and threshold dependent
BenchmarksFast signal of a jumpSaturation and contamination
Surveys and platformsCovers questions with no dataFraming effects, herding

The forecasting pipeline

Forecasts as a control input: evidence to tripwire to actionInput trendscompute, cost, dataCapability trendstime horizons, evalsJudgementsurveys, platformsQuestion bankresolvable, dated, sourcedAggregatepooled probabilityTripwiresp above thresholdfiresControl changeeval cadence, limits, budgetResolve and scoreBrier, log score, calibrationoutcomereweightQuarterly review: retire resolved questions, add new ones, publish the calibration chartforecasters who score well get more weight next cycle
Evidence feeds a bank of resolvable questions; aggregated probabilities cross tripwires that change controls, and resolution scores feed back into forecaster weights.

The pipeline has a feedback loop on purpose. A forecast that is never resolved and scored is an opinion with a number attached. The diagram's left side gathers evidence, the middle turns it into probabilities on specific questions, and the right side does the only thing the security programme actually needs: when a probability crosses a threshold agreed in advance, a named control changes. The bottom loop resolves questions as their dates pass, scores everyone who forecast them, and reweights.

Writing questions that can be resolved

Most forecasting failures are question failures. "Will AI be able to hack systems by 2027?" cannot be resolved, so nobody can be wrong, so nobody learns. A good question has four parts: a measurable event, a date, a named resolution source, and a rule for ambiguous cases. Write it so two people who disagree today would agree on the answer once the date passes.

question:
  id: FQ-014
  text: >
    By 2027-06-30, will any publicly available model, using the vendor's own
    agent scaffold, solve at least 60% of our internal exploit-chain suite
    (v3, 40 tasks) within the standard 2-hour budget?
  resolution_source: internal eval harness, run quarterly on new releases
  ambiguous_if: suite changed by more than 10% of tasks -> resolve on overlap
  owner: appsec-eval
  linked_controls: [CTRL-agent-egress, CTRL-eval-cadence]

Note that the question is about your suite, not a public benchmark. Public results are evidence for the forecast; the resolution is something you control and can re-run. Keep the bank small, ten to thirty questions, each tied to a control you would actually change. A question that would not change any decision is trivia.

Fitting a trend honestly

Trend extrapolation is the workhorse, and the honest version reports an interval, not a line. Capability measures such as time horizons grow roughly exponentially, so fit a straight line to the logarithm and convert the slope to a doubling time. Then bootstrap the data to see how much the doubling time and the projected crossing date move.

import numpy as np

def fit_doubling(t_years, metric, n_boot=5000, seed=0):
    """Fit log2(metric) = a + b*t; doubling time is 1/b years."""
    t = np.asarray(t_years, float)
    y = np.log2(np.asarray(metric, float))
    b, a = np.polyfit(t, y, 1)
    rng = np.random.default_rng(seed)
    boots = []
    for _ in range(n_boot):
        idx = rng.integers(0, len(t), len(t))
        if len(set(t[idx])) < 3:          # degenerate resample
            continue
        bb, aa = np.polyfit(t[idx], y[idx], 1)
        boots.append((aa, bb))
    return (a, b), np.array(boots)

def crossing_year(params, threshold):
    a, b = params
    return (np.log2(threshold) - a) / b

# Illustrative data: horizon in minutes for successive model releases.
t = [2021.0, 2021.8, 2022.6, 2023.3, 2024.0, 2024.7, 2025.4]
h = [2.0,    3.5,    6.0,    11.0,   18.0,   35.0,   60.0]
(a, b), boots = fit_doubling(t, h)
print(f"doubling time: {12 / b:.1f} months")
years = [crossing_year(p, 8 * 60) for p in boots]   # 8-hour tasks
lo, mid, hi = np.percentile(years, [10, 50, 90])
print(f"8h horizon crossing: {mid:.1f} (80% interval {lo:.1f}-{hi:.1f})")

Three cautions apply to every fit like this. First, the bootstrap only captures noise in the points you have; it says nothing about a regime change, so treat the interval as a floor on uncertainty. Second, the threshold matters: a model's horizon at 80% success is much shorter than at 50%, and a security control usually cares about the reliable rate, so state which one you used. Third, prefer fits to the frontier series (the best model at each date) over all models, or a flood of small releases drags the slope down.

Scoring and aggregating forecasts

A forecast is useful only if you can tell good forecasters from bad ones. Proper scoring rules make honesty the best strategy: a forecaster maximises expected score by reporting their true belief. The two standard rules are the Brier score, the mean squared error between probability and outcome (lower is better, 0.25 is what always answering 50% earns on binary questions), and the log score, which punishes confident wrong answers much more heavily.

import math
from collections import defaultdict

def brier(p, outcome):           # outcome is 0 or 1
    return (p - outcome) ** 2

def log_score(p, outcome, eps=1e-4):
    p = min(max(p, eps), 1 - eps)
    return -math.log(p if outcome else 1 - p)

def calibration(rows, bins=5):
    """rows: (prob, outcome). Returns (bin_mid, mean_prob, freq, n)."""
    groups = defaultdict(list)
    for p, o in rows:
        groups[min(int(p * bins), bins - 1)].append((p, o))
    out = []
    for k in sorted(groups):
        g = groups[k]
        out.append(((k + 0.5) / bins,
                    sum(p for p, _ in g) / len(g),
                    sum(o for _, o in g) / len(g),
                    len(g)))
    return out

def pooled(probs, weights=None):
    """Weighted mean of log-odds, a common aggregation that beats the plain mean."""
    weights = weights or [1.0] * len(probs)
    z = sum(w * math.log(p / (1 - p)) for p, w in zip(probs, weights)) / sum(weights)
    return 1 / (1 + math.exp(-z))

Calibration asks whether events you call 70% happen about 70% of the time; plot mean forecast against observed frequency per bin. Resolution asks whether your forecasts separate events from non-events at all. A forecaster who always says the base rate is calibrated and useless. With only a few dozen resolved questions, neither measure is precise, so compare forecasters on the same questions and expect it to take several quarters before weights mean much.

Worked example: tripwires for coding agents

Suppose a platform team runs coding agents with shell access in CI. Today's controls assume agents cannot reliably chain a known vulnerability into lateral movement, so egress is open to package registries and evaluations run twice a year. The team writes question FQ-014 above and links it to two controls.

Five people forecast. Their probabilities for resolution by mid-2027 are 0.15, 0.30, 0.25, 0.40 and 0.20. Pooling in log-odds with equal weights gives about 0.25. The team agreed thresholds in advance: at 0.20 the evaluation cadence moves from twice a year to every new model release; at 0.40 agent egress moves to an allow-list proxy and build credentials become short-lived; at 0.60 the agent tier loses shell access in production-adjacent pipelines until a sandbox review completes.

So the first tripwire fires immediately: evaluations now run on each release, which is cheap. A quarter later a new model release scores 45% on the suite in the team's own harness. That is evidence, not resolution, and forecasters update; the pooled number reaches 0.48 and the egress change ships. The point is that the expensive control was built when the evidence justified it, and the decision was mechanical because the threshold was agreed before anyone was anxious. When the question resolves, everyone is scored and next cycle's weights shift toward whoever was closer.

Pooled pActionCost
0.20Evaluate every new model releaseLow: harness time
0.40Egress allow-list, short-lived credentialsMedium: platform work
0.60Remove shell access pending sandbox reviewHigh: developer friction

Failure modes

Forecasting programmes fail in recognisable ways.

  • Unresolvable questions. Vague wording means no scores and no learning. Fix the resolution source before collecting a single number.
  • Extrapolating a saturated benchmark. A score near 95% cannot keep doubling; switch to a harder suite or a time-horizon style measure.
  • Single-source anchoring. One famous chart becomes the forecast. Require at least one input trend and one judgement source per question.
  • Threshold confusion. Quoting a 50% success horizon as if it were the reliable rate overstates what an attacker can depend on.
  • Survey framing. Asking when machines can do every task and when every occupation is automated produced very different dates in the same survey. Treat survey numbers as soft evidence.
  • No tripwires. Forecasts nobody acts on drift into slide decks. Every question needs a linked control and an owner.
  • Silent retraction. Teams quietly drop questions that resolved badly. Publish the full calibration record, including misses.

Trade-offs

Explicit forecasting costs analyst time and exposes people to being visibly wrong, which some cultures resist. The alternative is implicit forecasting: controls that silently assume capabilities will not change, which is also a prediction, just an unscored one. Trend fits are cheap and transparent but blind to breaks; judgement sources catch breaks but drift with sentiment; aggregation in log-odds across both usually beats either. Fast evaluation cadence reduces the need for long-range forecasts, which is why the cheapest tripwire is usually re-measuring more often. For context on how competitive pressure shapes release timing, see the AI safety race; for turning forecasts into planned capacity rather than controls, the same methods appear in LLM capacity forecasting.

What to do next

  1. Write five resolvable questions about capabilities that would change one of your controls, each with a date, a resolution source and an owner.
  2. Link each question to an entry in your AI risk register and agree action thresholds before collecting forecasts.
  3. Build the resolution harness first; the capability regression gate is a good template.
  4. Collect independent forecasts from at least four people and pool them in log-odds.
  5. Fit a trend with a bootstrap interval for any question that has a time series, and state the success threshold you used.
  6. Feed fired tripwires into risk monitoring so control changes are tracked like any other alert.
  7. Score resolved questions every quarter and publish the calibration chart.
Key takeaway: AI forecasting earns its place in a security programme only when it is explicit, scored and wired to decisions. Use several evidence classes, ask questions you can resolve on your own harness, report intervals rather than lines, pool forecasters in log-odds and reweight them by Brier score, and agree the thresholds that change controls before the numbers arrive.