LLM spend fails differently from most cloud spend. A misconfigured virtual machine wastes money at a steady rate; an agent stuck in a tool-calling loop, a prompt template that accidentally inlines a whole document, or a leaked API key can spend a month's budget in an afternoon. The provider dashboard will show it, hours later. A spend alerting system exists to close that gap: detect abnormal spend within minutes, route it to someone who can act, and, when nobody acts, limit the damage automatically.

This article designs that system. It covers why billing data is the wrong primary signal, metering at the gateway with a versioned price table, four families of alert and the burn-rate math that makes them precise, idle spend on self-hosted GPU fleets, an enforcement ladder, a worked incident, and the failure modes. It assumes the cost model from LLM cost analysis and the allocation practices in LLM FinOps; here the focus is the alerting itself.

Why billing data is the wrong primary signal

Provider usage dashboards, cloud billing exports and invoices are authoritative, but they arrive late and coarse. Cloud billing exports typically land hours after the usage and are revised afterwards; provider usage reports aggregate by hour or day and by API key or project, not by your customer or feature. Budgets set in a cloud console notify you when a threshold is crossed, and on most platforms they notify rather than stop spending: Google Cloud budgets, for example, do not cap usage on their own.

So the primary signal must be one you produce yourself, at the moment of each call. Every LLM response carries a usage object with input and output token counts, often split into cached and uncached input. Multiply by the price for that model and you have the cost of the request, attributed to whoever made it, within milliseconds. The billing data is then used for what it is good at: confirming, once a day, that your meter agrees with the bill.

Architecture

Spend alerting: meter at the gateway, alert on the stream, reconcile against the billAI gatewayevery LLM callmeter eventtokens, tenant, modelprice tableversioned, datedcost streamper-minute rollupsdollarsGPU fleet metricsnode-hours, utilizationself-hosted costalert evaluatorburn, forecast, anomalyrouterpage, ticket, digestenforcerthrottle, downgrade, cappolicylimitsreconciliation jobbilling export vs meter, dailydrift alertAlerts fire from the meter in minutes; the invoice only checks that the meter is right.
Gateway metering feeds a cost stream; the evaluator alerts and enforces; a daily job reconciles the meter with billing.

The architecture has one rule: every model call goes through a component you own. Usually that is the AI gateway, which already handles keys, routing and retries. If some teams call providers directly, their spend is invisible until the invoice, and the first incident will come from them.

The gateway emits one meter event per completed call, including failed and cancelled ones, because a call that streamed half its output before the client disconnected may still be billed for the tokens produced. A stream processor rolls events into per-minute cost buckets keyed by tenant, feature, model and API key. The evaluator reads those buckets, and the enforcer pushes limits back into the gateway.

Metering every request

The meter event needs enough fields to attribute and to re-price. Keep raw token counts and the price version, not just a dollar figure; when a price changes, or when you discover a pricing bug, you can recompute history.

from dataclasses import dataclass
from decimal import Decimal

# Hypothetical prices in dollars per million tokens. Load from config, never hard-code.
PRICES = {
    ("model-large", "2026-09-01"): {"in": Decimal("3.00"), "cached_in": Decimal("0.30"), "out": Decimal("15.00")},
    ("model-small", "2026-09-01"): {"in": Decimal("0.25"), "cached_in": Decimal("0.03"), "out": Decimal("1.25")},
}

@dataclass
class MeterEvent:
    ts: float; tenant: str; feature: str; api_key: str; model: str
    price_version: str; input_tokens: int; cached_input_tokens: int
    output_tokens: int; conversation_id: str; status: str   # ok | error | cancelled

def cost(e: MeterEvent) -> Decimal:
    pr = PRICES[(e.model, e.price_version)]
    uncached = e.input_tokens - e.cached_input_tokens
    return (uncached * pr["in"] + e.cached_input_tokens * pr["cached_in"]
            + e.output_tokens * pr["out"]) / Decimal(1_000_000)

Field names in provider usage objects differ, and some count cached tokens inside the input total while others report them separately. Normalise once in the gateway adapter, and test the adapter against a real response from each provider, because a double-counted cache field quietly inflates every figure downstream. Batch and discounted tiers need their own price entries.

Four families of alert

Four kinds of alert catch different failures. Most teams start with the first and are surprised by incidents only the others would catch.

FamilyQuestion it answersCatchesMisses
Threshold and forecastWill this month end over budget?Slow growth, new features launchingAnything fast; fires too late
Burn rateIs the budget being consumed too fast right now?Loops, leaked keys, traffic spikesSmall sustained overspend
AnomalyIs this hour unusual for this tenant and feature?Spend shifts hidden inside a stable totalGradual drift that becomes the new baseline
Unit costDid each request, conversation or task get more expensive?Prompt bloat, model routing changes, retry stormsVolume-driven growth

Forecasts can be simple: month-to-date spend plus the trailing seven-day daily average times the days remaining. Anomaly detection should compare against the same hour of the week, using a robust spread such as the median absolute deviation, because LLM traffic follows business hours and a plain z-score over the last day pages every Monday morning. Unit-cost alerts watch ratios: dollars per conversation, tokens per request, calls per agent task. A conversation that makes 400 model calls is a loop whatever the total spend.

Burn-rate math in dollars

Burn-rate alerting comes from SRE error budgets, described in LLM error budgets, and transfers to dollars directly. Define the burn rate as actual spend in a window divided by the spend that would exactly exhaust the monthly budget at a constant pace. A burn rate of 1 lands on budget; a burn rate of 14.4 sustained for one hour consumes 14.4 * 1 / 720 = 2% of a 30-day budget.

Pair a long window with a short one so that alerts fire fast and also stop fast once the problem is fixed. The thresholds below are the ones the Google SRE workbook suggests for error budgets, and they work unchanged for money:

Burn rateLong windowShort windowBudget consumedAction
14.41 hour5 minutes2%Page
66 hours30 minutes5%Page
13 days6 hours10%Ticket
def burn_rate(spend, window_hours, monthly_budget, hours_in_month=720):
    allowed = monthly_budget * window_hours / hours_in_month
    return spend / allowed

RULES = [(14.4, 1, 5 / 60, "page"), (6, 6, 0.5, "page"), (1, 72, 6, "ticket")]

def evaluate(spend_in, budget, scope):
    """spend_in(hours) returns spend for this scope over the trailing window."""
    for rate, long_h, short_h, action in RULES:
        if (burn_rate(spend_in(long_h), long_h, budget) >= rate and
                burn_rate(spend_in(short_h), short_h, budget) >= rate):
            return {"scope": scope, "action": action, "rate": rate}
    return None

Run the evaluator per scope: the whole organisation, each tenant, each feature, and each API key, each with its own budget. Organisation-wide alerts alone hide a tenant that is burning its share at 50 times the expected rate while the total looks normal.

Self-hosted GPU fleets: alert on idle dollars

On a self-hosted fleet the cost model flips. Tokens are nearly free at the margin and GPUs cost the same whether they serve traffic or not, so the waste to alert on is idle capacity. Convert fleet telemetry into dollars: node-hours times the hourly rate, split into busy and idle by utilisation. NVIDIA's DCGM exporter publishes per-GPU utilisation, for example DCGM_FI_DEV_GPU_UTIL, which is enough for a first version even though it measures time with a kernel running rather than how hard the GPU works.

  • Idle dollars: a GPU allocated to a serving or training job with utilisation below 10% for two hours. Route to the owning team, not to on-call.
  • Orphans: nodes with no scheduled workload, typically forgotten development or notebook instances. These are the single largest source of waste in many fleets.
  • Autoscaler pinned at maximum: replicas at the ceiling for hours with falling tokens per GPU is a scaling-policy bug or a slow request flood.
  • Reserved capacity unused: committed GPUs below their expected occupancy, which money already spent cannot recover but next quarter's commitment can.

The enforcement ladder

Alerts that nobody acts on at 3 a.m. need a backstop. An enforcement ladder escalates only as the evidence and the cost grow:

  1. Notify. Burn-rate ticket, anomaly digest, forecast warning to the budget owner.
  2. Throttle. Lower the requests-per-minute and tokens-per-minute limits for the offending key or tenant. Legitimate users slow down; loops starve.
  3. Degrade. Route the scope to a cheaper model, cap max output tokens, or disable optional features such as long-context retrieval.
  4. Hard cap. Reject requests for that scope with a clear error once a daily cap is reached. Reserve this for keys and tenants, not for the whole product.

Each rung needs an owner who agreed to it in advance and a documented way to lift it. A hard cap that takes down a paying customer's integration is itself an incident; the ladder trades money against availability, and that trade must be a business decision written into the policy, not a threshold an engineer guessed.

Worked example: an agent retry loop

A team runs an assistant with a monthly LLM budget of $30,000, so the even pace is $1,000 a day or about $41.67 an hour. A Tuesday release changes an agent's tool error handling: when a search tool times out, the agent now retries by re-planning, and each re-plan resends the full conversation.

Normal spend is about $40 an hour. At 14:00 a few hundred conversations hit a slow search backend, those conversations jump from 6 model calls to over 100, and within ten minutes spend runs at $1,300 an hour. The 5-minute window crosses 14.4 almost at once, but the page needs both windows: the trailing hour must hold 14.4 * 41.67 = $600, which it first does at 14:33 with $618 (a burn rate of 14.8). The unit-cost alert on calls per conversation has already fired as a ticket, so the responder looks at the agent, not at traffic.

Policy lets on-call throttle the feature's API key and lower the per-conversation call limit to 20; by 14:38 spend is back near $40 an hour and the 5-minute window clears the page. Excess spend was about $700. Without burn-rate alerts the first signal would have been the next morning's billing export; by 09:00 the loop would have spent about $24,000. The postmortem adds a permanent cap of 25 model calls per conversation, enforced in the gateway, and a synthetic test from synthetic monitoring that runs the agent against a deliberately slow tool before every release.

Failure modes

  • Alerting only on billing data. The alert fires correctly and a day late.
  • Meter blind spots. Direct provider calls, batch jobs, embeddings and evaluation harnesses bypass the gateway. Reconciliation drift is the symptom; treat drift over a few percent as a bug.
  • Stale prices. A price change, a new model or a missing cached-token rate makes every dollar figure wrong. Fail loudly on unknown model and price-version pairs.
  • Percentage thresholds early in the month. 80% of budget means nothing on the 2nd and everything on the 29th. Burn rate is pace-aware; fixed percentages are not.
  • Alert storms. One incident triggers organisation, tenant, feature and key alerts at once. Group by incident and page once, at the narrowest scope that explains it.
  • Leaked keys. Spend appears on a key from unfamiliar IPs or models. Per-key budgets and automatic key suspension at the hard-cap rung limit the blast radius.
  • Caps that hide incidents. A throttle stops the spend and the loop keeps failing silently. Enforcement actions must open a ticket of their own.

Trade-offs

ChoiceBenefitCost
Gateway meteringMinutes of latency, per-tenant attributionBuild and maintain adapters; must cover every path
Billing-only alertsAuthoritative, no codeHours to a day late; coarse attribution
Automatic throttlingBounded loss when nobody respondsCan degrade legitimate users
Hard capsAbsolute ceiling per keyTurns a cost incident into an availability incident
Fine-grained scopesCatches local runawaysMore alerts to tune and route

What to do next

  1. List every path that calls a model, and route all of them through one gateway.
  2. Emit a meter event per call with raw tokens and a price version; load prices from config.
  3. Build a daily reconciliation against billing exports and alert on drift.
  4. Give every tenant, feature and API key a budget, and run multiwindow burn-rate rules per scope.
  5. Add unit-cost alerts on calls per conversation and tokens per request.
  6. For self-hosted GPUs, alert on idle dollars and orphaned nodes.
  7. Write the enforcement ladder down, name the owner for each rung, and rehearse lifting a cap.
Key takeaway: Meter every model call yourself, price it with a versioned table, and alert on the resulting stream with multiwindow burn rates per tenant, feature and key, plus unit-cost and idle-GPU alerts. Use billing data to reconcile, not to detect, and back the alerts with a pre-agreed ladder of throttle, downgrade and cap.