LLM spend fails differently from most cloud spend. A misconfigured virtual machine wastes money at a steady rate; an agent stuck in a tool-calling loop, a prompt template that accidentally inlines a whole document, or a leaked API key can spend a month's budget in an afternoon. The provider dashboard will show it, hours later. A spend alerting system exists to close that gap: detect abnormal spend within minutes, route it to someone who can act, and, when nobody acts, limit the damage automatically.
This article designs that system. It covers why billing data is the wrong primary signal, metering at the gateway with a versioned price table, four families of alert and the burn-rate math that makes them precise, idle spend on self-hosted GPU fleets, an enforcement ladder, a worked incident, and the failure modes. It assumes the cost model from LLM cost analysis and the allocation practices in LLM FinOps; here the focus is the alerting itself.
Why billing data is the wrong primary signal
Provider usage dashboards, cloud billing exports and invoices are authoritative, but they arrive late and coarse. Cloud billing exports typically land hours after the usage and are revised afterwards; provider usage reports aggregate by hour or day and by API key or project, not by your customer or feature. Budgets set in a cloud console notify you when a threshold is crossed, and on most platforms they notify rather than stop spending: Google Cloud budgets, for example, do not cap usage on their own.
So the primary signal must be one you produce yourself, at the moment of each call. Every LLM response carries a usage object with input and output token counts, often split into cached and uncached input. Multiply by the price for that model and you have the cost of the request, attributed to whoever made it, within milliseconds. The billing data is then used for what it is good at: confirming, once a day, that your meter agrees with the bill.
Architecture
The architecture has one rule: every model call goes through a component you own. Usually that is the AI gateway, which already handles keys, routing and retries. If some teams call providers directly, their spend is invisible until the invoice, and the first incident will come from them.
The gateway emits one meter event per completed call, including failed and cancelled ones, because a call that streamed half its output before the client disconnected may still be billed for the tokens produced. A stream processor rolls events into per-minute cost buckets keyed by tenant, feature, model and API key. The evaluator reads those buckets, and the enforcer pushes limits back into the gateway.
Metering every request
The meter event needs enough fields to attribute and to re-price. Keep raw token counts and the price version, not just a dollar figure; when a price changes, or when you discover a pricing bug, you can recompute history.
from dataclasses import dataclass
from decimal import Decimal
# Hypothetical prices in dollars per million tokens. Load from config, never hard-code.
PRICES = {
("model-large", "2026-09-01"): {"in": Decimal("3.00"), "cached_in": Decimal("0.30"), "out": Decimal("15.00")},
("model-small", "2026-09-01"): {"in": Decimal("0.25"), "cached_in": Decimal("0.03"), "out": Decimal("1.25")},
}
@dataclass
class MeterEvent:
ts: float; tenant: str; feature: str; api_key: str; model: str
price_version: str; input_tokens: int; cached_input_tokens: int
output_tokens: int; conversation_id: str; status: str # ok | error | cancelled
def cost(e: MeterEvent) -> Decimal:
pr = PRICES[(e.model, e.price_version)]
uncached = e.input_tokens - e.cached_input_tokens
return (uncached * pr["in"] + e.cached_input_tokens * pr["cached_in"]
+ e.output_tokens * pr["out"]) / Decimal(1_000_000)Field names in provider usage objects differ, and some count cached tokens inside the input total while others report them separately. Normalise once in the gateway adapter, and test the adapter against a real response from each provider, because a double-counted cache field quietly inflates every figure downstream. Batch and discounted tiers need their own price entries.
Four families of alert
Four kinds of alert catch different failures. Most teams start with the first and are surprised by incidents only the others would catch.
| Family | Question it answers | Catches | Misses |
|---|---|---|---|
| Threshold and forecast | Will this month end over budget? | Slow growth, new features launching | Anything fast; fires too late |
| Burn rate | Is the budget being consumed too fast right now? | Loops, leaked keys, traffic spikes | Small sustained overspend |
| Anomaly | Is this hour unusual for this tenant and feature? | Spend shifts hidden inside a stable total | Gradual drift that becomes the new baseline |
| Unit cost | Did each request, conversation or task get more expensive? | Prompt bloat, model routing changes, retry storms | Volume-driven growth |
Forecasts can be simple: month-to-date spend plus the trailing seven-day daily average times the days remaining. Anomaly detection should compare against the same hour of the week, using a robust spread such as the median absolute deviation, because LLM traffic follows business hours and a plain z-score over the last day pages every Monday morning. Unit-cost alerts watch ratios: dollars per conversation, tokens per request, calls per agent task. A conversation that makes 400 model calls is a loop whatever the total spend.
Burn-rate math in dollars
Burn-rate alerting comes from SRE error budgets, described in LLM error budgets, and transfers to dollars directly. Define the burn rate as actual spend in a window divided by the spend that would exactly exhaust the monthly budget at a constant pace. A burn rate of 1 lands on budget; a burn rate of 14.4 sustained for one hour consumes 14.4 * 1 / 720 = 2% of a 30-day budget.
Pair a long window with a short one so that alerts fire fast and also stop fast once the problem is fixed. The thresholds below are the ones the Google SRE workbook suggests for error budgets, and they work unchanged for money:
| Burn rate | Long window | Short window | Budget consumed | Action |
|---|---|---|---|---|
| 14.4 | 1 hour | 5 minutes | 2% | Page |
| 6 | 6 hours | 30 minutes | 5% | Page |
| 1 | 3 days | 6 hours | 10% | Ticket |
def burn_rate(spend, window_hours, monthly_budget, hours_in_month=720):
allowed = monthly_budget * window_hours / hours_in_month
return spend / allowed
RULES = [(14.4, 1, 5 / 60, "page"), (6, 6, 0.5, "page"), (1, 72, 6, "ticket")]
def evaluate(spend_in, budget, scope):
"""spend_in(hours) returns spend for this scope over the trailing window."""
for rate, long_h, short_h, action in RULES:
if (burn_rate(spend_in(long_h), long_h, budget) >= rate and
burn_rate(spend_in(short_h), short_h, budget) >= rate):
return {"scope": scope, "action": action, "rate": rate}
return NoneRun the evaluator per scope: the whole organisation, each tenant, each feature, and each API key, each with its own budget. Organisation-wide alerts alone hide a tenant that is burning its share at 50 times the expected rate while the total looks normal.
Self-hosted GPU fleets: alert on idle dollars
On a self-hosted fleet the cost model flips. Tokens are nearly free at the margin and GPUs cost the same whether they serve traffic or not, so the waste to alert on is idle capacity. Convert fleet telemetry into dollars: node-hours times the hourly rate, split into busy and idle by utilisation. NVIDIA's DCGM exporter publishes per-GPU utilisation, for example DCGM_FI_DEV_GPU_UTIL, which is enough for a first version even though it measures time with a kernel running rather than how hard the GPU works.
- Idle dollars: a GPU allocated to a serving or training job with utilisation below 10% for two hours. Route to the owning team, not to on-call.
- Orphans: nodes with no scheduled workload, typically forgotten development or notebook instances. These are the single largest source of waste in many fleets.
- Autoscaler pinned at maximum: replicas at the ceiling for hours with falling tokens per GPU is a scaling-policy bug or a slow request flood.
- Reserved capacity unused: committed GPUs below their expected occupancy, which money already spent cannot recover but next quarter's commitment can.
The enforcement ladder
Alerts that nobody acts on at 3 a.m. need a backstop. An enforcement ladder escalates only as the evidence and the cost grow:
- Notify. Burn-rate ticket, anomaly digest, forecast warning to the budget owner.
- Throttle. Lower the requests-per-minute and tokens-per-minute limits for the offending key or tenant. Legitimate users slow down; loops starve.
- Degrade. Route the scope to a cheaper model, cap max output tokens, or disable optional features such as long-context retrieval.
- Hard cap. Reject requests for that scope with a clear error once a daily cap is reached. Reserve this for keys and tenants, not for the whole product.
Each rung needs an owner who agreed to it in advance and a documented way to lift it. A hard cap that takes down a paying customer's integration is itself an incident; the ladder trades money against availability, and that trade must be a business decision written into the policy, not a threshold an engineer guessed.
Worked example: an agent retry loop
A team runs an assistant with a monthly LLM budget of $30,000, so the even pace is $1,000 a day or about $41.67 an hour. A Tuesday release changes an agent's tool error handling: when a search tool times out, the agent now retries by re-planning, and each re-plan resends the full conversation.
Normal spend is about $40 an hour. At 14:00 a few hundred conversations hit a slow search backend, those conversations jump from 6 model calls to over 100, and within ten minutes spend runs at $1,300 an hour. The 5-minute window crosses 14.4 almost at once, but the page needs both windows: the trailing hour must hold 14.4 * 41.67 = $600, which it first does at 14:33 with $618 (a burn rate of 14.8). The unit-cost alert on calls per conversation has already fired as a ticket, so the responder looks at the agent, not at traffic.
Policy lets on-call throttle the feature's API key and lower the per-conversation call limit to 20; by 14:38 spend is back near $40 an hour and the 5-minute window clears the page. Excess spend was about $700. Without burn-rate alerts the first signal would have been the next morning's billing export; by 09:00 the loop would have spent about $24,000. The postmortem adds a permanent cap of 25 model calls per conversation, enforced in the gateway, and a synthetic test from synthetic monitoring that runs the agent against a deliberately slow tool before every release.
Failure modes
- Alerting only on billing data. The alert fires correctly and a day late.
- Meter blind spots. Direct provider calls, batch jobs, embeddings and evaluation harnesses bypass the gateway. Reconciliation drift is the symptom; treat drift over a few percent as a bug.
- Stale prices. A price change, a new model or a missing cached-token rate makes every dollar figure wrong. Fail loudly on unknown model and price-version pairs.
- Percentage thresholds early in the month. 80% of budget means nothing on the 2nd and everything on the 29th. Burn rate is pace-aware; fixed percentages are not.
- Alert storms. One incident triggers organisation, tenant, feature and key alerts at once. Group by incident and page once, at the narrowest scope that explains it.
- Leaked keys. Spend appears on a key from unfamiliar IPs or models. Per-key budgets and automatic key suspension at the hard-cap rung limit the blast radius.
- Caps that hide incidents. A throttle stops the spend and the loop keeps failing silently. Enforcement actions must open a ticket of their own.
Trade-offs
| Choice | Benefit | Cost |
|---|---|---|
| Gateway metering | Minutes of latency, per-tenant attribution | Build and maintain adapters; must cover every path |
| Billing-only alerts | Authoritative, no code | Hours to a day late; coarse attribution |
| Automatic throttling | Bounded loss when nobody responds | Can degrade legitimate users |
| Hard caps | Absolute ceiling per key | Turns a cost incident into an availability incident |
| Fine-grained scopes | Catches local runaways | More alerts to tune and route |
What to do next
- List every path that calls a model, and route all of them through one gateway.
- Emit a meter event per call with raw tokens and a price version; load prices from config.
- Build a daily reconciliation against billing exports and alert on drift.
- Give every tenant, feature and API key a budget, and run multiwindow burn-rate rules per scope.
- Add unit-cost alerts on calls per conversation and tokens per request.
- For self-hosted GPUs, alert on idle dollars and orphaned nodes.
- Write the enforcement ladder down, name the owner for each rung, and rehearse lifting a cap.