Most LLM cost incidents are not caused by one expensive prompt. They are caused by a chain: an agent loop that calls an application retry wrapper, which calls a gateway with its own retries, which calls a provider SDK that retries too, against a model whose alias, price or tokenizer changed upstream last week. Each link was configured by someone different, and nobody owns the product of all of them.

This article treats LLM spend as a supply chain problem with a security edge. It maps every hop that can add or multiply cost, shows the threats that exploit those hops, and builds four controls: a cost bill of materials (BOM) per feature, a usage ledger that prices every call from a versioned price table, a cost regression test in CI, and unit-cost anomaly alarms. Request-time admission control, the reserve-then-settle token budget, is covered in LLM denial of service and is not repeated here.

The cost supply chain

Start from first principles. Each model call costs roughly input tokens times the input price, plus output tokens times the output price, with adjustments for cached input and, on reasoning models, for thinking tokens, which providers bill as output. Everything else is about how many calls happen and how many tokens each one carries.

That number is shaped by things you own and things you do not. You own the prompt, the step limit and your retry policy. You partly own the libraries you depend on, which set defaults on your behalf. You do not own the provider's price sheet, which model a floating alias points to, how the tokenizer splits your text, or the default reasoning effort on a new model version. A change in any of them moves your bill with no diff in your repository.

The LLM cost supply chain: each hop can add spend or multiply itUser / job1 taskAgent loopx stepsApp retriesx attemptsGatewayx retries, fallbackProvider SDKx retries (default 2)Provider model + price sheetalias, tokenizer, prices changeTools with paid APIssearch, OCR, MCP serversBackground spendembedding re-index, evalsCost BOM (pinned models, budgets)checked in CI, owned per featureUsage ledger + unit-cost alarmscost per successful task, reconciled monthlyRed and amber boxes are where spend enters without a code change in your repository.
Multipliers stack along the call path; side branches add spend that never passes through the main model call.

Threats to the bill

The security framing matters because some of these hops are reachable by an attacker, and the rest fail in ways that look like attacks.

ThreatHop it usesTypical signal
Denial of wallet via prompt injectionAgent loop, toolsSteps and tokens per task jump for one tenant
Leaked or stolen provider keyProviderSpend from unknown IPs or models you do not use
Retry stacking during an outageApp, gateway, SDKAttempts per task far above 1 while success falls
Dependency update changes defaultsSDK, frameworkTokens per call change on deploy day
Upstream alias or tokenizer changeProviderSame prompts, different token counts, no deploy
Prompt cache silently brokenYour promptCache-read share drops to near zero
Paid tool fan-outTools, MCP serversThird-party invoice grows faster than LLM spend
Runaway background jobRe-index, evalsSpend outside business traffic patterns

Two of these have their own articles: key theft in API key management for LLM providers, and non-terminating agents in Agent DoS. The rest are covered below.

Retry amplification

Retry stacking is the most common hidden multiplier, and it is a supply chain problem because each layer's default was chosen by a different project. The official OpenAI and Anthropic Python SDKs both default to max_retries=2, which is three attempts. Many gateways and agent frameworks add their own retries on top.

# Worst-case attempts for one logical model call = product of attempts per layer.
layers = {"agent tool retry": 2, "app wrapper": 3, "gateway": 3, "provider SDK": 3}
worst = 1
for name, attempts in layers.items():
    worst *= attempts
print(worst)   # 2 * 3 * 3 * 3 = 54 requests for one step

# Policy: exactly one layer retries, with a budget, and every other layer is set to zero.
#   SDK:      client = Anthropic(max_retries=0)   # or OpenAI(max_retries=0)
#   gateway:  retries=2, retry only 429/5xx/connection errors, honour Retry-After
#   app:      no retry loop; surface the error to the agent as a tool failure

Most failed attempts are cheap, because a 429 or a connection error is not billed. The expensive case is a client-side timeout on a long generation: the client gives up and retries, but the provider may finish the first generation and bill it anyway. During a provider slowdown, every layer times out at once and each retry starts a fresh long generation. Choose one retry owner, usually the gateway, set the others to zero, and make the timeout at each layer longer than the one beneath it so the inner layer always gives up first.

A cost bill of materials

A software BOM lists what you ship. A cost BOM lists what a feature is allowed to spend and which upstream versions that expectation depends on. It is a small file per feature, checked in next to the code and reviewed like code.

# cost-bom/support_summariser.yaml  (one file per feature, reviewed like code)
feature: support_summariser
owner: team-support-ai
model: claude-sonnet-4-5-20250929   # dated snapshot, never a floating alias
price_table: prices/2026-09.yaml   # versioned; a price change is a reviewed diff
per_task:
  max_steps: 4
  max_output_tokens: 1200
  expected_input_tokens: {p50: 9000, p95: 22000}
  expected_cost_usd: {p50: 0.035, p95: 0.090}
  min_cache_read_share: 0.6        # system prompt + policy docs are cached
retries:
  owner: gateway
  sdk_max_retries: 0
paid_tools:
  - {name: web_search, max_calls_per_task: 2}
monthly_budget_usd: 4000

The model name in the example is a placeholder for whatever dated snapshot your provider offers; use the exact identifier from its model list, and check your provider's current price sheet rather than any number here. The point is the structure. Pinning a snapshot turns a silent upstream model change into a planned migration. Pointing at a versioned price table turns a price change into a diff someone reviews. Writing down expected tokens and cache share gives CI and monitoring something to compare against. Naming the retry owner makes stacking visible in review.

A priced usage ledger

You cannot manage what you cannot attribute. Price every call at the point it returns, from the usage object the provider sends back, and tag it with tenant, feature, task and attempt number. Do not estimate from your own token counts: your tokenizer may not match the provider's.

import yaml

def price_call(usage, model, table):
    """usage: the provider's usage object; field names follow the Anthropic Messages API."""
    p = table[model]   # USD per million tokens
    cached_write = usage.get("cache_creation_input_tokens", 0)
    cached_read = usage.get("cache_read_input_tokens", 0)
    return (usage["input_tokens"] * p["input"]
            + cached_write * p["cache_write"]
            + cached_read * p["cache_read"]
            + usage["output_tokens"] * p["output"]) / 1_000_000

def record(ledger, ctx, usage, model, table_version):
    table = yaml.safe_load(open(f"prices/{table_version}.yaml"))
    ledger.append({
        "trace_id": ctx.trace_id, "tenant": ctx.tenant, "feature": ctx.feature,
        "task_id": ctx.task_id, "attempt": ctx.attempt, "model": model,
        "price_table": table_version, **usage,
        "cost_usd": price_call(usage, model, table),
    })

Other providers name these fields differently, so normalise at the gateway and keep the raw object too. Recording the price table version on each row means a later price change does not rewrite history. Recording the attempt number is what lets you see retry stacking. Once a month, reconcile the ledger total against the provider invoice per project. A gap of more than a few percent means spend is reaching the provider through a path that bypasses the gateway, such as a leaked key or a team calling the API directly.

Cost regression tests in CI

Run a fixed set of golden tasks against the real model on every change to prompts, dependencies or the BOM, and fail the build when cost or cache share drifts. Thirty to fifty tasks is usually enough to see a shift in the 95th percentile, and the run costs a few dollars.

def test_cost_regression(bom, golden_tasks, run_feature):
    rows = [run_feature(t) for t in golden_tasks]       # recorded in a sandbox project
    costs = sorted(r.cost_usd for r in rows)
    p95 = costs[int(0.95 * (len(costs) - 1))]
    assert p95 <= 1.2 * bom["per_task"]["expected_cost_usd"]["p95"], f"p95 cost {p95:.3f}"
    cache = sum(r.cache_read_input_tokens for r in rows) / max(1, sum(r.total_input for r in rows))
    assert cache >= bom["per_task"]["min_cache_read_share"], f"cache share {cache:.2f}"
    assert max(r.steps for r in rows) <= bom["per_task"]["max_steps"]

This test is where dependency updates get caught. A framework release that adds a default system preamble, changes the order of messages, or raises a default output cap shows up as a token or cache-share change before it ships. Treat the dependency like any other supply chain input, as the AI supply chain program does for models and packages.

Worked example: a timestamp that tripled input cost

A team runs the support summariser above. Its system prompt plus policy documents are about 18,000 tokens and normally served from the prompt cache. A logging library update adds the current timestamp to the first line of the system prompt for debugging. Every request now has a unique prefix, so every request is a cache miss.

Work through the arithmetic with a cached read priced at one tenth of uncached input, which is the ratio Anthropic publishes for cache reads; your provider's ratio may differ. Before the change, 18,000 cached tokens cost the same as 1,800 uncached tokens. After it, they cost the full 18,000, and each miss also pays the cache write premium for a prefix that is never reused. With a few thousand variable tokens per task, input cost per task roughly triples overnight.

Total spend rises, but it also rises on busy days, so a total-spend alarm is slow to fire. The unit-cost alarm, cost per successful task for this feature, fires within an hour. The CI test would have caught it earlier, because the cache-share assertion falls from about 0.8 to 0.

Unit-cost alarms and hard limits

Alert on cost per successful task per feature, not on total spend. Total spend tracks traffic; unit cost tracks the supply chain. Use a median-based score so one expensive outlier hour does not become the new normal, and require an absolute increase too so cheap features do not page on noise.

import statistics

def unit_cost_alarm(hourly, window=168, k=4.0, floor_usd=0.01):
    """hourly: list of (cost_usd, successful_tasks) for one feature, oldest first."""
    unit = [c / s for c, s in hourly if s > 0]
    if len(unit) < 24:
        return False
    history, latest = unit[-window - 1:-1], unit[-1]
    med = statistics.median(history)
    mad = statistics.median(abs(x - med) for x in history) or 1e-9
    return (latest - med) / (1.4826 * mad) > k and latest - med > floor_usd

Add three simpler alarms next to it: attempts per task above 1.5, spend on any model not listed in a BOM, and spend in a provider project with no gateway traffic. Then place hard limits where they exist: provider-side spend limits on each project or workspace as a backstop that may lag by minutes, gateway budgets per virtual key, and per-request output caps and step limits in the application.

Failure modes

  • Alarms on totals only. A doubling of unit cost hides inside a quiet week.
  • Floating aliases in production. The model changes under you; pin and migrate on purpose.
  • Pricing from local token counts. Tokenizers differ by model version; use returned usage.
  • Budgets without owners. A cap that nobody is paged for gets raised, not investigated.
  • Hard caps on user-facing paths with no degraded mode. Fall back to a smaller model or a shorter answer rather than failing the request outright.

Trade-offs

Single-layer retries mean some transient errors reach users that a stacked policy would have hidden. That is the right trade, because a retry storm during an outage costs more than the errors it hides and slows recovery. Pinned snapshots delay quality improvements until you migrate, and they retire on the provider's schedule, so pinning must come with a calendar. The cost BOM and CI test add a little process to every prompt change. In return, cost becomes a reviewed property of a feature rather than a monthly surprise.

What to do next

  1. Inventory every path that can spend money: models, gateways, paid tools, MCP servers, batch jobs.
  2. Pick one retry owner per call path and set max_retries=0 everywhere else.
  3. Write a cost BOM per feature with a pinned model snapshot and a versioned price table.
  4. Price every call from returned usage into a ledger tagged by tenant, feature, task and attempt.
  5. Add a golden-task cost and cache-share test to CI.
  6. Alarm on unit cost per feature, attempts per task, and unknown models; reconcile invoices monthly.
Key takeaway: LLM cost is the product of many hops, and several of them change without a commit in your repository. Give each feature a cost BOM with a pinned model and a versioned price table, let exactly one layer retry, price every call from the provider's returned usage, test cost and cache share in CI, and alarm on cost per successful task rather than on total spend.