Ask a finance team what one answer from your assistant costs and you will usually get a price per million tokens. That is the right unit for a GPU or an API contract, but it is the wrong unit for a product. A user does not buy tokens; they send one query, and that query fans out into an embedding, a vector search, a rerank, one or more model calls, tool calls, a guardrail check and, sometimes, a retry. The bill for the query is the sum of all of it, including the work whose output the user never saw.

This page prices the query as a graph of spans. It builds on LLM Cost Analysis, in depth, which turns a GPU-hour into a cost per token, and on LLM Cost Attribution, in depth, which splits a shared batched step among requests. Here we go one level up: a cost model you can run over traces, a worked example for a retrieval pipeline and an agent, the distribution rather than the mean, and the metric that matters most, cost per successful query. Every price on this page is an illustrative assumption, labelled as one; substitute your contract prices and measured throughputs before you quote any number.

Query, request and token are different units

Three units get confused. Cost per token is a property of a model on a piece of hardware at a given batch size. Cost per request is the cost of one call to one model. Cost per query is what one user intent costs end to end, and it is the only one of the three that maps to revenue, to a free-tier budget or to a price per seat.

The difference is not academic. A retrieval pipeline that makes one model call has a query cost close to its request cost. An agent that plans, calls tools and re-reads its growing context can make a dozen calls, so a cheap per-token price hides an expensive query. A pipeline that retries on a failed schema check pays twice for some queries and once for others. None of this is visible in a per-token dashboard.

One user query fans out into priced spansUser query1 requestEmbedGPU-secondsVector searchfixed per callRerankGPU-secondsLLM call (planner)in, cached, out tokenstool callTool / APIfixed or GPUresult appendedContext grows+tokens per turnLLM call (answer)cache hit on prefixGuardrail checkmay block or retryretry on failureRetryrepeats spansanswerQuery cost = sum of span costs, including retries and spans whose answer was never shownPlanner loop repeats until done or a step budget stops it
A query is a tree of spans. Each span has its own pricing basis: tokens for hosted models, GPU-seconds for self-hosted models, a fixed fee for search or tool calls.

Three pricing bases

Every span falls into one of three pricing bases, and the model needs all three.

  • Token-priced spans. Hosted model calls bill fresh input tokens, cached input tokens (when the provider supports prompt caching, at a discount) and output tokens, each at its own rate. Output is usually several times the input rate, so verbose answers cost more than long prompts.
  • GPU-time spans. Self-hosted embedders, rerankers and models cost GPU-seconds. The honest conversion divides by utilization: a span that kept a GPU busy for 4 ms on a fleet running at 50 percent average utilization really consumed 8 ms of paid time. cost = gpu_seconds / utilization * price_per_gpu_hour / 3600.
  • Fixed-fee spans. Vector database queries, web search APIs, code sandboxes and third-party tools bill per call or per unit of their own.

The fully loaded GPU-hour price should come from your infrastructure cost model, not from a cloud list price; LLM TCO, in depth builds one. For batched self-hosted LLM calls, use the attribution rule from the attribution article to get each call's share of a step, then treat the result as GPU-seconds here.

A per-query cost model in code

The model is deliberately small: a price table, one function per pricing basis, and a sum. It runs over spans you already emit for tracing.

PRICE = {  # illustrative USD per million tokens, not any provider's list price
    "llm_large": {"in": 3.00, "cached_in": 0.30, "out": 15.00},
    "llm_small": {"in": 0.25, "cached_in": 0.025, "out": 1.25},
}
GPU_HOUR = 4.00  # assumed fully loaded USD per GPU-hour for self-hosted spans

def span_cost(span):
    kind = span["kind"]
    if kind in PRICE:
        p = PRICE[kind]
        cached = span.get("cached_tokens", 0)
        fresh = span["in_tokens"] - cached
        return (fresh * p["in"] + cached * p["cached_in"]
                + span["out_tokens"] * p["out"]) / 1e6
    if kind == "gpu":
        return span["gpu_seconds"] / span["utilization"] * GPU_HOUR / 3600
    if kind == "fixed":
        return span["usd"]
    raise ValueError(f"unpriced span kind: {kind}")

def query_cost(spans):
    return sum(span_cost(s) for s in spans)

Two design choices matter. First, the function raises on an unknown span kind instead of pricing it at zero; a new tool that silently costs nothing is how cost reports drift away from invoices. Second, retries are just more spans with the same query id, so they are counted without special handling.

Worked example: retrieval Q&A and an agent

Take two products built on the same large model, with the prices above.

Retrieval Q&A. Embed the query (4 ms of GPU at 50 percent utilization), one vector search (assume 0.00002 USD), rerank 50 passages (30 ms of GPU), then one model call with 6,000 input tokens of which 1,500 are a cached system prompt, and 350 output tokens. If your provider bills cache writes above the base input rate, add a cache-write rate to the price table; terms differ by provider, so take them from your contract.

SpanBasisCost (USD)
Embed0.004 GPU-s / 0.50.000009
Vector searchfixed0.000020
Rerank0.030 GPU-s / 0.50.000067
Model call4,500 fresh + 1,500 cached in, 350 out0.019200
Query0.019296

The retrieval machinery is half a percent of the query. The model call is everything, and within it the 350 output tokens cost 0.00525 USD, more than a quarter of the total. Optimising the reranker would be wasted effort here.

Agent. Start with a 3,000-token context. Each step makes a model call that emits 200 tokens and receives an 800-token tool result, so the context grows by 1,000 tokens per step; a final call writes a 400-token answer. Assume prefix caching turns everything except the newest 800 tokens into a cache hit on each call after the first; that includes the model's own earlier output, which only holds if the history is resent byte for byte.

StepsModel callsWith caching (USD)Without caching (USD)
120.01730.0300
340.03090.0690
670.05360.1500
12130.10710.3930

Without caching the cost grows roughly with the square of the step count, because every call re-reads the whole history: twelve steps read 117,000 input tokens in total. Caching turns that into near-linear growth, a 3.7x saving at twelve steps. The same product, the same model and the same per-token price give query costs that differ by more than twenty times depending on how many steps the agent takes and whether its prefix stays stable enough to hit the cache. Prefix Caching in Depth explains what a hit skips on a self-hosted GPU.

The distribution, not the mean

Because query cost depends on behaviour, it is a distribution, and the mean hides its shape. Simulate 100,000 agent queries whose step count is one plus the floor of an exponential draw with mean two, capped at 25:

StatisticCost per query (USD)
Mean0.0282
p500.0240
p950.0536
p990.0881
Share of spend from the top 5% of queries13.3%

With caching on, this tail is mild: p99 is about 3.7 times the median. Run the same simulation with caching off and the mean rises to 0.0646 USD, p99 to 0.30 USD, and the top 5 percent's share of spend from 13.3 to 18.8 percent, because long queries pay the quadratic re-read. Real traffic usually has a heavier tail than this exponential step count: a few users paste whole documents and a few prompts send the agent into loops, and those rare queries are what come to dominate a bill.

Report at least p50, p95, p99 and the spend share of the top percentiles, per product surface and per model route. Set budgets on the tail, not the mean: a per-query cap of, say, five times the p95 catches runaway loops without touching normal traffic.

Cost per successful query

The metric to put in front of the business is cost per successful query:

cost_per_success = sum(cost of every span of every query) / count(queries judged successful)

Everything in the numerator counts: queries that timed out, streams the user abandoned, answers a guardrail blocked, schema failures that were retried, and the first attempt of every retry. The denominator counts only queries that delivered something useful, which you define once and measure consistently: a thumbs-up, a task completion, an answer that passed an automated grader, or simply a non-error response that was displayed.

This metric exposes trade-offs the per-token view cannot. A smaller model may halve the cost per call and raise the failure rate from 5 to 20 percent; if failures trigger a retry on the large model, cost per success can go up. A stricter guardrail lowers risk and raises cost per success, because blocked answers were still paid for. Write both numbers down before switching.

Measuring it in production

Do not estimate query cost from averages; measure it from traces.

  1. Propagate a query id. Generate it at the edge and pass it to every model call, tool call and retry, including background calls such as summarising history.
  2. Record the billing inputs on each span. Model name, input, cached and output token counts as reported by the provider or the server, GPU time for self-hosted spans, and the unit count for fixed-fee tools. Store counts, not costs, so a price change can be replayed over history.
  3. Price in a batch job. Join spans to a versioned price table, sum per query, and write one row per query with its outcome.
  4. Reconcile monthly. Total priced spans should match provider invoices and the GPU bill within a small tolerance; the gap is untraced traffic or idle capacity. LLM Chargeback, in depth covers who carries idle capacity.

Levers, ranked by where the cost is

Rank levers by where the money is, which the span breakdown tells you.

LeverMovesWatch for
Stable prompt prefixcached share of inputdynamic content (dates, ids) early in the prompt breaks the cache
Step budget for agentstail of the distributiontoo low a cap lowers success rate
Context trimminginput tokens per calldropping facts the answer needed
Shorter outputsoutput tokens, the expensive sideterse answers that fail the user
Route easy queries to a small modelprice per callfailure-triggered retries on the large model
Response cachewhole queries skippedstale or personalised answers served to the wrong user

Exact repeats are cheaper still: Response Caching for LLM APIs covers keys and invalidation. Batching and utilization only help self-hosted spans, and in the retrieval example above those were half a percent of the query.

Failure modes

  • Pricing from estimated tokens. Client-side tokenizer counts drift from what the provider bills, especially for images and tool schemas. Use the usage numbers returned with each response.
  • Dropping failed spans. Traces that end in an error are often sampled out or never closed. They were billed; keep them.
  • Forgetting hidden calls. Query rewriting, conversation summarisation, moderation and evaluation calls ride along with user traffic and are easy to leave untraced.
  • Ignoring utilization on self-hosted spans. Pricing busy GPU-seconds only makes a half-idle fleet look twice as cheap as it is.
  • Mean-only reporting. A stable mean can hide a growing tail from a new agent behaviour until the monthly invoice arrives.

Trade-offs

Per-query costing needs tracing discipline and a pricing job; per-token dashboards are free from your provider. The return is that you can answer product questions directly: what a free-tier user costs per month, whether a feature pays for itself, and which change moved the bill. Pricing from stored counts makes history replayable but means two systems must agree on model names and price versions. Budgets on the tail stop runaway queries but will occasionally cut off a legitimate long task, so return a clear message, not a silent truncation. Finally, cost per success depends on how you define success; pick a definition you can measure every day, even if it is crude.

What to do next

  1. Add a query id to every span, including retries, rewrites and moderation calls.
  2. Record provider-reported input, cached and output tokens, GPU time and tool units on each span.
  3. Build a versioned price table and a batch job that writes one priced row per query with its outcome.
  4. Plot p50, p95 and p99 cost per query and the top-5-percent spend share per product surface.
  5. Define success, then report cost per successful query next to the success rate.
  6. Check your cached-token share; move dynamic content to the end of the prompt if it is low.
  7. Set a per-query step and spend cap from the measured p95, and alert when a query hits it.
  8. Reconcile priced spans against invoices and the GPU bill each month.
Key takeaway: Price the query, not the token: sum token, GPU-time and fixed-fee spans over every call a query makes, including retries and failures, report the distribution and its tail, and judge changes by cost per successful query.