An agent that uses tools fails in more ways than a model that only writes text, and the failures are harder to see. The final answer can read perfectly while the refund was never issued, the order id was invented, or the search error was quietly ignored. Teams then respond to every bad run with the same fix, more prompt instructions, and the failure rate barely moves, because many of these failures are not prompt problems at all.

This article gives you a taxonomy organised by where in the tool loop a failure happens, shows what each class looks like in a trace, provides detectors you can run automatically, and maps each class to the layer that should fix it: the tool definition, the runtime gate, the orchestration logic or the model. How to design tool definitions well is covered in tool calling best practices, and timeouts and side-effect safety in tool calling reliability. This page is about recognising and measuring failure.

Advertisement

Seven stages, seven places to fail

Tool use fails at a stage; name the stage and you know which layer to fix1. Decidecall or answer?2. Selectwhich tool?3. Fill argumentsgrounded values?4. Gatevalid, allowed, now?5. Executeruntime, not model6. Interpretread result right?7. Continue or stopprogress, budgetnext stepModel failuresskip, wrong tool, invented ids, misread, claim without callRuntime failuresno validation, no ordering rule, no loop limitEnvironment failurestimeouts, partial data, injected text in resultsThe trace records every stage, so every failure in this picture is detectable after the fact.
The tool loop. Stages 1-3 and 6-7 are model decisions; 4 and 5 belong to your runtime, which is where many fixes should land.

Every tool call passes through the same stages. The model decides whether a tool is needed at all, selects one, and fills its arguments. Your runtime then gates the call, checking schema, permissions and preconditions, and executes it. The model interprets the result and decides whether to continue, call again or stop. The tool use runtime article describes how to build that loop; the point here is that each stage has characteristic failures, and the stage tells you who can fix them.

The taxonomy

StageFailureWhat the trace showsFix layer
DecideAnswered from memoryNo call; answer contains facts only a tool could knowPrompt, tool description, eval
DecideUnnecessary callCall whose result is unusedDescription scope, cost signal
SelectWrong toolA plausible but wrong tool for the intentNames, descriptions, fewer overlapping tools
SelectUnknown toolName not in the offered setRuntime rejects; constrained decoding
ArgumentsSchema-invalidValidation error on the callSchema, enums, runtime validation
ArgumentsUngrounded valueAn id or amount never seen in contextRuntime grounding check
GateWrong orderAction before its prerequisite or confirmationRuntime preconditions, state machine
InterpretError ignoredError result, then the answer proceeds as if successError payload design, verification
InterpretClaim without callFinal answer asserts an action no successful call performedClaim-against-trace verifier
InterpretTruncation misreadPartial list treated as completeExplicit truncation markers
InterpretInstruction in resultBehaviour changes after reading tool outputData and instruction separation, permissions
ContinueLoopSame call with same arguments repeatedLoop detector, step budget
ContinueGave up earlyStops after one empty result with the goal unmetPrompt, recovery hints in errors

Two properties make this table useful. Every row has a trace signature, so you can count it automatically. And the fix layer column is not always the model: roughly half the rows are best fixed by deterministic code that does not depend on the model getting it right.

Advertisement

Decision and selection failures

Answering from memory is the quiet one. The user asks for an order status, and the model, primed by earlier turns, writes a confident status without calling get_order. The output looks fine; only the trace shows no call. Detect it by tagging tasks that require fresh data and flagging runs with no relevant call, and reduce it with descriptions that say when the tool must be used, for example that order status must always come from this tool and never from memory. For hard requirements, have the orchestration force a tool call when the router classifies the intent as data-dependent.

Wrong-tool selection grows with the number and overlap of tools. Two search tools whose descriptions differ only in nouns will be confused, and the confusion rate rises as you add tools. The fix is in the toolset rather than the prompt: merge overlapping tools, name them by task, and offer only the tools relevant to the current step. A model that names a tool you did not offer is a separate failure: reject it in the runtime with an error listing the valid names.

Argument failures: invalid is easy, ungrounded is dangerous

Schema-invalid arguments are the easy case, because validation catches them and the model can retry with the validation message. Use enums, formats and required fields so more mistakes become invalid rather than wrong. The dangerous case is an argument that is perfectly valid and not true: an order id with the right shape that the user never mentioned and no tool ever returned, an amount the model computed instead of read, a date resolved against the wrong time zone. These pass every schema check.

The defence is a grounding check in the gate: identifiers in arguments must have appeared earlier in the conversation or in a tool result. It is a cheap string check and catches the most damaging class of argument error.

import json, re

ID_PATTERN = re.compile(r"^(ord|cus|inv)_[A-Za-z0-9]{8,}$")

def ungrounded_ids(call_args: dict, context_text: str) -> list[str]:
    """IDs in the arguments that never appeared in the user turn or an earlier tool result."""
    found = []
    def walk(v):
        if isinstance(v, dict): [walk(x) for x in v.values()]
        elif isinstance(v, list): [walk(x) for x in v]
        elif isinstance(v, str) and ID_PATTERN.match(v) and v not in context_text:
            found.append(v)
    walk(call_args)
    return found

# In the gate, before execution:
bad = ungrounded_ids(call.args, transcript_text(session))
if bad:
    return tool_error(call, f"Unknown id(s) {bad}. Look them up with search_orders first.")

The error message matters as much as the check. Telling the model which tool produces valid ids turns a hard failure into a recovery on the next step. Apply the same idea to amounts on side-effecting tools: require that a refund amount equals a value from the order record, or route it to human approval.

Sequencing and state failures

Some calls are valid alone and wrong in order: issuing a refund before checking eligibility, sending an email before the user confirmed the draft, writing a file before reading its current version. Models follow ordering instructions most of the time, and most of the time is not good enough for actions with consequences. Encode ordering as preconditions the runtime enforces, such as requiring that issue_refund follows a successful check_eligibility for the same order in this session, and that a confirmation token from the user turn is present. When the precondition fails, return an error naming the missing step. Parallel calls add a variant: two calls issued together where the second depends on the first's result, which the runtime should reject or serialise.

Interpretation failures: the ones that lie

The most harmful failures are in reading results, because they produce answers that are confidently wrong. An error result followed by a success message is common when errors are returned as friendly prose the model can skim past; make them structured and unmistakable, with a status field and a next step. Truncated lists read as complete unless the result says so explicitly, with a total count and a marker that more exist. Empty results are often over-interpreted, so no matching orders becomes the customer has no orders, when the search filter was wrong.

A claim without a call is the extreme case: the final answer says the refund has been issued, and no successful refund call exists. Detect it with a verifier that extracts action claims from the final answer and checks each against successful calls in the trace, and block the answer or rewrite it when a claim has no support. The output verification article covers verifiers in general; this one is the cheapest high-value instance. Finally, results can carry instructions, text in a web page or ticket telling the model to do something else. That is prompt injection through tools, covered in the indirect injection article; for failure analysis, flag runs whose actions change right after reading untrusted content.

Control-flow failures: loops and early stops

Loops happen when a call fails in a way the model cannot fix, and it retries the same call with the same arguments. Detect identical consecutive calls and stop after the second repeat with an explicit message, and enforce a per-run step and cost budget so a loop cannot run up a bill. Giving up early is the opposite: the model treats one empty search as final. Recovery hints in empty and error results, such as suggesting a broader query or a different tool, cut this sharply. Both show up clearly in step-count distributions, so a sudden rise in runs hitting the budget is a regression signal.

Detecting failures from traces

If each tool call is a span with name, arguments, validation outcome and result status, most of the taxonomy becomes a function over traces. This classifier labels runs; the labels feed dashboards and eval reports.

from collections import Counter

def classify(trace) -> list[str]:
    """Label one agent run from its spans. Each label maps to a fix layer."""
    labels, calls = [], [s for s in trace.spans if s.kind == "tool_call"]
    sig = Counter((c.tool, json.dumps(c.args, sort_keys=True)) for c in calls)
    if any(n >= 3 for n in sig.values()):
        labels.append("loop.repeat_identical_call")
    for c in calls:
        if c.tool not in trace.available_tools:
            labels.append("select.unknown_tool")
        if c.validation_error:
            labels.append("args.schema_invalid")
        if ungrounded_ids(c.args, trace.context_before(c)):
            labels.append("args.ungrounded_id")
        if c.result_status == "error" and not trace.mentions_error_after(c):
            labels.append("interpret.error_ignored")
    for claim in trace.final_answer_action_claims():   # "I have refunded ..."
        if not any(c.tool == claim.tool and c.result_status == "ok" for c in calls):
            labels.append("interpret.claimed_without_success")
    if trace.task_requires_tools and not calls:
        labels.append("decide.answered_from_memory")
    if trace.hit_step_limit:
        labels.append("control.step_budget_exhausted")
    return sorted(set(labels))

Some classes need a model to judge, such as whether an unused call was unnecessary or whether the wrong tool was chosen, and those should be sampled and reviewed rather than trusted blindly. The deterministic labels, repeats, unknown tools, invalid and ungrounded arguments, ignored errors and unsupported claims, are reliable enough to alert on. The tracing setup this depends on is described in the agent observability article.

Worked example: a support agent's failure budget

Take a customer support agent with eight tools, evaluated on 400 recorded tasks. Illustrative first-pass numbers: 9 percent of runs answered order questions without calling the order tool, 4 percent used an ungrounded order id, 3 percent claimed a refund that had not succeeded, 2 percent looped, and wrong-tool selection was 6 percent, concentrated on two overlapping search tools.

Fixes went to their layers. Merging the two search tools and rewriting descriptions addressed selection. A grounding check in the gate turned ungrounded ids into recoverable errors, which the model fixed on the next step most of the time. A claim verifier blocked unsupported refund claims outright, so that class could no longer reach users even when the model made the mistake. A loop detector and a twelve-step budget capped loops. Only the answered-from-memory class needed prompt work, plus a router that forces the order tool for status intents. The lesson generalises: runtime checks turn a model error rate into a user-visible error rate near zero for the classes they cover, while prompt changes only lower the model rate.

What to do next

  1. Make sure every tool call is traced with arguments, validation outcome and result status.
  2. Implement the deterministic classifier labels and report their rates per day and per release.
  3. Add a grounding check for identifiers and amounts on every side-effecting tool.
  4. Encode ordering rules as runtime preconditions instead of prompt instructions.
  5. Return structured errors with a next step, and explicit truncation markers on lists.
  6. Add a claim-against-trace verifier before final answers that report actions.
  7. Build an eval set with tasks targeted at each failure class, and track each class separately.
Key takeaway: Tool use fails at identifiable stages: deciding, selecting, filling arguments, sequencing, interpreting results and controlling the loop. Each failure class leaves a trace signature, so most can be counted automatically, and many are better fixed by deterministic runtime checks, grounding, preconditions, loop limits and claim verification, than by more prompt text. Classify failures, fix each at its own layer, and track every class separately.