When an agent calls the wrong tool, passes a malformed argument or loops on the same failing call, the usual reaction is to change the system prompt or switch to a bigger model. Most of the time the problem is the tool itself. A model chooses and fills tools using nothing but what you show it: the name, the description, the input schema, and whatever comes back when it calls. Those four surfaces are a prompt, read on every turn, and they deserve the same care and testing as any other prompt.

This article is about designing that surface. It covers how many tools to expose and at what granularity, how to name and describe them, how to constrain arguments, how to word errors so the model can recover, how to shape what comes back, how to control when tools are called, and how to evaluate tool definitions the way you evaluate code. The runtime that executes calls is covered in agent tool use architecture; this page is about what the model reads. API details are from the Claude API documentation, checked on 2026-10-01; other providers have equivalents with different field names.

Advertisement

The tool definition is a prompt

When you pass tools to a model API, the provider renders the names, descriptions and JSON schemas into the context the model reads before every response. The model has never seen your codebase, your database or your internal naming conventions. If two tools have overlapping descriptions, it has to guess between them. If an argument is called id with no further explanation, it has to guess which id. Every ambiguity becomes a probability of a wrong call, and in a multi-step task those probabilities multiply.

So design starts from the model's point of view. For each tool, ask: from the name and description alone, would a capable new colleague know when to use it, when not to, and exactly what to pass? If not, the model will not either.

The model only ever sees four surfaces of a tool; each one is a place to win or lose a callNamenamespaced verb_nounDescriptionwhat, when, when notInput schemaenums, required, strictResults and errorshigh signal, teachModel decideswhich tool, which arguments, what nextYour validatorschema plus business rulesYour handlerreal API, permissionscallEval harnesstasks, traces, metricslogsreviseTreat every surface as a prompt you test: the eval harness feeds back into names, descriptions, schemas and messages
The four surfaces the model reads: name, description, input schema, and results or errors. Your validator and handler sit behind them, and an evaluation harness feeds failures back into the definitions.

Granularity: design around tasks, not endpoints

The most common mistake is to wrap an existing REST API one endpoint per tool. A CRM with list_contacts, get_contact, get_contact_notes, get_contact_deals and get_deal produces five tools and forces the model to chain several calls, carrying ids between them, to answer one question. Each hop costs a turn, tokens and a chance of error.

Design tools around what a user would ask for. A single crm_lookup_contact that accepts a name or email and returns the contact with recent notes and open deals answers the common question in one call. Anthropic's tool guidance makes the same point: consolidate related operations into fewer tools, for example one tool with an action parameter instead of separate create, review and merge tools. The opposite failure is a do-everything tool with a free-text query argument, which pushes all ambiguity into one field. Aim for each tool to correspond to one recognisable intent.

Count matters too. Every tool adds tokens and another option to confuse. If you need many tools, group and namespace them, and consider exposing only the relevant subset per turn, as described in dynamic tool selection.

Advertisement

Names and descriptions

Names should be verbs plus nouns, namespaced by service when tools span several systems: github_list_prs and slack_send_message cannot be confused with each other or with a generic send_message. In the Claude API, a name must match the pattern ^[a-zA-Z0-9_-]{1,128}$, so keep to letters, digits, underscores and hyphens. Avoid near-duplicates such as search and find, because the model will treat them as interchangeable.

The description is the most important field. The Claude documentation calls detailed descriptions by far the most important factor in tool performance and suggests at least three to four sentences per tool. A good description says what the tool does, when to use it, when not to use it and which tool to use instead, what each parameter means, what the result contains, and any limits such as maximum results or data freshness.

# Weak: the model must guess scope, inputs and output.
{"name": "search", "description": "Searches records.",
 "input_schema": {"type": "object", "properties": {"q": {"type": "string"}}}}

# Strong: scope, when not to use it, parameter meaning, output and limits.
{"name": "crm_find_contacts",
 "description": ("Finds customer contacts in the CRM by name, email or company. "
                 "Use it when the user refers to a person or company and you need their "
                 "contact_id. Do not use it for deals; use crm_get_deal for those. "
                 "Returns at most `limit` contacts sorted by last activity, each with "
                 "contact_id, name, email, company and last_activity_date (ISO 8601). "
                 "Data is at most 15 minutes old."),
 "input_schema": {"type": "object",
     "properties": {
         "query": {"type": "string", "description": "Name, email or company; partial matches allowed."},
         "status": {"type": "string", "enum": ["active", "churned", "any"],
                    "description": "Filter by account status. Use 'any' if unsure."},
         "limit": {"type": "integer", "minimum": 1, "maximum": 25,
                   "description": "Maximum contacts to return. Default 10."}},
     "required": ["query"]}}

Schemas that make wrong calls impossible

Every constraint you put in the schema is one the model does not have to infer. Use enums for any argument with a closed set of values. Mark required fields as required. Put units and formats in names and descriptions: amount_cents and start_date with ISO 8601 leave no room for dollars or 03/04. Prefer stable identifiers the model got from a previous result over names it might misspell. Avoid arguments that are themselves JSON strings or free-form query languages unless the task truly needs them; a nested object with typed fields is easier to fill correctly.

Two API features help. Setting strict: true on a Claude tool definition enables strict tool use, which guarantees that tool inputs match the schema, eliminating missing parameters and type mismatches. The optional input_examples field accepts a list of example inputs, each validated against the schema, which helps with nested or format-sensitive arguments at a cost of some prompt tokens per example. Other providers have their own strict modes with their own schema restrictions; check them before relying on a particular schema feature.

Strict mode guarantees shape, not meaning. A schema-valid refund for the wrong order is still wrong, so keep server-side validation of business rules: ownership, limits, state transitions and permissions. When the schema changes later, follow tool schema versioning rather than editing a live definition in place.

Errors that teach the model to recover

When a call fails, the error message is the only thing the model has to decide what to do next. In the Claude API you return it in a tool_result block with is_error set to true. The documentation's advice is to write instructive errors: instead of a bare failed, say what went wrong and what to try next, such as a rate limit with the number of seconds to wait. It also notes that when a call is invalid or missing parameters, the model will typically retry a few times with corrections, so an error that names the fix is often enough for it to repair the call.

Unhelpful errorTeaching error
400 Bad Requeststart_date must be ISO 8601 (YYYY-MM-DD); got '03/04'.
Not foundNo contact with contact_id 'c_991'. Call crm_find_contacts with the person's name to get a valid id.
Permission deniedThis user cannot issue refunds above 50,000 cents. Ask a human approver or refund a smaller amount.
TimeoutThe billing service did not answer in 10 s. The refund was NOT created. Retry once; if it fails again, tell the user.

The last row shows a detail that matters for side effects: state whether the action happened. Retries, idempotency keys and what to do when the outcome is unknown are covered in tool-calling reliability. Never put stack traces, internal hostnames or secrets in errors; they waste tokens and leak information.

Return payloads: high signal, stable ids, explicit truncation

Whatever a tool returns is pasted into the context and read on every later turn. Return only the fields the model needs for its next step, use stable semantic identifiers such as slugs or ids rather than internal row numbers, and convert codes into words: status: churned is clearer than status: 3. When results are paged or truncated, say so in the payload, with the total and how to get more, otherwise the model will assume it saw everything.

A useful pattern is a response_format argument with values such as concise and detailed, so the model can ask for ids only when it is chaining and full records when it is answering. Treat tool results as untrusted data: a web page or email returned by a tool can contain instructions aimed at the model. Keep such content inside tool results and enforce permissions in code, as described under agent guardrails.

Controlling when tools are called

The Claude API's tool_choice has four values: auto lets the model decide and is the default when tools are present; any requires some tool; tool forces a named tool; none forbids tools. Forced choice is not available everywhere: the documentation lists the newest models, including Claude Opus 5.5 and Sonnet 5.5, as returning a 400 error for any and tool, and recommends auto combined with strict tool use, or structured outputs when you need a fixed JSON response. Do not build a design that depends on forcing a call without checking the model you target.

Models can request several tools in one turn. That is ideal for independent reads, such as fetching three contacts at once, and dangerous for dependent or side-effecting actions. Make the dependency visible in the schema, for example by requiring an id that only an earlier call can return, and require explicit confirmation arguments or a human approval step for irreversible actions.

A handler that enforces the contract

Behind the definitions, a thin wrapper turns every outcome into a result the model can use. This sketch validates with the jsonschema package, applies a business rule, and always returns either data or a teaching error.

import json
from jsonschema import Draft202012Validator

def run_tool(tool, args, user):
    errors = sorted(Draft202012Validator(tool["input_schema"]).iter_errors(args),
                    key=lambda e: list(e.path))
    if errors:
        e = errors[0]
        field = ".".join(map(str, e.path)) or "input"
        return {"is_error": True,
                "content": f"Invalid {field}: {e.message}. Fix this argument and call again."}
    try:
        data = HANDLERS[tool["name"]](args, user)    # your real implementation
    except PermissionError as e:
        return {"is_error": True, "content": f"Not allowed: {e}. Do not retry; tell the user."}
    except UpstreamTimeout:
        return {"is_error": True,
                "content": "Upstream timed out; no change was made. Retry once."}
    return {"is_error": False, "content": json.dumps(data, separators=(",", ":"))}

The returned dictionary maps onto a tool_result block with the tool_use_id of the call. Keep the wrapper generic so every tool gets the same error discipline.

Evaluating tool definitions

Tool definitions change model behaviour as much as prompts do, so test them the same way. Build a set of realistic tasks, ideally from real user requests, each with an expected outcome and, where it matters, the expected tool sequence. Run the agent on every task and log full traces. Track the rate of correct tool choice, the rate of schema-invalid or rejected calls, calls per task, tokens per task, recovery after an error, and task success.

Then read the failing transcripts. Repeated confusion between two tools means their descriptions overlap. Many invalid arguments for one field mean it needs an enum, a format or an example. Long chains of calls mean the granularity is too fine. Change one definition at a time and rerun the set. If you train or fine-tune a small model for tool calling, the same harness is the evaluation for it, as in small models for tool calling.

Worked example: refactoring a CRM toolset

A support agent was given eleven tools generated from a CRM's REST API: list, get, create and update for contacts, notes and deals, plus a search endpoint. Traces showed typical failures: the model called list_contacts and paged through results instead of searching, passed contact names where ids were required, and confused update_deal with create_note when a user asked to log a call.

The refactor produced four tools. crm_find_contacts searches by name, email or company. crm_get_contact returns a contact with recent notes and open deals in one payload. crm_log_activity takes contact_id, an enum activity_type of call, email or meeting, and a summary. crm_update_deal takes deal_id and an enum stage, with a required confirm boolean for closing a deal. Every description now says when not to use the tool and points to the right one. Not-found errors tell the model to call crm_find_contacts. The team reran the same task set and compared the metrics above before and after, and kept the change only because the numbers improved on their own data; you should expect to do the same rather than assume a gain.

Failure modes

  • Overlapping tools. Two tools that can answer the same question split the model's choices. Merge them or make the boundary explicit in both descriptions.
  • Silent truncation. A result cut at fifty rows with no marker leads to confident wrong answers.
  • Opaque errors. A generic failure message produces loops of identical retries.
  • Schema drift. The handler accepts arguments the schema does not describe, so the model never learns them, or rejects ones it does describe.
  • Bloated results. Full records returned on every call fill the context and push out earlier instructions.
  • Trusting the schema for safety. Strict validation does not check permissions or intent.

What to do next

  1. List your tools and the user intents they serve; merge tools that serve one intent and split tools that serve many.
  2. Rewrite each description to cover what, when, when not, parameters, output and limits.
  3. Add enums, required fields, formats and units to every schema, and enable strict mode where your provider supports it.
  4. Route every call through one wrapper that validates, enforces business rules and returns teaching errors with is_error set.
  5. Trim result payloads to high-signal fields with stable ids and explicit truncation markers.
  6. Build a task set with traces and metrics, and change one definition at a time against it.
Key takeaway: A model calls tools using only their names, descriptions, schemas and results, so those surfaces are prompts. Design tools around user intents, name and describe them precisely, constrain arguments with schemas and strict mode while still validating business rules in code, return errors that tell the model what to do next and payloads that carry only what it needs, and measure every change on a task set.