A language model can only produce text. The tools pattern is how that text turns into action: the application describes a set of functions, the model replies with a structured request to call one, the runtime decides whether to run it, runs it, and feeds the result back. Every agent framework implements this loop, and every one of them hides the same small set of design decisions behind its own vocabulary.

This article is about those decisions, independent of framework. The runtime mechanics, meaning the registry, policy gate, parallel dispatch and result shaping, are covered in agent tool use architecture, and how to write names, descriptions and schemas that a model uses well is in tool calling best practices. Here the focus is the shape of each tool: the six kinds of thing a tool can be, why each needs a different contract, and how to compose cross-cutting behaviour without rewriting every tool. You will finish with a working, provider-neutral tool layer you can adapt.

Advertisement

The pattern from first principles

Strip the pattern to its essentials and three parties remain. The model sees a specification: a name, a description and a JSON Schema for arguments. It proposes a call by emitting a name and arguments. The runtime owns everything else: it validates the arguments, checks policy, executes code with real credentials, and returns a result envelope the model can read. The model never executes anything; it only asks.

That division gives the pattern its two rules. First, everything the model emits is untrusted input, so validation and authorisation happen in the runtime, every time, no matter how capable the model is. Second, the result is a message to a reader with a limited context window, so it must be compact, explicit about success or failure, and carry the ids the model needs for the next call.

The tools pattern: the model proposes a call, the runtime decides and executesModelreads specs, emits callValidateschema, args, policyMiddlewareauth, audit, timeoutTool bodyyour codecallThe tool body takes one of six shapesQueryread-only, cacheableCommandside effect, dry-run firstJobstart now, poll laterHandlecursor or resource idAgent as toola sub-agent behind a specHuman as toola question that suspendsResult envelopestatus, data or error, ids for follow-up calls, truncation flags; back into the model's contextnext turn
The tools pattern. The model only proposes calls; validation, middleware and execution belong to the runtime. The body behind a spec can take one of six shapes, each with a different contract, and every result returns as the same envelope.

Six shapes of tool

Most tool bugs come from treating every tool as a synchronous function that returns data. In practice a tool is one of six shapes, and each shape implies a contract for retries, caching, approval and the result.

ShapeExampleContract
Querysearch orders, read a fileread-only and idempotent; safe to retry, cache and run in parallel
Commandsend email, refund, deployside effect; needs authorisation, an idempotency key and usually a dry run
JobOCR a batch, run a reportreturns a job id at once; the model polls or the runtime resumes on completion
Handlepaged search, open datasetreturns a cursor or resource id instead of the whole payload
Agent as toola research sub-agenttakes a task, returns a summary; costly and non-deterministic, needs a budget
Human as toolask the user to choosesuspends the run; resumes with the answer, possibly hours later

Declare the shape explicitly in your registry. A runtime that knows a tool is a query can cache it and run it in parallel; one that knows a tool is a command can require approval and attach an idempotency key; one that knows a tool is human-backed can checkpoint and stop the loop rather than spin.

Advertisement

A minimal, provider-neutral tool layer

Model APIs differ in how they want tool specifications wrapped, but they all take a name, a description and a JSON Schema. Keep a neutral registry and write thin adapters per provider, so tools are written once and survive a model switch.

from dataclasses import dataclass, field
from typing import Any, Callable
import jsonschema

@dataclass
class Tool:
    name: str
    description: str
    input_schema: dict
    fn: Callable[..., dict]
    shape: str                    # query | command | job | handle | agent | human
    read_only: bool = False
    idempotent: bool = False
    middleware: list = field(default_factory=list)

REGISTRY: dict[str, Tool] = {}

def tool(name, shape, schema, read_only=False, idempotent=False, middleware=()):
    def register(fn):
        REGISTRY[name] = Tool(name, (fn.__doc__ or "").strip(), schema, fn, shape,
                              read_only, idempotent, list(middleware))
        return fn
    return register

def specs_for_model():
    # Provider adapters translate this neutral form into each API's tool format.
    return [{"name": t.name, "description": t.description, "input_schema": t.input_schema}
            for t in REGISTRY.values()]

def invoke(name: str, args: dict, ctx: dict) -> dict:
    t = REGISTRY.get(name)
    if t is None:
        return {"status": "error", "error": f"unknown tool {name}"}
    try:
        jsonschema.validate(args, t.input_schema)
    except jsonschema.ValidationError as e:
        return {"status": "error", "error": f"invalid arguments: {e.message}"}
    call = t.fn
    for layer in reversed(t.middleware):          # first listed runs outermost
        call = layer(t, call)
    return call(ctx=ctx, **args)

Three details matter. Validation happens before any tool code runs, and a validation failure becomes an error result the model can read and correct, not an exception that kills the run. Unknown tool names are handled the same way, because models do occasionally invent tools. And the per-tool middleware list is composed at call time, so behaviour like auditing is attached declaratively instead of being copied into every function.

Middleware: compose rather than copy

Cross-cutting concerns such as audit logging, scope checks, caching, timeouts, rate limits and redaction are the same for many tools. Write each once as a layer that wraps the next callable, the same idea as HTTP middleware.

import time, hashlib, json

def audit(t, nxt):
    def wrapped(ctx, **args):
        started = time.time()
        result = nxt(ctx=ctx, **args)
        ctx["audit"].append({"tool": t.name, "args": args, "status": result.get("status"),
                             "ms": int((time.time() - started) * 1000), "run": ctx["run_id"]})
        return result
    return wrapped

def require_scope(scope):
    def layer(t, nxt):
        def wrapped(ctx, **args):
            if scope not in ctx["user_scopes"]:
                return {"status": "error", "error": f"{t.name} needs scope {scope}; ask the user"}
            return nxt(ctx=ctx, **args)
        return wrapped
    return layer

def cache_reads(ttl_s=60):
    store = {}
    def layer(t, nxt):
        if not t.read_only:
            return nxt                            # never cache side effects
        def wrapped(ctx, **args):
            ident = [t.name, ctx["tenant"], ctx["user"], args]
            key = hashlib.sha256(json.dumps(ident, sort_keys=True).encode()).hexdigest()
            hit = store.get(key)
            if hit and time.time() - hit[0] < ttl_s:
                return hit[1]
            result = nxt(ctx=ctx, **args)
            store[key] = (time.time(), result)
            return result
        return wrapped
    return layer

Notice the cache layer refuses to wrap anything that is not read-only, and keys on the tool, tenant and user as well as the arguments; a cache that ignores identity leaks one user's data to another. Order matters: put audit outermost so that denied calls are logged too, then authorisation, then caching, so that a cached answer is never served to a caller who lacks the scope.

The shapes in code

Here are four shapes from an expense-reporting agent, built on the layer above.

@tool("find_expenses", "query", {"type": "object", "properties": {
        "month": {"type": "string", "pattern": "^[0-9]{4}-[0-9]{2}$"},
        "cursor": {"type": "string"}}, "required": ["month"]},
      read_only=True, idempotent=True, middleware=[audit, cache_reads(60)])
def find_expenses(ctx, month, cursor=None):
    "List the user's expense lines for a month. Returns at most 50 lines and a next_cursor."
    page, next_cursor = ctx["db"].expenses(ctx["user"], month, cursor, limit=50)
    return {"status": "ok", "lines": page, "next_cursor": next_cursor}

@tool("submit_report", "command", {"type": "object", "properties": {
        "line_ids": {"type": "array", "items": {"type": "string"}, "minItems": 1},
        "dry_run": {"type": "boolean"}, "confirm_token": {"type": "string"}},
        "required": ["line_ids", "dry_run"]},
      middleware=[audit, require_scope("expenses:submit")])
def submit_report(ctx, line_ids, dry_run, confirm_token=None):
    "Submit expense lines for approval. Call with dry_run=true first, then pass its confirm_token."
    if dry_run:
        plan = ctx["expenses"].plan_report(ctx["user"], line_ids)   # totals, policy flags
        return {"status": "ok", "plan": plan, "confirm_token": plan["token"]}
    if not confirm_token:
        return {"status": "error", "error": "run with dry_run=true first and pass its confirm_token"}
    report_id = ctx["expenses"].submit(confirm_token, idempotency_key=confirm_token)
    return {"status": "ok", "report_id": report_id}

@tool("start_receipt_ocr", "job", {"type": "object", "properties": {
        "receipt_ids": {"type": "array", "items": {"type": "string"}}}, "required": ["receipt_ids"]},
      middleware=[audit])
def start_receipt_ocr(ctx, receipt_ids):
    "Start reading receipts. Returns a job_id; call get_job after poll_after_s seconds."
    job_id = ctx["jobs"].enqueue("ocr", receipt_ids, owner=ctx["user"])
    return {"status": "running", "job_id": job_id, "poll_after_s": 20}

@tool("ask_user", "human", {"type": "object", "properties": {
        "question": {"type": "string", "maxLength": 300},
        "options": {"type": "array", "items": {"type": "string"}, "maxItems": 5}},
      "required": ["question"]}, middleware=[audit])
def ask_user(ctx, question, options=None):
    "Ask the user a question you cannot answer from tools. The run pauses until they reply."
    ticket = ctx["inbox"].post(ctx["user"], question, options or [])
    return {"status": "suspended", "ticket": ticket}            # runtime checkpoints and stops here

The query returns a page and a cursor rather than every line in a month, so the model's context does not fill with data it will not read; that is the handle shape folded into a query. The command makes dry_run a required argument, so the model cannot skip it by omission: the first call returns a plan with totals and policy flags for the user to see, and only a second call carrying its confirm token commits exactly that plan, using the token as the idempotency key so a retry cannot submit twice. Compensation, timeouts and retries for commands are covered in depth in tool calling reliability.

The job returns immediately with a job id and a hint for when to poll. Never let a tool block for minutes inside the model loop: the request times out, the model retries, and you start the work twice. The human tool returns a suspended status, and the runtime treats that as a signal to checkpoint the conversation and stop until the answer arrives. When the decision is an approval rather than a question, use the gate patterns in human approval gates.

Agent as tool

Wrapping a sub-agent behind a tool specification is the cleanest way to compose agents: the parent sees one tool with a task argument and gets back a short summary, while the sub-agent runs its own loop with its own tools and context. It keeps the parent's context small and lets each agent be tested alone. The ADK version of this composition is described in agent-as-tool.

Treat the sub-agent like the most expensive tool you own. Give it a budget in steps, tokens and wall time, and return a structured result with a status, the answer, and the evidence it relied on, so the parent can tell a confident answer from a guess. Give the sub-agent only the scopes its task needs, and log its trace under the parent's run id.

Exposing tools over MCP

The Model Context Protocol standardises the specification and result so tools can be served to any compatible client. A tool there has a name, a description, an inputSchema, an optional outputSchema and optional annotations; a result carries content, optional structuredContent and an isError flag. The annotations map neatly onto shapes:

{
  "name": "submit_report",
  "description": "Submit expense lines for approval. Call with dry_run=true first.",
  "inputSchema": {"type": "object",
                  "properties": {"line_ids": {"type": "array", "items": {"type": "string"}},
                                 "dry_run": {"type": "boolean"}},
                  "required": ["line_ids", "dry_run"]},
  "annotations": {"title": "Submit expense report", "readOnlyHint": false,
                  "destructiveHint": false, "idempotentHint": true, "openWorldHint": false}
}

readOnlyHint marks a query. destructiveHint distinguishes commands that can destroy from those that only add, idempotentHint says a repeat call has no further effect, and openWorldHint says whether the tool reaches outside a closed system. The spec defaults are cautious: a tool with no annotations is treated as not read-only, possibly destructive, not idempotent and open-world. They are hints, though. A client may use them to decide when to ask the user, but the server must enforce authorisation itself, because a hint from an untrusted server is not a guarantee.

Worked example: the expense agent

A user writes: submit my September expenses. The model calls find_expenses for 2026-09 and gets 50 lines and a cursor, then calls again with the cursor and gets 12 more. Six lines have receipt images but no amounts, so it calls start_receipt_ocr and receives a job id with a 20-second poll hint; the runtime parks the run and resumes it when the job finishes, rather than letting the model poll in a tight loop. Two lines are ambiguous between travel and meals, so the model calls ask_user with two options and the run suspends. The user answers the next morning; the runtime restores the checkpoint and the model calls submit_report with dry_run set to true, gets a plan totalling the lines with one policy flag, shows it, and on confirmation calls again with dry_run false and the confirm token. A network retry of that final call returns the same report id because of the idempotency key.

Every call passed through the same validation and middleware, and no tool body contains audit, scope or caching code.

Failure modes

  • Blocking jobs. A long call inside the loop times out and is retried, duplicating work. Use the job shape.
  • Optional dry runs. If dry_run defaults to false, models skip it. Make it required, or split plan and commit into two tools.
  • Tenant-blind caches. Cached query results cross tenants. Key on identity.
  • Exceptions as control flow. A raised exception ends the run; a returned error lets the model recover.
  • Unbounded sub-agents. An agent-as-tool loops until the budget of the whole run is gone. Give it its own limits.
  • Trusting annotations. A server marks a destructive tool read-only, and a client auto-approves it.
  • Too many tools at once. Past a few dozen, selection accuracy drops; retrieve a relevant subset per turn, as in tool selection.

Trade-offs

Fine-grained tools give precise control but cost turns and tokens; coarse tools are cheap but hide decisions the model cannot see. Two-call commands are safer and slower. Agent-as-tool simplifies the parent but adds cost and latency. Suspension needs checkpoint storage. Choose per tool, by the cost of being wrong.

What to do next

  1. List your agent's tools and label each with one of the six shapes.
  2. Build or adopt a neutral registry with schema validation that returns errors as results, and thin adapters per model provider.
  3. Move audit, scope checks and caching into middleware, ordered audit, authorisation, cache.
  4. Make every command take a required dry_run, or split it into plan and commit, and attach idempotency keys.
  5. Convert any call that can exceed a few seconds into a job with a poll hint, and resume runs on completion.
  6. Give agent-as-tool calls their own step, token and time budgets and structured results.
  7. If you serve tools over MCP, set annotations honestly and enforce authorisation on the server regardless.
Key takeaway: The tools pattern is simple: the model proposes, the runtime validates, decides and executes. What makes it robust is recognising that tools come in shapes, queries, commands, jobs, handles, sub-agents and humans, each with its own contract for retries, caching, approval and time. Encode the shape in your registry, put cross-cutting behaviour in middleware, make side effects take two deliberate steps, never block the loop on slow work, and treat every hint and every argument from outside your runtime as untrusted.