A model that "uses tools" never executes anything. It emits a structured request that says, in effect, "call search_orders with these arguments". Everything after that, deciding whether the call is allowed, running it, bounding how long it takes, cutting its output down to something the model can read, and recording what happened, is ordinary software that you write. That software is the tool runtime, and it decides most of what users experience: whether the agent is fast, whether it can be tricked into deleting data, and whether it drowns in its own tool output.

This article is about that runtime. The protocol and the craft of writing tool definitions are covered in function calling and tool use, and deadlines, idempotency and compensation for side effects are covered in tool calling reliability. Here we build the layer in between: a registry, a selector, a validation and policy gate, a dispatcher, executors, a result shaper and a tracer, with working code and the failure modes of each.

Advertisement

The loop, and who owns each step

Every tool-using agent runs the same loop. The runtime sends the model the conversation plus a list of tool definitions. The model replies with text, one or more tool calls, or both. Each call carries a name, a JSON arguments object and an id. The runtime executes the calls and appends one result per call, tagged with the same id, and sends everything back. The loop ends when the model replies without calls or when the runtime stops it.

The model owns one decision: which call to propose next. The runtime owns everything else and should treat the model as a capable but untrusted client that can invent tool names, repeat failed calls, or follow instructions it read in a web page.

The tool runtime: the model proposes calls, the runtime decides and executesModelemits tool callsTranscriptcalls + results by idTool selectorwhich tools this turnRegistryname, schema, riskValidate + policyallow, confirm, denyDispatcherparallel reads, serial writesExecutorssandbox, scoped credsResult shapertruncate, page, redactTracerone span per calldefinitionstool listcallsapprovedraw outputshaped resultsnext turnEvery arrow into the model is budgeted in tokens; every arrow out of it is untrusted input.
Components of a tool runtime. The model only proposes; selection, policy, execution and shaping are code you control.

The registry: one source of truth per tool

The registry maps a tool name to everything the runtime needs: the argument schema, the function, a risk class, a timeout and a result size limit. Keeping these together matters because the same metadata drives several layers. The schema feeds both the definition the model sees and the validator. The risk class feeds the policy gate and the dispatcher. The size limit feeds the shaper. If the risk class lives in a separate config file, sooner or later someone adds a write tool and forgets to classify it.

Tools can come from three places: functions in your own process, remote MCP servers, and other agents exposed as tools. Treat them uniformly in the registry but record the origin. A remote MCP tool advertises hints such as read-only or destructive through tool annotations, but those hints come from the server, so use them as defaults and let your own classification win. Namespacing by origin, for example github.create_issue, prevents two servers from shadowing each other's names.

Advertisement

Selection: which tools the model sees this turn

Every definition you send costs input tokens on every turn, and accuracy drops as the list grows because the model has more near-miss options to confuse. A runtime with five tools can send all of them. A runtime with two hundred cannot. There are three common strategies, and larger systems combine them.

  • Static profiles. Each agent role gets a fixed subset: the billing agent sees billing tools. This is simple, cacheable and easy to audit.
  • Retrieval. Embed tool descriptions, retrieve the top few for the current request, and always include a small core set. This scales, but a retrieval miss means the model cannot even try the right tool, so log what was offered next to what was needed.
  • Discovery tools. Expose a find_tools(query) tool that returns definitions, and add the chosen ones to the next turn. The model does its own retrieval, at the cost of an extra round trip.

Whichever you choose, keep the list stable within a conversation. Changing definitions mid-conversation breaks prefix caching and confuses a model that planned around a tool that then vanishes.

The gate: validate, then decide

Validation comes first and is mechanical. Check the arguments against the JSON Schema, then apply semantic checks the schema cannot express, such as that a date range is not reversed or that an account id belongs to the current tenant. On failure, return a result with the error flag set and a message that tells the model what to fix, rather than raising into the loop.

Policy comes second and returns one of three verdicts: allow, confirm (ask the human) or deny. The inputs are the tool's risk class, the user's grants, the arguments, and the state of the conversation. The last input is the one teams forget. Once the agent has read untrusted content, such as a fetched web page, an inbound email or a file from a shared drive, any instruction in that content may now be steering the model. Marking the turn as tainted and requiring confirmation for writes after that point is a cheap, effective defence against prompt injection turning into action.

async def policy(tool, args, ctx):
    if tool.name not in ctx.granted_tools:                  # per-user, per-tenant grants
        return "deny", "tool not granted to this user"
    if tool.risk == "destructive":
        return "confirm", "destructive"
    if tool.risk == "write" and ctx.tainted:                # untrusted text entered this turn
        return "confirm", "write after reading untrusted content"
    if tool.name == "send_email" and not args["to"].endswith("@" + ctx.org_domain):
        return "confirm", "external recipient"
    return "allow", ""

Keep policy in code. A system prompt saying "never email external addresses" is a request; a function that returns confirm is enforcement.

Dispatch: parallel calls and result ids

Models can return several calls in one turn, for example three lookups for three customers. Running them concurrently cuts latency to roughly the slowest call instead of the sum. Two rules keep it safe. First, run only read-only calls in parallel. Calls with side effects run one at a time in the order the model gave, because the model may have assumed that order: "create the folder, then upload the file". Second, return the results matched by call id and in the original order. Some providers reject a turn whose results do not cover every call id, and every model reasons better when result three follows call three.

Bound concurrency with a semaphore, plus a limit per downstream service so one agent cannot exhaust a connection pool. Give every call a timeout, and turn a timeout into an error result that says the outcome is unknown, never a silent success.

Shaping results to fit the context budget

The most common way a tool-using agent degrades is a tool that returns 60,000 characters of JSON, burying the instructions, raising cost on every later turn and sometimes overflowing the context. The shaper enforces a budget per tool.

  • Project, don't dump. Return the fields the task needs. A customer lookup should return name, status and plan, not the full CRM record.
  • Paginate explicitly. Return the first page plus a cursor, and say in the result how to get the next page. The model will ask if it needs more.
  • Truncate with a note. If you must cut, say what was cut and how to narrow the request. Silent truncation leads the model to conclude that data does not exist.
  • Redact. Strip secrets, tokens and personal data the task does not require before they enter the transcript, because the transcript is logged, cached and sometimes shown to users.
  • Mark errors as errors. An error string returned as a normal result reads as a successful answer. Use the error flag so the model knows to recover.

Watch the cumulative budget too: replace results older than a few turns with a one-line summary and a reference id.

Worked example: the runtime in code

The runtime below puts the layers together. It looks the tool up, validates, applies policy, runs reads in parallel and writes in order, maps every failure to an error result with a message aimed at the model, and shapes the output. Unexpected exceptions are recorded on the trace with full detail but reach the model only as a generic message, because stack traces leak internals and invite the model to "debug" your server.

import asyncio, json
from dataclasses import dataclass
from typing import Any, Awaitable, Callable

class UserFacingError(Exception):
    """Raised by tools when the model can fix the problem with different arguments."""

@dataclass
class Tool:
    name: str
    schema: dict                               # JSON Schema for the arguments
    fn: Callable[..., Awaitable[Any]]
    risk: str = "read"                         # read | write | destructive
    timeout_s: float = 20.0
    max_chars: int = 8000

@dataclass
class Call:
    id: str
    name: str
    args: dict

def result(call, content, is_error=False):
    return {"call_id": call.id, "is_error": is_error, "content": content}

class Runtime:
    def __init__(self, tools, validate, policy, tracer, max_parallel=4):
        self.tools = {t.name: t for t in tools}
        self.validate, self.policy, self.tracer = validate, policy, tracer
        self.sem = asyncio.Semaphore(max_parallel)

    async def run_one(self, call, ctx):
        tool = self.tools.get(call.name)
        if tool is None:
            return result(call, f"Unknown tool {call.name!r}. Available: {sorted(self.tools)}", True)
        problems = self.validate(tool.schema, call.args)
        if problems:
            return result(call, "Invalid arguments: " + "; ".join(problems[:5]), True)
        verdict, reason = await self.policy(tool, call.args, ctx)
        if verdict == "deny":
            return result(call, f"Not permitted: {reason}", True)
        if verdict == "confirm" and not await ctx.ask_user(tool.name, call.args):
            return result(call, "The user declined this action. Do not retry it.", True)
        async with self.sem:
            with self.tracer.span("tool", name=call.name, call_id=call.id) as span:
                try:
                    out = await asyncio.wait_for(tool.fn(ctx=ctx, **call.args), tool.timeout_s)
                except asyncio.TimeoutError:
                    return result(call, f"Timed out after {tool.timeout_s}s; outcome unknown.", True)
                except UserFacingError as e:
                    return result(call, str(e), True)
                except Exception as e:
                    span.record_exception(e)             # full detail goes to the trace only
                    return result(call, f"{call.name} failed with an internal error.", True)
        return result(call, shape(out, tool.max_chars))

    async def run_batch(self, calls, ctx):
        is_read = lambda c: getattr(self.tools.get(c.name), "risk", "read") == "read"
        reads = [c for c in calls if is_read(c)]
        rest = [c for c in calls if not is_read(c)]
        done = list(await asyncio.gather(*(self.run_one(c, ctx) for c in reads)))
        for c in rest:                                    # side effects: one at a time, in order
            done.append(await self.run_one(c, ctx))
        order = {c.id: i for i, c in enumerate(calls)}
        return sorted(done, key=lambda r: order[r["call_id"]])

def shape(out, limit):
    text = out if isinstance(out, str) else json.dumps(out, default=str)
    if len(text) <= limit:
        return text
    return (text[: limit - 160] + f"\n[truncated: showing {limit - 160} of {len(text)} chars;"
            " narrow the filter or request the next page]")

Walk one turn through it. A support agent receives "Refund order 1182 and tell the customer". The model returns get_order(order_id=1182) and refund_order(order_id=1182, amount=49.00). The read runs immediately. The refund is a write, so it runs afterwards, once policy allows it: the user holds the refunds grant and the turn is not tainted. Both results go back in call order. Next turn the model proposes send_email to an external address, policy returns confirm, and the human sees the draft first. All of it sits in one trace.

Executors, sandboxes and credentials

Where a tool runs matters as much as whether it may run. In-process functions are fastest but share the agent's memory, credentials and failure domain. Tools that execute model-written code, such as a Python interpreter or shell, belong in a sandbox: a container or microVM with no ambient credentials, a read-only base image, CPU, memory and wall-clock limits, and egress restricted to an allowlist. Remote tools, including MCP tools, run in another process by design, which gives isolation but adds a network hop and a server you must trust.

Credentials should belong to the user, not the agent. The executor attaches a token scoped to the current user and task, and the model never sees it. With one powerful service account, every successful injection inherits that power.

Observability and failure modes

Trace every model turn and every tool call as spans in one trace, recording tool name, argument hash, verdict, latency, result size, truncation and error class. The metrics that reveal trouble early are calls per task, error rate per tool, repeated identical calls, and the share of context filled by tool results.

SymptomLikely causeFix
Same call repeated five timesError result was vague, or the error flag was not setActionable error text; cap identical calls per task
Model invents a tool nameToo many or overlapping tools offeredSmaller profile; return the valid names in the error
Answers get worse late in long tasksTool output crowding the contextPer-tool size limits; summarise old results
Unexpected write after browsingInjected instructions in fetched contentTaint tracking; confirm writes after untrusted reads
Provider rejects the turnMissing or reordered result idsAlways return one result per call id, in order
Slow turns with many callsSerial execution of independent readsParallel reads behind a semaphore

Trade-offs

Each layer trades autonomy against control. Confirmations stop injected actions but annoy users if they fire constantly, so reserve them for destructive, external or tainted writes. Truncation saves tokens but can hide the row that mattered, so pair it with pagination. Retrieval-based selection scales but adds a failure mode. For the model-side view of the same loop, see LLM tool use.

What to do next

  1. Put name, schema, function, risk class, timeout and size limit for every tool into one registry.
  2. Validate arguments against the schema before anything runs, and return fixable errors as error results.
  3. Write a policy function with allow, confirm and deny, and add taint tracking for untrusted content.
  4. Run read-only calls in parallel behind a semaphore; run writes serially in the model's order.
  5. Set a size budget per tool, paginate large results and say so whenever output is truncated.
  6. Move code execution into a sandbox and switch tools to per-user scoped credentials.
  7. Trace each call and alert on repeated identical calls and on per-tool error rates.
Key takeaway: Tool use is a request from the model and an action taken by your runtime. Build that runtime as a pipeline you control: a registry that holds each tool's schema, risk and limits; a selector that keeps the offered list small; a gate that validates arguments and applies policy, including taint after untrusted reads; a dispatcher that runs reads in parallel and writes in order and returns results matched by id; a shaper that keeps output within budget and marks errors as errors; sandboxed executors with user-scoped credentials; and one trace across all of it. Most agent failures in production are failures of this layer, and all of them are fixable in ordinary code.