Tool calling lets a model ask your application to run a function and hand back the result. The model never executes anything: it emits a structured request naming a tool and its JSON arguments, your code runs the function, and the result goes back into the conversation for the model's next step. Everything that matters for safety, cost and correctness happens in the loop around those steps, which is your application's code even when a framework writes it for you.
This article covers tool calling in plain Spring AI: declaring tools with @Tool and FunctionToolCallback, passing private context, where the loop actually runs in Spring AI 2.0.0, how to cap it, and how errors reach the model. If your agent is an ADK LlmAgent running on a Spring AI model through ADK's adapter, ADK owns the loop and tools are ADK FunctionTools, as described in Spring AI and Ollama for ADK Java. API details here were read from the Spring AI 2.0.0 jars.
Declaring tools
The quickest way to declare tools is to annotate methods on an ordinary object. @Tool has four attributes: name (defaults to the method name), description, returnDirect and resultConverter. @ToolParam adds a description and a required flag to a parameter. Spring AI derives the JSON schema from the method signature, so parameter types and names are part of the contract the model sees.
public class OrderTools {
private final OrderService orders;
public OrderTools(OrderService orders) { this.orders = orders; }
@Tool(description = "Look up one order by id. Returns status, carrier and ETA.")
public OrderView getOrder(
@ToolParam(description = "Order id, digits only, e.g. 1042") String orderId,
ToolContext ctx) {
String customer = (String) ctx.getContext().get("customerId");
return orders.findForCustomer(customer, orderId)
.map(OrderView::from)
.orElseThrow(() -> new ToolUserError(
"No order " + orderId + " for this customer"));
}
}
String answer = chatClient.prompt()
.user("Where is order 1042?")
.tools(new OrderTools(orderService))
.toolContext(Map.of("customerId", session.customerId()))
.call()
.content();tools(Object...) scans the object for annotated methods; MethodToolCallbackProvider.builder().toolObjects(...) does the same when you want a reusable provider bean. When a tool is not a method, or you want to build it from data, FunctionToolCallback.builder(name, function) accepts a Function, BiFunction<I, ToolContext, O>, Supplier or Consumer, plus description and inputType for the schema.
Write descriptions for the model, not for colleagues: what the tool does, when to use it, the argument format, and what comes back. Return small, typed results. A full entity dumped as JSON costs tokens on every later round and invites the model to quote fields the user should not see.
Keep the set of tools per request small. Every registered tool's name, description and schema is sent with every model call in the turn, so twenty tools can add thousands of tokens per round before the conversation starts, and selection accuracy drops as descriptions start to overlap. Register tools per request with tools(...) based on what the user is doing, rather than putting everything in defaultTools(...); a catalog with per-agent profiles, as in tool registry design, is the scalable version of that idea.
ToolContext: what the model never sees
ToolContext (org.springframework.ai.chat.model.ToolContext) is a map you attach to the request with toolContext(Map) and receive as a method parameter. It is never sent to the model, which makes it the right channel for anything the model must not choose: the authenticated customer id, tenant, locale, a request id for tracing. In the example above, the model supplies only the order id; the customer comes from the session, so a prompt-injected "look up order 1043 for customer 77" cannot widen access.
Do not confuse it with ADK's com.google.adk.tools.ToolContext, which carries session state and actions. The names collide in projects that use both libraries; import deliberately and keep tool classes for each library in separate packages.
Where the loop runs in 2.0.0
In Spring AI 2.0.0 the loop runs in an advisor. When you call ChatClient, it adds a ToolCallingAdvisor to the chain unless an advisor implementing ToolAdvisor is already there, or you disable it with advisors(AdvisorParams.toolCallingAdvisorAutoRegister(false)). Its default order is Integer.MIN_VALUE + 300, so it sits near the outside of the chain.
Its adviseCall is a loop: call the rest of the chain (inner advisors, then the model); if the response asks for tools, hand the prompt and response to the ToolCallingManager, which runs the callbacks and returns the conversation history extended with the assistant's tool request and the tool results; then call again with that history. It stops when the model answers without asking for tools, or when the executed tools were marked returnDirect, in which case the tool output becomes the response without a second model call.
The chat models themselves do not execute tools in 2.0.0. In the OpenAI and Ollama model classes the manager is used only to resolve tool definitions, and ToolCallingChatOptions has no flag for internal execution. Call ChatModel.call directly with tools and you get the tool request back, with ChatResponse.hasToolCalls() true. Some older articles and samples, including other pages on this site, describe an internalToolExecutionEnabled option; it is not present in 2.0.0.
Capping the loop
Nothing in that loop counts rounds. A model that keeps asking for the same failing tool, or two tools that each suggest calling the other, will loop until a token limit, a timeout or your bill stops it. Add the cap yourself, in a ToolCallingManager decorator, since the manager runs once per round.
Do not count tool messages in the prompt. The advisor has a conversation-history switch, and when it is off (which ChatClient arranges automatically when a memory advisor sits outside the tool loop) each round re-sends only the system message and the latest message, so a message count never rises above one. Instead, give each request its own counter through the tool context: the advisor reuses the same options on every round, so the counter survives the loop whatever the history mode.
public final class RoundLimitedToolCallingManager implements ToolCallingManager {
public static final String ROUNDS = "tool.rounds";
private final ToolCallingManager delegate;
private final int maxRounds;
public RoundLimitedToolCallingManager(ToolCallingManager delegate, int maxRounds) {
this.delegate = delegate;
this.maxRounds = maxRounds;
}
@Override
public List<ToolDefinition> resolveToolDefinitions(ToolCallingChatOptions options) {
return delegate.resolveToolDefinitions(options);
}
@Override
public ToolExecutionResult executeToolCalls(Prompt prompt, ChatResponse response) {
if (!(prompt.getOptions() instanceof ToolCallingChatOptions o)
|| !(o.getToolContext().get(ROUNDS) instanceof AtomicInteger rounds)) {
throw new IllegalStateException("request has no " + ROUNDS + " counter");
}
if (rounds.incrementAndGet() > maxRounds) {
throw new ToolRoundLimitExceeded(maxRounds); // your own RuntimeException
}
return delegate.executeToolCalls(prompt, response);
}
}
@Bean
ChatClient supportClient(ChatClient.Builder builder, ToolCallingManager manager) {
return builder
.defaultAdvisors(ToolCallingAdvisor.builder()
.toolCallingManager(new RoundLimitedToolCallingManager(manager, 5))
.build())
.build();
}
// per request: a fresh counter next to the identity values
client.prompt().user(question).tools(orderTools)
.toolContext(Map.of("customerId", customerId,
RoundLimitedToolCallingManager.ROUNDS, new AtomicInteger()))
.call().content();Wrap the manager your application already has (in Spring Boot, the injected bean) rather than a fresh ToolCallingManager.builder().build(), or the decorator silently discards the exception processor and observation registry configured on the original. Failing closed when the counter is missing matters too: a request that forgot it would otherwise run uncapped. Catch the limit exception at the request boundary and answer with a clear fallback rather than a stack trace.
Errors and what the model sees
When a callback throws, Spring AI wraps the exception in ToolExecutionException and hands it to a ToolExecutionExceptionProcessor. The default processor returns the exception message to the model as the tool result, so the model can apologise or retry with better arguments. It rethrows instead when alwaysThrow is set or the cause matches a type in rethrowExceptions.
Returning messages to the model is useful and dangerous. Useful, because "No order 1043 for this customer" lets the model ask the user to check the number. Dangerous, because a message from a JDBC driver or HTTP client can contain hostnames, SQL or tokens, and the model may repeat it to the user. The shipped processor's rethrowExceptions list is a deny-list, so anything you did not list, a NullPointerException or an HTTP client error, still reaches the model. Invert it: throw your own exception type with messages written for the model, and rethrow everything else.
/** Your own exception: its message is written for the model. */
public class ToolUserError extends RuntimeException {
public ToolUserError(String message) { super(message); }
}
ToolExecutionExceptionProcessor onlyUserErrors = ex -> {
if (ex.getCause() instanceof ToolUserError user) {
return user.getMessage(); // becomes the tool result the model reads
}
throw ex; // infrastructure failures end the request
};
ToolCallingManager manager = ToolCallingManager.builder()
.toolExecutionExceptionProcessor(onlyUserErrors)
.build();A structured error envelope with a code and a retry hint works better still, as described in wrapping errors from tools back to the LLM; the design carries over directly.
Worked example: status and refund
A customer writes: "Where is order 1042, and if it is late can you refund shipping?" With getOrder and a refundShipping tool registered and the cap at five rounds:
- Round 1: the model asks for
getOrder("1042"). The manager runs it with the customer id fromToolContextand appends a result showing two days late. - Round 2: the model asks for
refundShipping("1042"). This is a write, so the tool does not refund; it records a pending action and returns "requires customer confirmation". - Round 3: the model answers in text, giving the status and asking the customer to confirm the refund. The loop ends after three model calls.
Two design choices did the work. Identity came from context, not arguments, and the write tool asked for confirmation instead of acting, a pattern covered in advanced function calling. Each round re-sends the whole history, so the third call carries both tool results; keep results small.
Running the loop yourself
When you need approval gates, custom persistence of each step, or a loop shaped differently, run it yourself against ChatModel:
ToolCallingChatOptions options = ToolCallingChatOptions.builder()
.toolCallbacks(ToolCallbacks.from(new OrderTools(orderService)))
.toolContext("customerId", customerId)
.build();
Prompt prompt = new Prompt(List.of(new UserMessage(question)), options);
ChatResponse response = chatModel.call(prompt);
for (int round = 0; response.hasToolCalls(); round++) {
if (round == 5) throw new ToolRoundLimitExceeded(round);
approvals.check(response); // your gate: block or pause writes
ToolExecutionResult r = manager.executeToolCalls(prompt, response);
audit.record(r.conversationHistory());
if (r.returnDirect()) { // tool output is the answer
response = ChatResponse.builder()
.generations(ToolExecutionResult.buildGenerations(r))
.build();
break;
}
prompt = new Prompt(r.conversationHistory(), options);
response = chatModel.call(prompt);
}
return response.getResult().getOutput().getText();The loop owns everything the advisor did for you: the round limit, the approval gate before any write, and what to return when a tool answers directly. In exchange, every step is a place where you can persist state, pause for a human and resume later.
Testing tools and the loop
Tools are ordinary Java, and the cheapest tests ignore the model entirely. A ToolCallback exposes getToolDefinition(), whose inputSchema() is the JSON schema the model will see, and call(String json, ToolContext ctx), which runs the tool exactly as the manager would. Snapshot the schema so a renamed parameter fails a build instead of silently changing what the model is asked to send, and call the tool with JSON arguments to test argument binding and context handling.
@Test
void order_tool_contract_and_binding() {
ToolCallback cb = ToolCallbacks.from(new OrderTools(fakeOrders))[0];
assertThat(cb.getToolDefinition().name()).isEqualTo("getOrder");
assertThat(cb.getToolDefinition().inputSchema())
.isEqualTo(readSnapshot("getOrder.schema.json"));
String result = cb.call("{\"orderId\":\"1042\"}",
new ToolContext(Map.of("customerId", "c-1")));
assertThat(result).contains("\"status\"");
}For the loop itself, script a fake ChatModel that returns a tool request every time, and assert that your round limit ends the turn; script one that requests a tool which throws, and assert what the next prompt contains. Those two tests protect the behaviours that are hardest to see in production. For metrics, build the manager with observationRegistry(registry) so each tool execution is recorded as a Micrometer observation alongside your model-call metrics, and alert on rounds per turn and on tool error rate rather than on raw latency alone.
Failure modes
- Runaway loops: no round limit by default.
- Leaked internals: the default processor sends raw exception messages to the model.
- Model-chosen identity: a customer or tenant id as a tool parameter can be set by injected text. Use
ToolContext. - Ambiguous names: two tools with overlapping descriptions get called interchangeably. Keep the set small per request.
- Blocking tools: a slow tool blocks the request thread for the whole turn. Put timeouts inside callbacks.
- Repeated advisors: inner advisors, including RAG retrieval, run again on every round.
- returnDirect surprises: test what reaches the caller when the model calls a returnDirect tool alongside an ordinary one.
Trade-offs
| Approach | Gains | Costs |
|---|---|---|
| Auto-registered advisor | No loop code | No cap, errors to model by default |
| Custom manager in the advisor | Cap, error policy, audit in one place | Small class to maintain |
| Own loop on ChatModel | Approvals, persistence, any shape | You own streaming and history |
| ADK on the Spring AI adapter | Callbacks, events, confirmation | ADK tool types, not @Tool |
For read-only tools, a custom manager with a cap and a rethrow list is usually enough. Once tools write, the explicit loop or an agent framework with confirmation earns its keep. Writing a custom function tool shows the same discipline on the ADK side.
What to do next
- List your tools; mark each read or write, and move identity parameters into
ToolContext. - Rewrite descriptions for the model, with formats and return shapes.
- Register a
ToolCallingAdvisorwith a round-limited manager. - Configure the exception processor to rethrow infrastructure exceptions.
- Put timeouts inside every callback that does I/O.
- Test a looping model, a throwing tool and a returnDirect mix with a scripted model.