An ADK Java agent can use tools served by any Model Context Protocol (MCP) server. The client side of that is covered in MCP integration in ADK Java: transports, timeouts, filters, credentials and callbacks. This article covers the other half, the part your team controls when it owns the tools: how to build an MCP server in Java that an ADK agent uses well, and how to test the pair before a model ever sees it.

That half matters more than it looks. ADK passes your tool names, descriptions and schemas to the model unchanged, and turns your results into the map the model reads. A vague description, a loose schema or a result in the wrong content type degrades every agent that connects, and no client-side setting can repair it. The code targets ADK Java 1.11.0, which depends on version 2.0.0 of the MCP Java SDK; check both versions before copying, since both libraries move quickly.

The contract ADK imposes

What crosses the boundary: your server's metadata and results become the model's viewLlmAgentADK Java 1.11McpToolset / McpToolfilter, declaration, wrapYour MCP server (Java SDK 2.0)Tool: name, descriptioninputSchema (JSON Schema)callHandler(exchange, request)CallToolResult: text JSON, isErrorstdout = protocol onlylogs go to stderrtools/listname, description, schematools/call (retried on error)What ADK hands the modeldeclaration: parametersJsonSchema = inputSchemaresult: {text_output: [parsed JSON]} or {error: ...}structuredContent and images are not passed
The contract between an ADK agent and an MCP server you own.

Start from what ADK actually does with your server, read from the ADK Java source. Design every tool against these rules rather than against the MCP specification in general.

  • Declarations are pass-through. AbstractMcpTool builds a FunctionDeclaration whose name and description are your tool's, whose parametersJsonSchema is your inputSchema, and whose responseJsonSchema is your outputSchema if you set one. Nothing is renamed, shortened or linted.
  • Results are converted to a map. If isError is true, the model gets an error entry built from your text. Otherwise each text block is parsed as JSON if it parses, or wrapped as {"text": ...}, and the list goes under text_output.
  • Only text survives. Images, audio and embedded resources are not passed to the model, and structuredContent is not read. Put your JSON in a text block.
  • Calls are retried. McpTool.runAsync retries a failed tools/call up to three times, 100 ms apart, reinitialising the session each time. A call that timed out on the client may have completed on the server, so the server will see it again.

The last rule is the most expensive to learn in production. It turns every side-effecting tool into a tool that must be idempotent, which is a server design decision, not a client setting.

Designing tools for a model

A model chooses tools by reading text and fills arguments by reading schemas, so treat both as the interface.

  • Names are verbs on nouns: get_invoice, list_invoices, refund_invoice. Avoid names that collide with common local tools, because two tools with one name in one agent is an error.
  • Descriptions say when, not just what: "Fetch one invoice by id. Use list_invoices first if you only have a customer id." The second sentence prevents the guessing loop where a model invents ids.
  • Schemas are tight: required fields listed, additionalProperties false, enum for closed sets, pattern for id formats, and a description on every property. Amounts in integer minor units, never floats.
  • Tools are coarse enough to finish a step: one call that returns an invoice with its line items beats three calls the model must chain, because every extra call is latency, tokens and another chance to go wrong.
  • Outputs are small: return the fields the task needs, cap lists and say so ("truncated": true), because the whole result lands in the prompt.

Building the server with the MCP Java SDK

Here is a billing server with three tools, using the SDK's synchronous API over stdio. The tool schema is a JSON string parsed by the SDK's mapper; Tool.builder(name, jsonMapper, inputSchema) accepts exactly that.

import io.modelcontextprotocol.json.McpJsonDefaults;
import io.modelcontextprotocol.json.McpJsonMapper;
import io.modelcontextprotocol.server.McpServer;
import io.modelcontextprotocol.server.McpServerFeatures.SyncToolSpecification;
import io.modelcontextprotocol.server.McpSyncServer;
import io.modelcontextprotocol.server.transport.StdioServerTransportProvider;
import io.modelcontextprotocol.spec.McpSchema.ServerCapabilities;
import io.modelcontextprotocol.spec.McpSchema.Tool;

public final class BillingToolsServer {
  public static void main(String[] args) {
    McpJsonMapper json = McpJsonDefaults.getMapper();
    BillingHandlers h = new BillingHandlers(InvoiceStore.fromEnvironment());

    Tool getInvoice = Tool.builder("get_invoice", json, """
        {"type": "object",
         "properties": {"invoice_id": {"type": "string", "pattern": "^INV-[0-9]{8}$",
                                       "description": "Invoice id, for example INV-00012345"}},
         "required": ["invoice_id"], "additionalProperties": false}""")
        .description("Fetch one invoice by id: status, currency, total and line items in minor units. "
            + "If you only have a customer id, call list_invoices first.")
        .build();

    Tool refundInvoice = Tool.builder("refund_invoice", json, """
        {"type": "object",
         "properties": {
           "invoice_id":   {"type": "string", "pattern": "^INV-[0-9]{8}$"},
           "amount_minor": {"type": "integer", "minimum": 1,
                            "description": "Refund amount in minor units, at most the unrefunded total"},
           "ticket_id":    {"type": "string", "description": "Support ticket that approved this refund"}},
         "required": ["invoice_id", "amount_minor", "ticket_id"], "additionalProperties": false}""")
        .description("Refund part or all of an invoice. Requires an approved support ticket. "
            + "Repeating the same call returns the original refund instead of refunding twice.")
        .build();

    McpSyncServer server = McpServer.sync(new StdioServerTransportProvider(json))
        .serverInfo("billing-tools", "1.4.0")
        .capabilities(ServerCapabilities.builder().tools(false).build()) // static list
        .tools(
            SyncToolSpecification.builder().tool(getInvoice)
                .callHandler((exchange, req) -> h.getInvoice(req.arguments())).build(),
            SyncToolSpecification.builder().tool(refundInvoice)
                .callHandler((exchange, req) -> h.refundInvoice(req.arguments())).build())
        .build();
    // list_invoices is registered the same way; omitted for length.
  }
}

The capability flag passed to tools(...) advertises whether the tool list can change at run time. Keep the list static per release: an agent's prompt, evaluations and allowlists are all written against a fixed set, and ADK sends tools/list on every model call, so a changing list changes prompts mid-conversation.

The handlers are plain Java, which is what makes them testable without any protocol in the way:

final class BillingHandlers {
  private static final ObjectMapper JSON = new ObjectMapper();   // Jackson
  private final InvoiceStore store;
  BillingHandlers(InvoiceStore store) { this.store = store; }

  CallToolResult getInvoice(Map<String, Object> args) {
    String id = (String) args.get("invoice_id");
    return store.find(id)
        .map(inv -> ok(InvoiceView.of(inv)))                 // small, task-shaped view
        .orElseGet(() -> error("No invoice " + id
            + ". Call list_invoices with the customer id to get valid ids."));
  }

  CallToolResult refundInvoice(Map<String, Object> args) {
    String invoiceId = (String) args.get("invoice_id");
    long amount = ((Number) args.get("amount_minor")).longValue();
    String ticket = (String) args.get("ticket_id");
    String key = sha256(invoiceId + "|" + amount + "|" + ticket);   // same args => same key
    try {
      RefundOutcome r = store.refundOnce(key, invoiceId, amount, ticket);
      return ok(Map.of("refund_id", r.refundId(), "status", r.status(), "replayed", r.replayed()));
    } catch (RefundExceedsBalance e) {
      return error("Refund of " + amount + " exceeds the unrefunded balance " + e.balance()
          + ". Ask the user to confirm a smaller amount.");
    }
  }

  private static CallToolResult ok(Object body) {
    try {
      return CallToolResult.builder().addTextContent(JSON.writeValueAsString(body)).isError(false).build();
    } catch (JsonProcessingException e) {
      return error("Internal serialisation error");
    }
  }

  private static CallToolResult error(String message) {
    return CallToolResult.builder().addTextContent(message).isError(true).build();
  }
}

Notice what the error messages do: each one tells the model the next valid move. ADK forwards the first text block of an error, so that sentence is the model's only guidance for recovering. "Not found" produces a retry with an invented id; "call list_invoices with the customer id" produces the right next call.

Idempotency under retries

Because ADK retries tools/call on any error, including a client timeout after the server has committed, refund_invoice derives an idempotency key from its arguments. A retry sends identical arguments, so it gets the identical key. The store enforces it with a unique constraint:

CREATE TABLE refunds (
  idempotency_key text PRIMARY KEY,
  invoice_id      text   NOT NULL,
  amount_minor    bigint NOT NULL,
  ticket_id       text   NOT NULL,
  refund_id       text   NOT NULL,
  created_at      timestamptz NOT NULL DEFAULT now()
);
-- refundOnce: in one transaction
--   INSERT ... ON CONFLICT (idempotency_key) DO NOTHING RETURNING refund_id;
--   if nothing returned: SELECT refund_id FROM refunds WHERE idempotency_key = $1  -> replayed = true
--   else: call the payment provider with the same key, then commit

Pass the same key to the payment provider too, if it supports idempotency keys, so the guarantee holds end to end. Including ticket_id keeps two genuinely separate refunds of the same amount distinct, while still collapsing retries of one decision. Read-only tools need none of this, which is one more reason to keep reads and writes in separate tools.

Stdout belongs to the protocol

Over stdio, the server's standard output is the protocol: newline-delimited JSON-RPC, with the SDK sending diagnostics to its logger. A single stray line on stdout, such as a framework banner, a System.out.println left in for debugging, or a logging library defaulting to console output, corrupts the stream, and the client sees a parse failure or a hung initialise. Route all logging to stderr explicitly:

<!-- src/main/resources/logback.xml -->
<configuration>
  <appender name="STDERR" class="ch.qos.logback.core.ConsoleAppender">
    <target>System.err</target>
    <encoder><pattern>%d{HH:mm:ss.SSS} %-5level %logger{24} %msg%n</pattern></encoder>
  </appender>
  <root level="INFO"><appender-ref ref="STDERR"/></root>
</configuration>

Then add a build check that fails if System.out appears in server code. If you embed the server in a framework that prints a startup banner, turn the banner off; or run that server over Streamable HTTP instead, where stdout is just a log.

Connecting the agent

On the agent side, launch the packaged server as a child process and allowlist its tools:

StdioServerParameters billingParams = StdioServerParameters.builder()
    .command("java")
    .args(List.of("-jar", "/opt/tools/billing-tools-1.4.0.jar"))
    .build();

McpToolset billing = new McpToolset(
    billingParams.toServerParameters(), JsonBaseModel.getMapper(),
    List.of("get_invoice", "list_invoices", "refund_invoice"));

LlmAgent agent = LlmAgent.builder()
    .name("billing_support")
    .model("gemini-flash-latest")
    .instruction("Help customers with invoices. Refund only when a support ticket approves it.")
    .tools(billing)
    .build();

Pin the jar version in the path, exactly as you would pin an npm package, so a redeploy of the tools cannot change an agent without review. Keep secrets out of the command line, where process listings expose them. Approval for refund_invoice belongs in a before-tool callback, as described in ADK Java callbacks; the server-side ticket check is the second lock, not the only one.

Testing the pair

Test at three levels, cheapest first.

  1. Handler unit tests: call BillingHandlers directly with argument maps. Assert JSON shape, error text and, for refunds, that calling twice with the same arguments returns replayed: true and one row.
  2. Protocol contract test: start the real jar over stdio with the SDK client, list tools and compare against a committed golden file. A schema change then fails CI instead of surprising a model.
  3. Agent evaluation: run scripted conversations through the agent and score tool choice and arguments, as in the ADK Java eval framework.
@Test
void catalogMatchesGoldenFile() throws Exception {
  McpJsonMapper json = McpJsonDefaults.getMapper();
  ServerParameters params = ServerParameters.builder("java")
      .args("-jar", System.getProperty("billing.jar")).build();
  McpSyncClient client = McpClient.sync(new StdioClientTransport(params, json))
      .requestTimeout(Duration.ofSeconds(10)).build();
  try {
    client.initialize();
    Map<String, Tool> tools = client.listTools().tools().stream()
        .collect(Collectors.toMap(Tool::name, t -> t));
    assertEquals(Set.of("get_invoice", "list_invoices", "refund_invoice"), tools.keySet());
    for (Tool t : tools.values()) {
      assertEquals(Golden.schemaHash(t.name()), sha256(canonicalJson(t.inputSchema())), t.name());
    }
    CallToolResult missing = client.callTool(
        new CallToolRequest("get_invoice", Map.of("invoice_id", "INV-00000000")));
    assertTrue(missing.isError());
  } finally {
    client.closeGracefully();
  }
}

This test also catches the stdout bug: a banner on stdout makes initialize() fail, which is far better found here than in an agent's first production turn.

Worked example: a refund that times out

A customer writes that they were charged twice for invoice INV-00012345 and support ticket T-8812 approved a refund of 4,999 minor units. The agent calls get_invoice; the server returns one JSON text block, which ADK parses into text_output. The model sees a balance of 9,998 and calls refund_invoice with the invoice, 4,999 and T-8812. The server commits the refund, but its process is killed before the response is flushed to the pipe. ADK sees the call fail, reinitialises the session, which starts a fresh server process, and retries with the same arguments. The server computes the same key, finds the row, and returns the original refund_id with replayed: true. The customer is refunded once, and the agent can truthfully say so.

Without the key, the same sequence refunds twice, and every log line looks normal: two successful tool calls, two refund ids.

Failure modes

SymptomCauseFix
Initialise hangs or fails to parseSomething writes to stdoutLogs to stderr; banner off; contract test
Model sees nothing useful from a toolResult returned as image, resource or only structuredContentReturn JSON in a text block
Model invents ids and loopsError text gives no next stepErrors name the tool or input to use next
Duplicate refund after a timeoutADK retried a call the server completedArgument-derived idempotency key with a unique constraint
Arguments in wrong units or shapeLoose schema: floats, no pattern, extra propertiesIntegers in minor units, patterns, additionalProperties false
Agent behaviour changes after a tools deployUnpinned jar or schema changeVersion the jar path; golden-file schema test

Trade-offs

Stdio or HTTP for your own server? Stdio is simplest when one agent host uses the tools, but each toolset is a JVM process, and the stdout rule is strict. Streamable HTTP suits shared servers with their own scaling, auth and deploys, at the cost of a network hop on every list and call. Fine-grained or coarse tools? Fine-grained tools compose flexibly but cost calls and invite wrong chains; coarse tools finish steps in one call but grow schemas. An MCP server or a local Java FunctionTool? Choose MCP when other agents or clients will reuse the tools or another team owns the system; choose local tools for logic private to one agent. For keeping names unique across both, see tool registry design.

What to do next

  1. Check your ADK Java and MCP Java SDK versions, and read AbstractMcpTool in the version you run.
  2. Rewrite each tool description to say when to use it and what to call first.
  3. Tighten every input schema: required fields, patterns, enums, integer amounts, no extra properties.
  4. Make every result a JSON text block and every error a sentence naming the next valid step.
  5. Add argument-derived idempotency keys and unique constraints to every tool with side effects.
  6. Route logging to stderr and add a build check against System.out.
  7. Add the stdio contract test with a golden schema file to CI.
  8. Pin the server artefact version in each agent's toolset configuration.
Key takeaway: ADK Java passes an MCP server's names, descriptions and schemas straight to the model, keeps only text results, and retries failed calls. So the server decides tool quality: write descriptions that say when to use a tool, make schemas tight, return JSON in text blocks with errors that name the next step, make side effects idempotent with argument-derived keys, keep stdout for the protocol, and lock the catalog with a contract test.