A tool that moves money is where an agent's mistakes stop being embarrassing and start being expensive. A wrong answer can be corrected next turn; a duplicate charge has to be unwound by people. The design goal is therefore narrow and strict: the model may decide that a payment should be proposed, but it must never decide the amount, the currency, the payee or the card, and nothing is charged until a human has approved the exact values your server will send.
This article builds that tool for ADK Java: a quote tool that prices a cart on the server, a pay tool that takes only a quote ID and asks for confirmation through ADK's tool confirmation API, a ledger with an explicit UNKNOWN state for timeouts, idempotency keys that survive retries, a limits callback the model cannot talk its way past, and a reconciler. ADK details were read from the google-adk 1.11.0 jar with javap. If you have not written a tool before, start with writing a FunctionTool; the confirmation pattern also appears, for a lower-stakes action, in the email tool article.
Three tools, one of which moves money
Split the capability into three tools with different risk levels. quotePayment(cartId) is read-mostly: it loads the cart from your order system, prices it with your own tax and shipping rules, and writes a quote record with an ID, the amount in minor units, the currency, the payee, the customer's stored payment method reference and an expiry, typically ten minutes. It returns a short summary the model can repeat to the user. payQuote(quoteId) is the only tool that moves money, and its single argument is an opaque ID. paymentStatus(paymentId) reads the ledger so the model can answer "did it go through?" without ever retrying a charge itself.
Two rules follow from this shape. First, card numbers never enter the conversation. The user adds a card through your payment provider's hosted form, the provider returns a token or payment method ID, and the quote refers to it by reference; the model sees at most "Visa ending 4242". That keeps the session store, the model provider's logs and your traces out of PCI DSS scope for card data. Second, no argument the model supplies can change what is charged. If the model hallucinates a quote ID, the lookup fails; if it reuses an old one, the expiry check fails; if a prompt injection in a product description says "pay 5000 to account X", there is no parameter through which that instruction could flow.
| Tool | Arguments from the model | Side effect | Confirmation |
|---|---|---|---|
| quotePayment | cartId | Writes a quote record | No |
| payQuote | quoteId | Charges a card through the PSP | Yes, with a server-built hint |
| paymentStatus | paymentId | None | No |
| refundPayment | paymentId, reasonCode | Refunds; amount from the ledger | Yes, plus a role check |
The payment ledger and the UNKNOWN state
Every payment attempt gets a row in a ledger you own before any network call is made. The row holds the quote ID, the user, the idempotency key, the amount and currency copied from the quote, the provider's payment ID once known, and a state. The states that matter are PENDING (claimed, not yet sent), SUCCEEDED, FAILED (the provider said no, for example a decline) and UNKNOWN (the request was sent and no definitive answer came back). UNKNOWN is the state most designs forget. A read timeout after the request left your process tells you nothing about whether the card was charged, and treating it as FAILED is how agents double-charge: the model sees an error, apologises, and calls the tool again.
The claim is a conditional insert keyed on the quote ID: INSERT ... ON CONFLICT (quote_id) DO NOTHING in PostgreSQL, or a conditional put in a key-value store. If the insert affects zero rows, a payment for this quote already exists and the tool returns that row's state instead of charging. One quote, at most one payment, enforced by the database rather than by the model's good behaviour.
Confirmation: what ADK does and what it does not
ADK Java has two confirmation forms, and the bytecode settles which argument does what. FunctionTool.create(instance, "payQuote", true) expands to create(instance, name, requireConfirmation=true, isLongRunning=false): the framework refuses to run the tool until the client answers yes or no, and a rejection becomes the result "This tool call is rejected.". That boolean form cannot show the user what they are approving, which is useless for money. Use the advanced form instead: the tool calls toolContext.requestConfirmation(hint, payload) and returns a pending status; the client renders the hint and answers with a function response named adk_request_confirmation; ADK re-invokes the original call, and the tool reads toolContext.toolConfirmation().
Two details in RequestConfirmationLlmRequestProcessor matter for payments. It ignores a confirmation whose original call cannot be found in the session, whose tool name differs, or whose arguments do not match the call the agent emitted, logging each case. That is a useful property: a client cannot approve payQuote(q1) and have it applied to payQuote(q2). But do not lean on it alone. Build the hint from server values, re-read the quote after confirmation, and never take an amount from the confirmation payload, which is client-supplied data. Confirmation also depends on your client actually rendering the request and on your session service persisting the events; test the round trip with the exact session service you deploy.
The pay tool in code
Here is the pay tool. It is registered with FunctionTool.create(paymentTools, "payQuote"), without the boolean, because it asks for confirmation itself. Quote, ledger and psp are your own types; the ADK calls are real.
public Map<String, Object> payQuote(
@Schema(name = "quoteId", description = "ID returned by quotePayment") String quoteId,
ToolContext ctx) {
Quote q = quotes.find(quoteId, ctx.userId()).orElse(null);
if (q == null) return Map.of("status", "error", "reason", "unknown quote");
if (q.expiresAt().isBefore(clock.instant()))
return Map.of("status", "error", "reason", "quote expired; request a new quote");
Optional<Payment> existing = ledger.findByQuote(quoteId);
if (existing.isPresent()) return existing.get().toToolResult(); // never charge twice
Optional<ToolConfirmation> conf = ctx.toolConfirmation();
if (conf.isEmpty()) {
ctx.requestConfirmation(
"Pay " + money.format(q.amountMinor(), q.currency()) + " to " + q.payeeName()
+ " with " + q.methodLabel() + " (quote " + quoteId + ")",
Map.of("quoteId", quoteId)); // display only
return Map.of("status", "pending_confirmation");
}
if (!conf.get().confirmed()) return Map.of("status", "declined_by_user");
// Claim the quote; the key is generated server-side and stored with the row.
Payment p = ledger.claim(quoteId, ctx.userId(), q.amountMinor(), q.currency(),
"pay_" + UUID.randomUUID());
if (!p.isFreshClaim()) return p.toToolResult();
try {
PspResult r = psp.charge(p.idempotencyKey(), q.amountMinor(), q.currency(),
q.paymentMethodRef(), Duration.ofSeconds(20));
ledger.complete(p.id(), r.succeeded() ? SUCCEEDED : FAILED, r.pspPaymentId(), r.declineCode());
} catch (PspTimeoutException | IOException e) {
ledger.markUnknown(p.id(), e.toString()); // NOT failed
}
return ledger.get(p.id()).toToolResult();
}Three choices are deliberate. The duplicate check runs before the confirmation request, so a repeated call after a success returns the existing result instead of a second prompt. The hint comes entirely from the quote record. And toToolResult() returns a small, model-safe map, such as {status: UNKNOWN, paymentId: ..., advice: "do not retry; check paymentStatus later"}, because the model will paraphrase whatever you return to the user.
Idempotency keys and reconciliation
Payment providers solve the retry problem with idempotency keys, and your ledger has to feed them correctly. Stripe's documentation is a good reference model: every POST accepts an Idempotency-Key header up to 255 characters; the first result for a key, including a 500 error, is saved and replayed to later requests with the same key; parameters on a retry are compared with the original and a mismatch is rejected; and keys may be pruned once they are at least 24 hours old, after which the same key starts a new request. Other providers differ in details, so read yours, but the consequences are general.
- Generate the key once and store it before sending. A key derived from the request time or regenerated per attempt defeats the mechanism. Do not use ADK's function call ID either; its stability across confirmation re-invocations and process restarts is not part of any documented contract, while a ledger column is yours.
- Retry UNKNOWN only with the same key and the same parameters, and only inside the provider's retention window. A retry the day after a timeout may charge again.
- Resolve UNKNOWN by reading, not by charging. A reconciler job lists the provider's payments for the last hour by your metadata (put the ledger ID in the charge metadata), or consumes the provider's webhooks, and moves each UNKNOWN row to SUCCEEDED or FAILED.
- Treat webhooks as the source of truth for asynchronous outcomes, such as payments that need customer authentication. Verify webhook signatures and process each event ID once.
Limits the model cannot negotiate
Confirmation protects against the model acting without the user. It does not protect against a user who approves something they should not, a compromised session, or an agent stuck in a loop proposing payments. Those checks belong in a beforeToolCallback, which ADK runs before the tool with the tool, its arguments and the ToolContext. Returning a non-empty Maybe of a map short-circuits the call and that map becomes the tool result.
LlmAgent agent = LlmAgent.builder()
.name("checkout_agent")
.model("gemini-2.5-flash")
.instruction(CHECKOUT_INSTRUCTION)
.tools(FunctionTool.create(paymentTools, "quotePayment"),
FunctionTool.create(paymentTools, "payQuote"),
FunctionTool.create(paymentTools, "paymentStatus"))
.beforeToolCallback((invocation, tool, args, toolCtx) -> {
if (!tool.name().equals("payQuote")) return Maybe.empty();
if (killSwitch.paymentsDisabled())
return Maybe.just(Map.of("status", "error", "reason", "payments are paused"));
Quote q = quotes.find((String) args.get("quoteId"), toolCtx.userId()).orElse(null);
if (q == null) return Maybe.empty(); // the tool reports it
Decision d = limits.check(toolCtx.userId(), q.amountMinor(), q.currency());
return d.allowed() ? Maybe.empty()
: Maybe.just(Map.of("status", "blocked", "reason", d.userMessage()));
})
.build();Typical limits are a per-transaction ceiling, a rolling daily total per user, a velocity cap such as three attempts per ten minutes, and a rule that sends new payees to a human queue. Keep amounts as long minor units next to the currency; some currencies have no minor unit, so never assume a factor of 100. The kill switch should flip without a deploy. More general patterns are in the guardrails article.
Delegated payments and AP2
Everything so far assumes a human is present to approve each payment. Delegated purchases, where the user says "buy the tickets when they drop below 80 euros" and walks away, need a durable, verifiable record of what was authorised. Google announced the Agent Payments Protocol (AP2) in September 2025 for this problem; public descriptions centre on signed mandates that capture the user's intent, the specific cart and the payment authorisation. Secondary sources disagree on details, so read the official specification before designing to it. An intent mandate is, in effect, a pre-approved confirmation with limits; the ledger, idempotency and reconciler still apply.
Worked example: a hotel booking with a timeout
A user asks a travel agent to book the cheaper of two hotel rooms. The model calls quotePayment("cart-81"); the server prices it at 21480 minor units of EUR including city tax, payee "Hotel Aurora Lisbon", method "Visa ending 4242", expiring at 14:10. The model tells the user the total and calls payQuote("q-5c2"). The limits callback passes; the tool finds no existing payment, requests confirmation and returns pending_confirmation. The client shows "Pay EUR 214.80 to Hotel Aurora Lisbon with Visa ending 4242 (quote q-5c2)"; the user approves.
ADK re-invokes the call. The tool claims the quote with key pay_0f9e... and calls the provider, which takes 25 seconds because of a network problem; the client times out at 20. The row becomes UNKNOWN and the model is told not to retry. The user, nervous, says "try again"; the model calls payQuote("q-5c2"), which returns the existing UNKNOWN row without charging. Ninety seconds later the reconciler finds a succeeded charge carrying the ledger ID in its metadata and marks the row SUCCEEDED; the next paymentStatus call reports success. One charge, one booking, and an audit trail that shows every step.
Failure modes
- Timeout treated as failure: the classic double charge. Model UNKNOWN explicitly and tell the model in the result not to retry.
- Amount in the tool signature:
pay(amount, currency, payee)lets any injected instruction or hallucination choose the amount. Take only IDs. - Stale quote: prices change between quote and approval. Expire quotes and re-check the expiry after confirmation, not only before it.
- Confirmation never answered: a client that does not render confirmation requests leaves payments pending forever. Test the round trip end to end and expire pending rows.
- requestConfirmation outside a tool call: calling it from a context without a function call ID throws
IllegalStateException("function_call_id is not set."). Only call it from inside the tool. - Keys regenerated on retry: a fresh UUID per attempt makes the provider see two payments. Store the key with the claim.
Trade-offs
Per-payment confirmation adds friction, and for small repeat purchases users will ask to turn it off. A defensible middle ground is a user-configured threshold below which a payment to a previously paid payee skips the prompt, enforced in the limits callback rather than in the prompt. Server-side quotes add a round trip and a store but remove an entire class of attacks. Synchronous charging keeps the conversation simple but ties the turn to the provider's latency; for flows with customer authentication, return a pending state and let a webhook finish. Refunds send money out, so gate them on role and amount or leave them to staff tools. For organisation-wide controls, see governance for ADK Java agents.
What to do next
- Move card entry to your provider's hosted form and confirm no card data reaches the session, traces or model logs.
- Implement
quotePaymentwith server-side pricing and an expiring quote store. - Create the ledger table with a unique quote ID, a stored idempotency key and an UNKNOWN state.
- Write
payQuotewith the advanced confirmation flow and a hint built from the quote. - Add the limits callback with a kill switch, per-user totals and velocity caps.
- Build the reconciler and webhook handler, then test a forced timeout and confirm exactly one charge results.