When one agent calls another over the Agent2Agent (A2A) protocol, each side makes a trust decision before it acts. The client must decide whether the remote agent is who it claims to be, whether its advertised skills are real, and whether its answer is safe to act on. The server must decide who is calling, on whose behalf, and with what permissions. Teams often collapse all of this into one question, "is it authenticated?", and then discover that a perfectly authenticated agent can still return a reply that talks their own agent into doing something harmful.
This article separates the trust decision into five questions, shows which A2A features provide evidence for each, compares the trust models you can build from that evidence, and ends with a policy function, a worked example and a checklist. Field-level detail for cards lives in the Agent Card specification and the mechanics of tokens and delegation in A2A authentication and authorization; here the focus is how to decide.
Trust is five questions, not one
Each question needs different evidence, and a yes to one says nothing about the others.
- Endpoint: am I connected to the host I meant to reach? TLS and DNS answer this, and nothing more.
- Card: did the organisation I think of as this agent's owner publish the Agent Card I am reading, and is it unmodified? A signature over the card answers this; TLS alone only says which host served it.
- Caller: from the server's side, which client is calling? Mutual TLS, OAuth client credentials or another scheme answers this.
- Authority: on whose behalf, and limited to what? A delegated, audience-restricted, short-lived token answers this.
- Content: may I act on what the remote agent sent back? No cryptography answers this. The signed, authenticated partner can still be compromised, mistaken, or relaying text injected by a third party.
What A2A gives you, and what it leaves to you
The A2A specification supplies the evidence; it does not supply a policy. As read on 2026-10-01 at a2a-protocol.org, the relevant pieces are these. Agents publish an Agent Card, conventionally at https://{domain}/.well-known/agent-card.json, and discovery can also go through curated registries or direct configuration. A card can carry a signatures list of JWS signatures computed over a canonical form of the card, which lets a client check who issued it independently of who served it. A card declares securitySchemes drawn from familiar families: API keys, HTTP authentication, OAuth 2.0, OpenID Connect and mutual TLS. An agent can expose an authenticated extended card with more detail for callers who have already authenticated, and the discovery guidance recommends out-of-band, dynamic credentials rather than secrets embedded in cards.
What the specification does not tell you is which signing keys to accept, which registries to believe, which callers may invoke which skills, or what your agent should do with a reply. Those decisions are your trust model. Writing them down explicitly is the difference between a system you can audit and one where trust is whatever the last integration happened to configure.
The trust models
Five models cover almost every deployment. They are not exclusive; most real systems combine two or three, with different models at different layers.
| Model | Evidence of identity | Strengths | Weaknesses |
|---|---|---|---|
| Static allowlist | Configured endpoint plus a known client credential or certificate per partner | Simple, auditable, no external dependency | Does not scale; rotation is manual; drift between config and reality |
| Web PKI plus signed cards | TLS to the domain, card signature by a key published under that domain | Open discovery; domain ownership is a familiar identity | Domain is not intent; a compromised domain or key signs anything |
| Curated registry | A registry you trust vouches for cards and keys, possibly with selective disclosure | Central governance, revocation, review before listing | Registry becomes a single point of failure and of compromise |
| Key pinning, trust on first use | The key seen at first contact, pinned afterwards | Detects later substitution cheaply | First contact is unprotected; rotation needs an out-of-band path |
| Per-request zero trust | Every call carries a short-lived, audience-bound token checked against policy | Limits blast radius; authority is explicit and expires | Needs an identity provider and token exchange across organisations |
A common, defensible combination for cross-organisation traffic is a curated registry or allowlist to decide which organisations are partners at all, signed cards verified against keys obtained through that channel, pinning of those keys so a silent change is noticed, and per-request delegated tokens for authority. Inside one organisation, a service mesh with mutual TLS and an internal registry often replaces card signatures, because the mesh already authenticates every workload.
Verifying a card is not trusting an agent
A valid signature means the holder of a key produced this card. It does not mean the key belongs to the party you think, that the skills listed are accurate, or that the agent behaves well. Three checks turn a signature into useful evidence. First, where did you get the key? A key fetched from the same host as the card proves little beyond what TLS already proved; a key delivered through a registry, a contract, or a previously pinned value adds independent evidence. Second, does the key's identifier match one you expect for this organisation, and has it been revoked? Third, does the card's content match what you agreed? A partner that suddenly advertises a new payments skill should trigger review, not automatic use.
Plan key rotation before you need it. Because a card may carry several signatures, a partner can sign with the old and new keys during an overlap window; accept a card if any signature verifies against a currently trusted key, and record which key identifier matched. A card that verifies only against an unknown key is not a rotation you learned about; it is a substitution until a channel you trust says otherwise. Push notification callbacks need the same care in the other direction: when a remote agent calls back into your webhook, authenticate that callback as strictly as any inbound request, rather than trusting it because it refers to a task you created.
Track behaviour separately from identity. An agent with a perfect signature that returns malformed artifacts or times out half the time deserves lower trust in routing; that is what agent reputation is for.
Authority: who the request is really for
When a client agent acts for a user, the server must know both the agent and the user, and it must see authority limited to the task. The usual mechanism is a token issued for the user, exchanged for a narrower token whose audience is the specific remote agent, with only the scopes the task needs and a short lifetime. OAuth 2.0 Token Exchange (RFC 8693) is one standard way to do that exchange; securing A2A endpoints with OAuth 2.0 covers scope design.
Two rules prevent most authority bugs. Never forward the user's original token to a remote agent; it would let that agent act as the user anywhere the token is accepted. And never let a remote agent's reply widen authority. If a supplier agent says it needs payment access to finish, that is a request for your policy and your user, not an instruction your agent executes.
Content trust: the reply is untrusted input
The hardest layer is the one cryptography cannot help with. A remote agent's message is text and data produced by a system you do not control. If your agent puts that text into its own model context and then has tools that can send email, move money or change records, the remote text can steer those tools. This is prompt injection across an agent boundary, and it works against fully authenticated partners because the partner may have ingested hostile content itself.
Defend by structure, not by asking the model to be careful. Mark everything from a remote agent as tainted. Parse structured artifacts against a schema and use the parsed fields, not free text, to drive decisions. Keep tainted text out of the instruction position of prompts. And gate any side-effecting action that was influenced by tainted content behind policy, and often behind the user.
A policy function
Policy is easiest to audit when it is one function that takes the evidence and returns the allowed actions. The sketch below is illustrative; adapt the evidence fields to what your stack actually verifies.
from dataclasses import dataclass
@dataclass
class Evidence:
tls_ok: bool
card_sig_kid: str | None # kid of a verified card signature, or None
kid_source: str | None # "registry", "pinned", "same_host", None
kid_revoked: bool
org_allowlisted: bool
card_changed_since_review: bool
reply_tainted: bool # remote free text reached the decision
action_side_effect: str # "none", "reversible", "irreversible"
def decide(ev: Evidence) -> str:
if not ev.tls_ok or ev.kid_revoked:
return "deny"
identity_strong = (ev.card_sig_kid is not None
and ev.kid_source in ("registry", "pinned")
and ev.org_allowlisted)
if not identity_strong:
return "read_only" # may query, may not act
if ev.card_changed_since_review:
return "read_only_and_flag"
if ev.action_side_effect == "none":
return "allow"
if ev.reply_tainted or ev.action_side_effect == "irreversible":
return "require_user_approval"
return "allow_with_audit"Notice that even the strongest identity evidence never makes an irreversible action automatic when tainted content influenced it. Log the evidence object with every decision so incidents can be reconstructed.
Worked example: procurement across two companies
A buyer's procurement agent needs quotes from a supplier's sales agent. The buyer onboarded the supplier through its partner registry, which recorded the supplier's card signing key and its OAuth client details. At call time the client fetches the supplier's card from the well-known path, verifies the signature against the registered key, and compares the card with the version reviewed at onboarding. It obtains a token for the buyer's purchasing team, exchanges it for one restricted to the supplier agent's audience with a quote-request scope and a lifetime of minutes, and sends the task.
The reply contains a structured quote artifact and a free-text note: "Price valid today only; please confirm the order immediately." The client parses the artifact against its quote schema and records the price. The note is tainted text. Policy marks placing the order as irreversible and influenced by tainted content, so the agent proposes the order to a human buyer with the parsed quote and the note shown as supplier text. A month later the supplier rotates its signing key. The registry entry is updated through the partner channel, so verification continues; a card signed with an unknown key would have dropped the supplier to read-only and alerted the integration owner.
Failure modes
- TLS mistaken for identity. Any host with a valid certificate is treated as a partner. Require organisation-level evidence such as a registry entry or allowlist.
- Keys fetched from the same place as the card. The signature adds nothing an attacker who controls the host cannot forge. Obtain keys through an independent channel or pin them.
- Token forwarding. The user's broad token reaches a remote agent. Exchange for audience-bound, minimal tokens.
- Card drift. A partner adds skills or changes endpoints and clients start using them automatically. Diff cards against the reviewed version.
- Injection through a trusted partner. Free text from the reply drives tool calls. Taint, parse, and gate side effects.
- No revocation path. A compromised key keeps working because nothing checks revocation. Make revocation part of the registry or allowlist update.
What to do next
- Write down, for each of the five questions, what evidence your system accepts today; gaps are usually obvious once listed.
- Decide where signing keys come from and stop trusting keys fetched only from the card's host.
- Store the reviewed version of every partner card and alert on differences.
- Replace forwarded user tokens with exchanged, audience-bound, short-lived tokens limited to the task.
- Taint every remote reply, parse artifacts against schemas, and route side-effecting actions influenced by tainted content through policy or a person.
- Put the decision in one function that logs its evidence, and add tests for revoked keys, changed cards and injected replies.