When one agent can delegate work to many others, it needs a way to prefer the ones that do good work. A signed Agent Card proves who published an agent. Authentication proves who is calling. Neither says whether the agent's translations are accurate, whether it meets its latency claims, or whether it fails on one skill while succeeding on another. That is a reputation problem, and it needs its own architecture.

The A2A protocol does not define a reputation mechanism, and nothing here is part of the specification. This article designs a reputation layer that sits beside A2A. It uses A2A task outcomes as raw material and feeds scores to your routing and registry. It covers what counts as evidence, how to turn evidence into a score that is honest about uncertainty, how to give new agents a fair start, and how to resist fake identities and colluding raters.

Advertisement

What reputation is, and what it is not

Reputation is a prediction: given past evidence, how likely is it that this agent will complete this kind of task well? Three properties follow from that definition. It is per skill, because an agent that is excellent at summarization can be poor at legal translation. It is uncertain, because ten successes say less than a thousand. And it is perishable, because agents are redeployed, models are swapped and last quarter's evidence describes a different system.

Keep reputation separate from its neighbours. Identity and provenance belong to the Agent Card and its signature. Permission to call belongs to A2A authentication and authorization. Discovery and liveness belong to the agent registry. Fast failure isolation belongs to a circuit breaker, which reacts in seconds to a burst of errors; reputation moves over days. A design that merges any of these into the score becomes hard to reason about and easy to game.

Reputation as a pipeline on top of A2A: evidence in, scores out, policy decidesCalling agentssigned outcome reportsGatewayobserved latency, errorsVerifiersschema checks, gradersEvidence logappend-only, per skillScorerdecay, Beta, credibilityScore storeper agent x skill x scopeRouter / registryrank, explore, gateAudit and appealsexplain every scorebatch or streamreadtask outcomesThe A2A protocol carries the tasks. Nothing in the specification defines reputation;every box on this diagram is a layer you build and operate yourself.
Evidence comes from callers, the gateway and verifiers; it is logged, scored with decay and credibility weights, and read by the router and registry. Every score must be explainable from its evidence.

Where evidence comes from

Each delegated task ends in a terminal state defined by A2A: completed, failed, canceled or rejected. The exact wire spelling of these states differs between specification versions, so normalize them at ingestion. Terminal state alone is weak evidence. An agent can report completion and return nonsense, and a failure can be the caller's fault. Use three sources, and trust them in this order:

  • Verifier results. Checks you run on the output: schema validation, unit tests for generated code, a grader model scoring a sample against a rubric. These are the most objective signals and the hardest to fake.
  • Gateway observations. If delegated calls pass through your gateway, it measures latency, error rates and timeouts directly. The agent under measurement cannot inflate these numbers.
  • Caller reports. The calling agent's judgement of the result. These are subjective, easy to fake and the only source for many tasks, so they must be weighted by the caller's credibility.

Every piece of evidence becomes an immutable record bound to a task id, signed by whoever produced it, and scoped to a tenant. The format below is our own; A2A does not define one:

{
  "evidence_id": "ev_01J9Z6...",
  "task_id": "the A2A task id this evidence is about",
  "subject": {"agent": "https://agents.example-vendor.com/translate", "skill": "translate-legal"},
  "rater":   {"agent": "https://agents.example-buyer.com/contracts", "org": "example-buyer.com"},
  "source": "caller_report",          // or "gateway_observed", "verifier"
  "outcome": "success",               // success | failure | neutral
  "reasons": ["output_valid", "latency_ok"],
  "observed_at": "2026-10-01T09:14:03Z",
  "scope": "tenant:acme",             // who may see and use this evidence
  "signature": "JWS over the fields above, by the rater's key"
}

Binding evidence to a task id that your gateway or tracing system has also seen prevents raters from inventing interactions that never happened. The scope field keeps one customer's experience from leaking into another's view unless both agreed to share it.

Advertisement

Scoring: Beta evidence, decay and a pessimistic bound

Treat each agent-skill pair as having an unknown success probability. With a uniform prior and s successes and f failures, the posterior is a Beta(1 + s, 1 + f) distribution, whose mean is (s + 1) / (s + f + 2). This is the core of the Beta reputation system proposed by Jøsang and Ismail. It handles small samples sensibly: one success gives a mean of 0.67, not 1.0.

Two refinements make it usable. First, decay: each piece of evidence is weighted by one half raised to its age divided by a half-life, so old behaviour fades smoothly and the score tracks the current system. Thirty days is a reasonable default for agents that deploy weekly. Second, rank on a lower bound, not the mean. The Wilson lower bound of the success rate penalizes small samples: it asks how bad the agent could plausibly be. Credibility weights multiply each record before counting, so a low-credibility report counts for a fraction of a success.

import math, random
from dataclasses import dataclass

HALF_LIFE_DAYS = 30
Z = 1.96

@dataclass
class Evidence:
    outcome: str        # "success" | "failure" | "neutral"
    age_days: float
    weight: float       # credibility of the rater or source, 0..1

def decayed_counts(evidence):
    s = f = 0.0
    for e in evidence:
        w = e.weight * 0.5 ** (e.age_days / HALF_LIFE_DAYS)
        if e.outcome == "success":
            s += w
        elif e.outcome == "failure":
            f += w
    return s, f

def wilson_lower(s, f, z=Z):
    """Pessimistic score for ranking: lower bound of the success rate."""
    n = s + f
    if n == 0:
        return 0.0
    p = s / n
    denom = 1 + z * z / n
    centre = p + z * z / (2 * n)
    margin = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n))
    return (centre - margin) / denom

def thompson_pick(candidates):
    """Exploration for routing: sample each agent's Beta(1+s, 1+f), take the max.
    New agents get traffic in proportion to how plausible it is that they are best."""
    best, best_draw = None, -1.0
    for name, (s, f) in candidates.items():
        draw = random.betavariate(1 + s, 1 + f)
        if draw > best_draw:
            best, best_draw = name, draw
    return best

Worked example: ranking a veteran and a newcomer

Two agents offer the same skill. Agent A has 95 successes in 100 recent tasks. Agent B is new, with 9 successes in 10. Their raw success rates are 0.95 and 0.90, and their Beta means are 96/102 = 0.94 and 10/12 = 0.83.

The Wilson lower bounds at 95 percent confidence separate them more clearly. For A, with n = 100 and p = 0.95, the bound is about 0.89. For B, with n = 10 and p = 0.90, it is about 0.60. A pure ranking on the lower bound would send every task to A, and B would never collect the evidence needed to prove itself. That is the cold-start problem.

Thompson sampling solves it. For each task, draw one sample from each agent's Beta distribution and route to the highest draw. A's samples cluster tightly around 0.94. B's spread widely, from about 0.6 to nearly 1.0, so B wins a minority of draws, gets real traffic, and its distribution narrows as evidence arrives. If B is as good as A, its share grows; if it is worse, its share falls on its own. Reserve this exploration for low-risk tasks, and route high-stakes tasks by lower bound only.

Rater credibility and propagation

Caller reports are only as good as the caller. A simple credibility model scores each rater by how often its reports agree with verifier results on the same tasks: a rater whose "success" reports usually match passing schema checks earns weight, one that marks failed outputs as successes loses it. Raters with no overlapping verifiable tasks start at a low default.

Where there is no verifier, propagate trust through the rating graph. EigenTrust, proposed by Kamvar, Schlosser and Garcia-Molina, computes global trust as the stationary vector of normalized local trust, mixed with a set of pre-trusted peers such as your own organization's agents. A rater's opinions count in proportion to its own global trust, so a ring of unknown agents praising each other gains little, because nobody trusted is pointing into the ring.

import numpy as np

def global_trust(local, pretrusted, alpha=0.15, iters=50):
    """EigenTrust-style propagation.
    local[i][j] >= 0: how much rater i trusts agent j from i's own experience.
    pretrusted: distribution over agents you trust by fiat (your own org's agents)."""
    C = np.array(local, dtype=float)
    pretrusted = np.asarray(pretrusted, dtype=float)
    rows = C.sum(axis=1, keepdims=True)
    C = np.divide(C, rows, out=np.tile(pretrusted, (C.shape[0], 1)), where=rows > 0)
    t = np.array(pretrusted, dtype=float)
    for _ in range(iters):
        t = (1 - alpha) * C.T @ t + alpha * np.array(pretrusted)
    return t     # a rater's opinions count in proportion to its own trust

Run propagation as a batch job, hourly or daily, not per request. It is a whole-graph computation, and scores that change slowly are easier to explain.

Sybil resistance and collusion

A Sybil attack creates many identities to inflate a score or to escape a bad one. The design counters it at three points. Identities must be costly: only agents with a verifiable identity, such as a signed card tied to a domain the publisher controls, can hold or give reputation. New identities start at the prior, so a fresh name can never beat an established good score; this also removes the reward for whitewashing, abandoning a bad identity for a new one. And influence is capped per organization: count distinct rater organizations, and limit how much any single organization can move a score.

Collusion between real organizations is harder to stop. Useful signals include reciprocal rating pairs that rate each other far above everyone else, ratings not backed by gateway-observed tasks, and sudden bursts from new raters. Treat them as flags for review, not automatic penalties, because legitimate partners also transact heavily with each other.

Using the score without being ruled by it

Routing is the main consumer: rank candidates for a skill, explore with Thompson sampling where risk allows, and exclude agents below a floor. The registry can display scores next to search results. Keep hard gates rare and coarse, such as a minimum lower bound for regulated workflows, because noisy scores used as fine thresholds cause oscillation: an agent drops below the line, loses traffic, cannot recover evidence, and stays excluded.

Every score must be explainable. Store which evidence records contributed and with what weights, and expose that to the agent's owner. Provide an appeal path for evidence the owner believes is wrong. Record score changes in your tracing or audit system, so a routing decision can be reconstructed later.

Failure modes

  • Rating inflation. Callers mark everything a success. Weight by agreement with verifiers, and prefer gateway and verifier evidence.
  • Blame misattribution. A caller sends a bad input and reports failure. Separate input-rejection outcomes, including rejected tasks, from quality failures.
  • Stale scores. An agent is redeployed with a new model and keeps its old reputation. Decay handles this slowly; also reset or widen uncertainty when the card's version changes.
  • Starvation. New agents never get traffic. Use exploration with a small, bounded share.
  • Cross-tenant leakage. Evidence from one customer shapes another's routing without consent. Enforce the scope field at scoring time.
  • Retaliation. An agent rates a competitor down after losing work to it. Cap per-organization influence and flag anomalous pairs.

What to do next

  1. Define the evidence record and require a task id, a signature and a scope on every record.
  2. Instrument your gateway to emit latency and error evidence for every delegated task.
  3. Add at least one verifier per high-volume skill, even if it only validates the output schema.
  4. Implement decayed Beta counts and the Wilson lower bound, and rank candidates with it.
  5. Add Thompson-sampling exploration for low-risk skills, capped at a small share of traffic.
  6. Build the explanation view and an appeal path before you let scores gate any workflow.
Key takeaway: The A2A protocol moves tasks between agents but defines no reputation, so you build it as a separate layer. Collect signed evidence tied to real task ids, and prefer verifier and gateway observations to caller reports. Score each agent per skill with decayed Beta counts, rank on a pessimistic lower bound, and use Thompson sampling to give newcomers a fair start. Weight raters by their credibility, propagate trust from pre-trusted peers, make identities costly and cap per-organization influence. Keep scores explainable and appealable, and keep them separate from identity, authorization and circuit breaking.