Every piece of content in the Agent2Agent protocol (A2A) travels in a Part: the instruction a client sends, the file it attaches, the structured result an agent returns. What a Part means, and which content field to choose, are covered in A2A message architecture and A2A artifact types. This article is about the layer underneath: the exact bytes. It covers how the one protobuf definition becomes JSON in two bindings and binary in a third, the encoding rules that cause real bugs, and how to convert between the 0.3 and 1.0 shapes when your agents do not upgrade together.
The definitions here were checked against the A2A 1.0 protobuf (a2a.proto) and the 0.3.0 specification in October 2026. The encoding rules come from the standard protobuf JSON mapping, which A2A's JSON bindings follow. The goal is that you can write a validator and a converter that accept everything valid, reject everything invalid, and never silently change a value.
The canonical definition
In A2A 1.0 the data model is defined in Protocol Buffers. The Part message is short:
message Part {
oneof content {
string text = 1;
bytes raw = 2;
string url = 3;
google.protobuf.Value data = 4;
}
google.protobuf.Struct metadata = 5;
string filename = 6;
string media_type = 7;
}Four observations matter for the wire. The content is a oneof, so a Part carries at most one of the four fields, and the specification requires exactly one. raw is a protobuf bytes field, which binary protobuf carries as-is but JSON must encode as text. data is a google.protobuf.Value, which can hold any JSON value: an object, but also an array, a string, a number, a boolean or null. And metadata is a Struct, a JSON object whose numbers are stored as doubles.
The same Part in JSON
The JSON-RPC binding carries messages inside params, and the HTTP+JSON binding carries them as request bodies. Both use the protobuf JSON mapping, so field names become camelCase (media_type becomes mediaType), and the content field that is present identifies the kind of part. There is no discriminator. One example of each content field:
{"text": "Summarise the attached contract.", "mediaType": "text/plain"}
{"raw": "JVBERi0xLjcKJ...", "mediaType": "application/pdf", "filename": "contract.pdf"}
{"url": "https://files.example.com/c/9a1?sig=EXPIRING", "mediaType": "application/pdf",
"filename": "contract.pdf"}
{"data": {"jurisdiction": "DE", "clauses": ["liability", "termination"]},
"mediaType": "application/json", "metadata": {"schema": "https://schemas.example.com/review/v1"}}Because the member that is present is the type, a parser must look at which key exists. A client library that serialises empty strings or nulls for unused members produces a Part with several content keys, which the specification forbids. Configure serialisers to omit unset fields, and validate on receipt rather than trusting that the sender did.
Validation rules worth enforcing
A Part arrives from another agent, so treat it as untrusted input. The protocol requires exactly one content field. The rest of the list below is policy that keeps a receiver safe and predictable; publish yours in your agent's documentation so senders know what will be rejected.
- Exactly one content field. Zero or two is a malformed request. Do not pick the first one found.
- A media type on every non-text part.
mediaTypeis what a receiver dispatches on. Requiring it for raw, url and data parts costs the sender nothing and removes guessing. - Decodable, bounded bytes. Decode raw parts before accepting the request, and reject them above a size limit you set, pointing the sender at url instead.
- Fetchable, safe URLs. Accept only the schemes you fetch, fetch through a guard with a host allowlist, timeout and size cap, and never let a peer point your agent at internal addresses.
- Numbers you can represent. See the next section.
- filename is a hint. Never use it as a path on disk. Sanitise it or generate your own name.
Numbers are doubles
This is the encoding rule that most often corrupts data without an error. google.protobuf.Value stores every number as a 64-bit IEEE double, and Struct is built from Values. A double represents every integer exactly only up to 2⁵³, which is 9,007,199,254,740,992. Beyond that, integers are rounded to the nearest representable value.
Worked example: an agent sends {"data": {"orderId": 9007199254740993}}. Any receiver that parses into a protobuf Value, and many JSON parsers in JavaScript, reads 9007199254740992. The order id changed by one, nothing failed, and the agent books or cancels the wrong order. Database ids, snowflake ids, nanosecond timestamps and monetary amounts in minor units are all at risk. The protobuf documentation also notes that NaN and Infinity cannot be serialised in a Value, so a model output containing them fails at the encoder.
The fix is a convention: send identifiers and any integer that might exceed 2⁵³ as strings, in data and metadata alike, and make receivers reject large integers instead of rounding them, as the codec below does.
Byte budgets: raw versus url
The JSON bindings carry raw as base64. The protobuf JSON mapping emits standard base64 with padding and accepts standard or URL-safe base64, with or without padding, on input. Base64 turns every 3 bytes into 4 characters, so n bytes become 4 × ⌈n / 3⌉ characters.
Take a 750,000-byte scanned contract. Inline, it becomes 1,000,000 characters of base64, parsed, held in memory and possibly logged at every hop from client to orchestrator to reviewer, and replayed whenever task history is fetched. The gRPC binding carries the 750,000 bytes without inflation, but the hops and the history still hold it.
Sent as a url, the Part is under 200 characters, and only the agent that needs the bytes fetches them. The cost moves elsewhere: the receiver needs network access and permission, and the URL must remain valid for as long as the task may be retried or inspected. A practical rule is inline below a limit of a few hundred kilobytes that you set and publish, url above it, and url always for sensitive files so access can be scoped and revoked.
Version skew: 0.3 and 1.0
Version 0.3 tagged each part with a kind discriminator and nested file content in a file object with bytes or uri, plus mimeType and name. The same content in 0.3 form:
{"kind": "text", "text": "Summarise the attached contract."}
{"kind": "file", "file": {"bytes": "JVBERi0xLjcKJ...", "mimeType": "application/pdf",
"name": "contract.pdf"}}
{"kind": "file", "file": {"uri": "https://files.example.com/c/9a1", "mimeType": "application/pdf"}}
{"kind": "data", "data": {"jurisdiction": "DE"}}| 0.3 | 1.0 | Note |
|---|---|---|
kind: "text", text | text | Same content |
file.bytes | raw | Base64 in JSON in both |
file.uri | url | Renamed |
file.mimeType | mediaType | Moved to the Part, applies to every kind |
file.name | filename | Moved to the Part |
kind: "data", data | data | 0.3 requires an object; 1.0 accepts any JSON value |
The asymmetry in the last row is the one converters miss. Converting 0.3 to 1.0 is lossless. Converting 1.0 to 0.3 is not when the data is an array or scalar, because 0.3 has nowhere to put it. The codec wraps such values under a key; that wrapping is a local convention, not part of either specification, so agree it with the old peer. The A2A 1.0 specification's appendix covers migration from 0.3; follow it for everything beyond parts. The cleanest design converts once at the edge, keeps one internal shape, and records which version each peer speaks.
import base64, binascii
CONTENT = ("text", "raw", "url", "data")
MAX_RAW = 512 * 1024 # our policy: larger files must use url
class PartError(ValueError):
pass
def from_v03(p):
# A2A 0.3 part -> 1.0 shape. Unknown kinds are rejected, not guessed.
kind, out = p.get("kind"), {}
if kind == "text":
out["text"] = p["text"]
elif kind == "data":
out["data"] = p["data"]
elif kind == "file":
f = p["file"]
if ("bytes" in f) == ("uri" in f):
raise PartError("file needs exactly one of bytes or uri")
out["raw" if "bytes" in f else "url"] = f.get("bytes", f.get("uri"))
if "mimeType" in f: out["mediaType"] = f["mimeType"]
if "name" in f: out["filename"] = f["name"]
else:
raise PartError(f"unknown 0.3 kind {kind!r}")
if "metadata" in p: out["metadata"] = p["metadata"]
return out
def to_v03(p):
# 1.0 part -> 0.3 shape for old peers. 0.3 data must be an object, so non-object
# values are wrapped under "value": a LOCAL convention, document it for the peer.
validate(p)
meta = {"metadata": p["metadata"]} if "metadata" in p else {}
if "text" in p:
return {"kind": "text", "text": p["text"], **meta}
if "data" in p:
d = p["data"]
return {"kind": "data", "data": d if isinstance(d, dict) else {"value": d}, **meta}
f = {"bytes": p["raw"]} if "raw" in p else {"uri": p["url"]}
if "mediaType" in p: f["mimeType"] = p["mediaType"]
if "filename" in p: f["name"] = p["filename"]
return {"kind": "file", "file": f, **meta}
def normalise(p):
return from_v03(p) if "kind" in p else p
def validate(p):
present = [k for k in CONTENT if k in p]
if len(present) != 1:
raise PartError(f"exactly one content field required, got {present}")
if "raw" in p:
try:
# ProtoJSON emits standard base64 with padding; accept URL-safe input too
s = p["raw"].replace("-", "+").replace("_", "/")
blob = base64.b64decode(s + "=" * (-len(s) % 4), validate=True)
except (binascii.Error, ValueError) as e:
raise PartError(f"raw is not valid base64: {e}")
if len(blob) > MAX_RAW:
raise PartError(f"raw part is {len(blob)} bytes; send a url instead")
if "url" in p and not p["url"].startswith("https://"):
raise PartError("only https urls are accepted")
if "data" in p:
check_numbers(p["data"])
return present[0]
def check_numbers(v, path="data"):
# Value/Struct numbers are IEEE doubles: integers beyond 2**53 are not exact.
if isinstance(v, bool):
return
if isinstance(v, int) and abs(v) > 2**53:
raise PartError(f"{path}: integer {v} exceeds 2**53; send it as a string")
if isinstance(v, dict):
for k, x in v.items(): check_numbers(x, f"{path}.{k}")
elif isinstance(v, list):
for i, x in enumerate(v): check_numbers(x, f"{path}[{i}]")
Order and streaming
Parts are an ordered list, and order carries meaning: an instruction followed by an attachment, or a summary followed by its data. Preserve order through every conversion and never deduplicate parts.
Streaming adds one more rule. A TaskArtifactUpdateEvent carries an artifact plus two flags: append, meaning the parts extend an artifact already sent with the same id, and lastChunk, meaning this is the final piece. A client must apply events in arrival order and must not treat an artifact as complete until the last chunk arrives. Whether to join consecutive text parts into one string for display is a rendering choice, not a protocol rule, so keep the parts as received and join only in the view.
def apply_artifact_update(artifacts, event):
# event: TaskArtifactUpdateEvent in JSON (artifact, append, lastChunk)
art = event["artifact"]
parts = [normalise(p) for p in art.get("parts", [])]
for p in parts:
validate(p)
current = artifacts.get(art["artifactId"])
if event.get("append") and current is not None:
current["parts"].extend(parts) # order of arrival is the order of parts
else:
artifacts[art["artifactId"]] = {**art, "parts": parts} # new or replaced
if event.get("lastChunk"):
artifacts[art["artifactId"]]["complete"] = TrueTransport-level details of the stream, including reconnects and ordering on the SSE connection, are covered in A2A streaming architecture.
Worked example: one request, checked
A procurement agent asks a contract-review agent for a risk summary. It sends three parts: a text instruction, a data part with the review scope, and the contract. The validator runs as follows. The text part has one content field and passes. The data part contains {"vendorId": "1180023377401299457", "maxLiability": 250000}; the vendor id is a string, so no rounding can occur, and 250,000 is far below 2⁵³. The contract is 750,000 bytes, above the 512 KiB inline policy, so the sender's own validator rejects the inline form before sending and the part goes out as a url to a signed link valid for 24 hours. The whole request is about 600 bytes instead of over a megabyte. The reviewer agent, still on 0.3, receives the parts through the edge converter: the url part becomes a 0.3 file part with uri, mimeType and name, and the data object passes through unchanged because it is already an object.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Rejected requests with two content keys | Serialiser emits empty strings or nulls for unused oneof members | Omit unset fields; validate before sending |
| Ids or amounts off by one or more | Integers above 2⁵³ in data or metadata became doubles | Send them as strings; reject large integers on receipt |
| Memory spikes and slow history fetches | Large raw parts replayed through every hop | Enforce an inline size limit; use url |
| Old peer drops structured results | 1.0 array or scalar data sent to a 0.3 peer | Wrap at the edge and document the convention |
| Base64 decode errors from some clients | URL-safe or unpadded base64 | Accept both alphabets and missing padding on input |
| Streamed artifacts shown half-finished | Client rendered before lastChunk | Track completion per artifact id |
Operational guidance and trade-offs
- Log shapes, not bytes. Log content field, media type and size for each part, never raw content, which may be private and is always large.
- Count by version. A metric of inbound parts by protocol version tells you when the last 0.3 peer has gone and the converter can be removed.
- Property-test the codec. Generate random valid parts and assert that 0.3 to 1.0 to 0.3 returns the input, and that every invalid part is rejected with a clear error.
- Publish limits on your Agent Card. Inline size limits, accepted URL schemes and media types belong in skill documentation; see the Agent Card specification.
What to do next
- Configure your serialiser to omit unset fields, then add the exactly-one-content-field check on send and on receive.
- Search your data and metadata payloads for integers that can exceed 2⁵³ and change them to strings.
- Set and publish an inline byte limit, and move anything larger to scoped, expiring URLs.
- If any peer still speaks 0.3, add an edge converter with round-trip property tests and a per-version metric.
- Make your streaming client track append and lastChunk per artifact id and render only completed artifacts as final.