Sooner or later an agent has to hand another agent a file: a scanned contract, a generated chart, a CSV export, an audio clip. The A2A protocol gives you one type for this, the Part, and two ways to carry the bytes: inline, or by reference to somewhere the receiver can fetch them. Both work in a demo. In production, the choice determines how large your requests get, what ends up in logs and task history, whether a link still works when someone reads the task tomorrow, and whether a hostile peer can make your agent fetch an internal address.
Which content field fits which kind of output is covered in A2A artifact types. This page is about the mechanics underneath for files specifically: encoding, size, transfer direction, chunking, link lifetime and safe retrieval. Field names follow the A2A specification version 1.0.0.
What the specification gives you
In A2A 1.0 a Part carries exactly one content field: text, raw, url or data. For files, two matter. raw holds the bytes themselves. url holds a string reference to content stored elsewhere. Alongside the content field, every Part may carry mediaType (the MIME type), filename and metadata. Parts appear in Messages, which is how a client sends input, and in Artifacts, which is how an agent returns results.
The canonical data model is defined in Protocol Buffers, where raw is a bytes field. Bindings then decide how that looks on the wire. In the JSON-RPC and HTTP+JSON bindings, bytes become a base64 string. In the gRPC binding they travel as native binary protobuf bytes with no expansion. The same agent can therefore have very different costs depending on which interface a client chose from its Agent Card.
If you still interoperate with version 0.3 peers, the older shape was different: a part with a kind discriminator of "file" and a nested file object holding either bytes or uri, plus mimeType and name. Convert at the edge into one internal representation and keep the version split out of business logic.
The real cost of inline bytes
Base64 turns every 3 bytes into 4 characters, so an n-byte file becomes 4 x ceil(n / 3) characters: a 33% expansion before JSON quoting. That is the visible cost. The larger cost is copies. A typical JSON stack reads the whole request body, parses it into a string, decodes that string into a byte buffer and may hand the buffer to a model client that encodes it again. A 9 MB PDF can easily occupy 40 MB or more of memory at peak in one request.
Then the bytes persist. The task store keeps the Message and Artifacts in history. Streaming servers often keep events so a client that reconnects can resubscribe to the task. Request logging, tracing payload capture and dead-letter queues all copy what they see. Inline files multiply through every one of them, and so do any secrets inside those files.
| File size | Base64 characters | Practical note |
|---|---|---|
| 50 KB icon | about 67 KB | inline is fine |
| 600 KB scan | about 800 KB | inline only if logs and history never capture payloads |
| 5 MB image | about 6.7 MB | over the 4 MiB default message size common in gRPC implementations, even as raw bytes |
| 9 MB PDF | about 12 MB | use a URL; many gateways cap bodies at 10 MB |
The specification itself does not set a file size limit, so limits come from your server framework, gateway and storage. Decide them explicitly, publish them in the skill description on your Agent Card, and reject oversize parts with a clear error rather than a timeout. A reasonable starting rule is: inline below a few hundred kilobytes, reference above.
Decoding base64 correctly
Because JSON bytes follow the proto3 JSON mapping, encoders emit standard base64 with padding, but conforming parsers accept both the standard and the URL-safe alphabets, with or without padding. A receiver built on a plain base64.b64decode will reject valid input from a peer that emits URL-safe characters, and one built on a lenient decoder may silently discard junk characters. Decode strictly, but accept both alphabets, and check the size before decoding: the decoded length is about three quarters of the string length, so you can refuse an oversized part without allocating it.
import base64, binascii
MAX_INLINE = 512 * 1024 # your published limit, in decoded bytes
def decode_raw(s: str) -> bytes:
if (len(s) * 3) // 4 > MAX_INLINE + 3:
raise ValueError("inline part exceeds limit; send a url part instead")
s = s.strip().replace("-", "+").replace("_", "/") # accept URL-safe alphabet
s += "=" * (-len(s) % 4) # accept missing padding
try:
data = base64.b64decode(s, validate=True) # reject any other characters
except binascii.Error as e:
raise ValueError(f"raw part is not valid base64: {e}")
if len(data) > MAX_INLINE:
raise ValueError("inline part exceeds limit")
return dataOne frequent interoperability bug: putting a data URL such as data:image/png;base64,iVBOR... into raw. The prefix is not base64 and a strict decoder rejects it. Put plain base64 in raw and the type in mediaType. If a peer insists on data URLs, they belong in url, and your fetcher must decide whether to accept the data: scheme at all.
URL parts: who fetches, with what authority
A url part moves the transfer out of the A2A call. The protocol does not define how the receiver authenticates to the URL, so there are two common patterns. Capability URLs are pre-signed object-storage links whose query string grants access to one object for a limited time; possession is authority. Authenticated URLs point at an endpoint that requires credentials the receiver already holds for the producer, such as the same OAuth token it uses for the A2A call. Capability URLs are simpler across organisations; authenticated URLs are revocable and auditable.
The trap is lifetime. A signed URL valid for 15 minutes is generated once, but the Part containing it lives as long as the task. A client that reads task history next morning, a push notification consumer that falls behind (see A2A push notifications), or a human reviewing an audit trail all find a dead link. Three ways out: issue URLs with an expiry longer than your task retention; point the url at a stable, authenticated resource and sign only when it is fetched; or document that file links expire and offer a skill that reissues them. Whatever you choose, state it, because the receiver cannot tell an expired link from a broken one.
Integrity is not in the specification either. A useful convention is to put a SHA-256 digest and the byte size in the Part's metadata, so the receiver can verify what it fetched and refuse a swapped or truncated object. Name these keys in your skill documentation; they are your contract, not the protocol's.
Sending files to an agent
Uploads run the other way and are less obvious. A2A has no separate upload operation: a client sends files as parts of the Message in SendMessage or SendStreamingMessage. So the client either inlines small files, or uploads large ones to storage the agent can reach and sends a url part. In a single organisation that is usually a shared bucket with per-task prefixes. Across organisations, the agent often exposes its own pre-signed upload endpoint, outside A2A, and documents it in the skill description.
Before sending anything, check what the agent accepts. Agent Cards declare defaultInputModes as media types, and skills can override them with their own inputModes. A client that sends image/heic to a skill declaring only image/png and image/jpeg should convert first, not hope.
Large outputs and chunked artifacts
For streaming, the agent sends TaskArtifactUpdateEvent objects. Each carries an artifact with an artifactId, and two flags: append, meaning the parts should be added to an artifact already received with that id, and lastChunk, meaning this is the final piece. That lets an agent emit a long file incrementally instead of buffering it.
Be precise about what the protocol promises. Appending adds parts to the artifact's list; the specification does not say that consecutive raw parts form one byte stream. If you split a file across chunks, declare that convention, for example a metadata key on the artifact saying its raw parts concatenate in order, and give the total size and digest on the last chunk. Each chunk's base64 is decoded independently, so chunk boundaries do not need to fall on 3-byte multiples. For anything larger than a few megabytes, emitting one url part when the file is complete is simpler and survives reconnects better than dozens of inline chunks, and keeps each event small, which matters for the reconnect and replay behaviour described in A2A streaming.
A receiver that fetches safely
A url part is an instruction from a peer to make your agent issue a request. Without guards, that is server-side request forgery: a malicious or compromised peer sends http://169.254.169.254/latest/meta-data/ or an internal admin address and reads the response through your agent's output. The fetcher below enforces scheme, resolves the host and refuses private, loopback and link-local addresses, follows no redirects, caps size while streaming and checks media type and digest.
import hashlib, ipaddress, socket
from urllib.parse import urlparse
import httpx
MAX_FETCH = 50 * 1024 * 1024
def _public_host(host: str) -> None:
for info in socket.getaddrinfo(host, 443, proto=socket.IPPROTO_TCP):
ip = ipaddress.ip_address(info[4][0])
if ip.is_private or ip.is_loopback or ip.is_link_local or ip.is_reserved or ip.is_multicast:
raise PermissionError(f"refusing non-public address {ip}")
def fetch_part(part: dict, allowed_types: set[str]) -> bytes:
u = urlparse(part["url"])
if u.scheme != "https" or not u.hostname:
raise PermissionError("only https urls are accepted")
_public_host(u.hostname)
meta = part.get("metadata", {})
h = hashlib.sha256()
buf = bytearray()
with httpx.Client(follow_redirects=False, timeout=httpx.Timeout(10, read=30)) as client:
with client.stream("GET", part["url"]) as r:
if r.status_code != 200:
raise IOError(f"fetch failed with HTTP {r.status_code} (expired link?)")
ctype = r.headers.get("content-type", "").split(";")[0].strip()
declared = part.get("mediaType", ctype).split(";")[0].strip()
if ctype not in allowed_types or ctype != declared:
raise ValueError(f"unexpected content type {ctype}")
for chunk in r.iter_bytes():
buf += chunk; h.update(chunk)
if len(buf) > MAX_FETCH:
raise ValueError("remote file exceeds limit")
want = meta.get("sha256")
if want and h.hexdigest() != want:
raise ValueError("digest mismatch")
return bytes(buf)Resolving and then connecting leaves a small window in which DNS can change (DNS rebinding). For high-risk deployments, run fetches through an egress proxy that enforces the same address policy at connect time, or pin the connection to the resolved IP. Treat the downloaded content as untrusted input as well: parse PDFs and images in a sandbox, and never pass fetched text to a model as instructions. The broader controls are in A2A security.
Worked example on the wire
A client asks an invoice agent to extract totals from a small receipt image and a large statement PDF. The image goes inline; the PDF goes by reference with a digest in metadata.
POST /a2a/v1 HTTP/1.1
Content-Type: application/json
A2A-Version: 1.0
{"jsonrpc": "2.0", "id": "req-31", "method": "SendMessage",
"params": {"message": {"messageId": "m-88", "role": "ROLE_USER", "parts": [
{"text": "Extract the total and currency from both documents."},
{"raw": "iVBORw0KGgoAAAANSUhEUgAA...", "mediaType": "image/png", "filename": "receipt.png"},
{"url": "https://files.example.com/u/t-77/statement.pdf?X-Expires=86400&sig=...",
"mediaType": "application/pdf", "filename": "statement.pdf",
"metadata": {"sha256": "9f2c...e1", "size": 9437184}}
]}}}The agent decodes the PNG with the size-checked decoder, fetches the PDF through the guarded fetcher, verifies the digest and returns a Task whose artifact holds a data part with the two totals. If the fetch fails because the link expired, the agent returns the task in a failed or input-required state with a message naming the file, so the client can re-upload instead of retrying blindly.
Failure modes and trade-offs
- Bodies rejected before your code runs. Gateways and frameworks cap request size; inline files produce 413 responses that look like agent failures.
- Files in logs. Payload logging captures every inline file, including personal data. Redact
rawfields at the logging layer. - Expired links in history. The task outlives the signature; readers see 403 errors with no explanation.
- Missing or wrong mediaType. Receivers sniff content and guess. Always set it, and verify it against the fetched Content-Type.
- SSRF through url parts. Any agent that fetches arbitrary URLs from peers is a proxy into your network until proven otherwise.
- Unstated chunk semantics. A consumer that concatenates appended parts and one that renders them separately will both be right by the specification and disagree with each other.
The trade-off in one line: inline is simpler, self-contained and atomic with the task, but expensive and sticky; URLs are cheap and scale to any size, but add a second system with its own authority, lifetime and attack surface.
What to do next
- Pick and publish an inline size limit, and enforce it before decoding.
- Make your decoder accept standard and URL-safe base64, with and without padding, and reject everything else.
- Move files above the limit to url parts, and put sha256 and size in part metadata.
- Choose a link-lifetime policy that matches task retention, and document it in the skill description.
- Route every url fetch through a guarded fetcher or egress proxy: https only, public addresses only, no redirects, size and type caps.
- Strip raw fields from logs and traces, and add a test that sends an oversize part, an expired link and an internal-address URL.