An A2A artifact is the deliverable a remote agent produces for a task: a report, an image, a table, a patch. Its structure is small. An Artifact has an artifactId, an optional name and description, metadata, and a list of Parts, and every Part carries exactly one of four content fields: text, raw, url or data. Alongside the content, a Part can carry a mediaType, a filename and metadata.
That small structure leaves the important decisions to you: which content field to use, which media type to declare, how several parts relate, and how a client and an agent agree on formats before work starts. This article treats those decisions as a type system. It covers the fields as defined in the A2A 1.0 specification, the conventions that make artifacts machine-consumable, negotiation, and a consumer that dispatches safely by type. Streaming assembly and the inline-versus-reference trade-off are covered in A2A artifact exchange architecture; this page is about what the parts mean.
Four content fields, one Part
The four content fields answer one question: how does the content travel? text is a string carried inline. raw is bytes; the specification's canonical model is Protocol Buffers, so in the JSON bindings bytes are base64-encoded, which inflates them by a third. url points at content stored elsewhere for the receiver to fetch. data is any JSON value carried inline. A Part with two content fields, or none, is invalid.
Version 1.0 flattened the earlier design. Version 0.3 used a kind discriminator and nested file objects with mimeType; in 1.0 the member that is present identifies the part, and file properties moved onto the Part as mediaType and filename. If you still talk to 0.3 peers, convert at the boundary and keep one internal shape. A2A message architecture covers the same Part type from the message side, where clients send it as input.
Choosing the content field
| Content | Use | Avoid when |
|---|---|---|
| Prose, Markdown, source code | text with text/markdown, text/x-python and so on | It is really structured data a program will parse |
| Structured results | data with application/json or a specific +json type | The payload is large or binary |
| Small binary: icons, short audio, thumbnails | raw with the real media type and a filename | Above a few hundred kilobytes; base64 adds a third |
| Large or long-lived files | url to short-lived, scoped storage | The receiver cannot reach your storage |
| Sensitive files | url with per-task authorisation | You would put the bytes in logs by inlining them |
Two rules settle most cases. If a program will parse it, use data, not JSON inside a text part: a data part can be validated against a schema before anything acts on it, while JSON in text invites regex extraction and silent breakage when a model adds a sentence before the brace. And if the content is large, prefer url: inline bytes travel through every proxy, log and task store that touches the task, and stream events carrying them become slow to replay.
Media types are the type system
The content field is transport; mediaType is meaning. A consumer should dispatch on media type, and the producer should always set it, even where a default seems obvious. Use registered types where they exist: text/markdown, text/csv, image/png, application/pdf. Parameters such as charset are legal and should be ignored for matching.
For structured data, application/json says only that the payload is JSON, which is not enough for a consumer to trust its shape. Two conventions help, neither of them mandated by A2A. One is a structured-syntax type such as application/vnd.acme.invoice+json, which says both that it is JSON and which document shape it is. The other is a schema URI in the Part's metadata, versioned in the path, which the consumer resolves and validates against. Whichever you choose, document it in the skill description on your Agent Card, version it, and reject unknown major versions explicitly rather than guessing. The Agent Card specification, field by field covers where those declarations live.
Multi-part artifacts
One artifact should be one deliverable; parts are its pieces. A variance report might be a Markdown summary for a person, a data part with the rows for a program, and a PDF by URL for the archive. The specification does not say whether such parts are components of one whole or alternative renditions of the same content, so a consumer cannot know whether to use all of them or pick one. State the relationship in artifact metadata and in your skill documentation, as in the example below, where partsAre is our own convention.
{
"artifactId": "q3-variance-report",
"name": "Q3 cost variance",
"description": "Actual versus budget by cost centre",
"parts": [
{"text": "## Q3 variance\nSpend ran 4.2% over budget, mostly cloud egress in CC-110.",
"mediaType": "text/markdown"},
{"data": {"currency": "EUR",
"rows": [{"costCentre": "CC-110", "budget": 120000, "actual": 131500}]},
"mediaType": "application/json",
"metadata": {"schema": "https://schemas.example.com/variance/v2.json"}},
{"url": "https://files.example.com/r/7f3c?sig=EXPIRING",
"mediaType": "application/pdf", "filename": "q3-variance.pdf"}
],
"metadata": {"partsAre": "components"}
}Keep separate deliverables in separate artifacts with stable ids, because artifact ids are what streaming updates and later turns refer to. A code-review agent that returns a summary and a patch should emit two artifacts, review-summary and patch, rather than one artifact whose second part happens to be a diff. Clients can then render, store or apply each independently, and a later refinement can replace one without resending the other.
Negotiating output types
Negotiation happens at three levels. The Agent Card declares defaultOutputModes, the media types the agent produces by default. Each skill may override them with its own outputModes. A request can carry acceptedOutputModes in its configuration, the types this client can handle, and the specification says agents SHOULD use it to tailor output. The input side mirrors this with defaultInputModes and per-skill inputModes, and a request whose parts use an unsupported type is answered with ContentTypeNotSupportedError.
The server-side logic is an intersection. Take the skill's output modes, or the card defaults if the skill sets none, keep those the client accepts, and produce parts only in the survivors. If nothing survives, refusing early with a clear error is better than doing the work and returning something the client will drop. The wildcard handling below borrows HTTP Accept conventions; A2A does not define matching rules, so publish the ones you implement.
class ContentTypeNotSupported(Exception): # surfaced as ContentTypeNotSupportedError (-32005)
pass
def media_match(offer, accepted):
"""HTTP-style matching: parameters ignored, 'type/*' and '*/*' wildcards.
A convention: the A2A spec lists media types but does not define wildcard rules."""
o = offer.split(";")[0].strip().lower()
for a in accepted:
a = a.split(";")[0].strip().lower()
if a in ("*/*", o) or (a.endswith("/*") and o.split("/")[0] == a[:-2]):
return True
return False
def plan_output(card, skill_id, accepted):
skill = next(s for s in card["skills"] if s["id"] == skill_id)
offered = skill.get("outputModes") or card["defaultOutputModes"]
if not accepted:
return offered # no stated preference
usable = [m for m in offered if media_match(m, accepted)]
if not usable:
raise ContentTypeNotSupported(f"offer {offered}, client accepts {accepted}")
return usable # produce parts only in these types, best first
Consuming parts safely
A consumer receives content from another agent, which makes every part untrusted input. Check the one-content-field invariant, resolve the media type, decode and size-check inline bytes, fetch URLs only through a guarded fetcher with a host allowlist, size cap and timeout so a peer cannot point you at internal services, and dispatch to a handler registered for the type. Parts with unknown types should be kept and reported, not coerced into text. Defaulting untyped parts to text/plain or application/json is a local convention; log it so producers can be asked to fix their types.
import base64
MAX_INLINE = 512 * 1024
DEFAULT_MEDIA = {"text": "text/plain", "data": "application/json"} # local convention for untyped parts
def consume_part(part, handlers, fetch):
kinds = [k for k in ("text", "raw", "url", "data") if k in part]
if len(kinds) != 1:
raise ValueError(f"a part must carry exactly one content field, got {kinds}")
kind = kinds[0]
media = (part.get("mediaType") or DEFAULT_MEDIA.get(kind, "application/octet-stream"))
media = media.split(";")[0].strip().lower()
if kind == "raw":
body = base64.b64decode(part["raw"], validate=True)
if len(body) > MAX_INLINE:
raise ValueError("inline part over the size limit")
elif kind == "url":
body = fetch(part["url"], expect=media) # allowlisted hosts, size cap, timeout, no redirects to private IPs
else:
body = part[kind]
handler = handlers.get(media) or handlers.get(media.split("/")[0] + "/*")
if handler is None:
return {"skipped": media, "filename": part.get("filename")} # keep it, do not guess
return handler(body, part.get("metadata", {}))Two more rules. Treat text parts as data, not instructions: a summary from a peer agent may contain text that tries to steer your model, and A2A security covers the controls. And validate data parts against their declared schema before anything acts on them; a validation failure is a bug report to the producer, not something to repair silently.
Worked example on the wire
A finance client asks a reporting agent for a Q3 variance report. Its card's skill declares text/markdown, application/json and application/pdf as output modes. The client cannot store PDFs, so it accepts only the first two. The server intersects, drops the PDF rendition, streams the Markdown summary first so a person sees progress, and appends the data part with lastChunk set.
POST /a2a (A2A-Version: 1.0)
{"jsonrpc": "2.0", "id": 7, "method": "SendStreamingMessage",
"params": {"message": {"messageId": "m-41", "role": "ROLE_USER",
"parts": [{"text": "Variance report for Q3, cost centres CC-1xx"}]},
"configuration": {"acceptedOutputModes": ["text/markdown", "application/json"]}}}
data: {"jsonrpc":"2.0","id":7,"result":{"artifactUpdate":{"taskId":"t-9","contextId":"c-2",
"artifact":{"artifactId":"q3-variance-report","parts":[{"text":"## Q3 variance\n...","mediaType":"text/markdown"}]}}}}
data: {"jsonrpc":"2.0","id":7,"result":{"artifactUpdate":{"taskId":"t-9","contextId":"c-2",
"append":true,"lastChunk":true,
"artifact":{"artifactId":"q3-variance-report","parts":[{"data":{"currency":"EUR","rows":[]},"mediaType":"application/json"}]}}}}The client's dispatcher sends the Markdown part to its renderer and validates the data part against the variance schema before writing rows to its ledger. Had the client accepted only text/csv, which the skill never offers, the server would have refused before starting work, and the client could have fallen back to another agent from its registry.
Failure modes and trade-offs
| Failure | What you see | Fix |
|---|---|---|
| JSON inside text parts | Parsers break when a model adds prose | Use data parts with a declared schema |
| Missing mediaType | Consumers guess, render bytes as text | Always set it; reject or quarantine untyped parts |
| Large raw parts | Slow streams, bloated task stores and logs | Size cap; switch to url above it |
| Unreachable or expired url | Artifact present, content missing | Scoped URLs valid for the task's retention; retry with GetTask |
| Ambiguous multi-part artifacts | Client shows three renditions as one document | Declare components versus alternatives |
| Ignored acceptedOutputModes | Work done, output dropped | Intersect before starting; refuse early |
| Schema drift | Silent field changes break consumers | Versioned media types or schema URIs; reject unknown majors |
The trade-offs are familiar from HTTP APIs. Inline parts are simple and self-contained but heavy; URL parts are light but add storage, authorisation and expiry to manage. Generic types are easy to produce and hard to trust; specific types cost a schema registry and pay back with validation. Producing several renditions widens compatibility and multiplies generation cost, so let negotiation decide which ones are worth producing.
What to do next
- Inventory every artifact your agent emits and give each an id, a media type per part, and a statement of whether its parts are components or alternatives.
- Move any JSON carried in text parts into data parts with a versioned schema URI or a specific +json media type.
- Set an inline size cap and switch larger content to short-lived, scoped URLs.
- Declare outputModes per skill on your Agent Card and implement the intersection with acceptedOutputModes, refusing early when it is empty.
- Build the consumer dispatcher: one-field check, media-type resolution, guarded fetcher, schema validation, and quarantine for unknown types.
- Add contract tests that replay recorded artifacts from each peer through the dispatcher on every release.