When a generative system produces an image or a voice clip, two parties need to know where it came from: the people who see it, and the platforms that decide how to label it. Content authentication is the engineering that lets a file carry verifiable claims about its origin and history, and lets anyone check those claims without trusting the channel that delivered it.
This article explains the mechanisms from first principles, then builds the two pipelines you actually run: a signing pipeline inside a generation service and a verification pipeline at ingestion. Details follow the C2PA specification version 2.2 (May 2025). The emphasis is on media; provenance of LLM text output, where the options are much weaker, is covered separately.
Three questions, three mechanisms
Authentication answers three questions: who made this claim, what did they say was done to the content, and has the content changed since. It does not answer whether the content is true. A camera can sign a photograph of a staged scene; a generator can sign an image honestly labelled as synthetic. Keep that boundary in mind, because user interfaces routinely blur it.
Three mechanisms exist, and serious systems combine them.
- Cryptographic provenance metadata. A signed manifest attached to the file states the claims and binds them to the exact bytes with a hash. It is precise and verifiable offline, but it lives in metadata, which many upload pipelines strip.
- Invisible watermarks. A signal embedded in the pixels or samples survives re-encoding, resizing and screenshots to varying degrees. It carries few bits, can be attacked, and detection is probabilistic.
- Fingerprints. A perceptual hash of the content, stored by the creator, lets a verifier look the content up later. Nothing is added to the file, but lookup needs a database and near-duplicates can collide.
C2PA calls the first a hard binding (a cryptographic hash of the bytes) and the other two soft bindings, which are not proofs on their own but keys for finding a manifest that has been separated from its file.
Anatomy of a C2PA manifest
A C2PA-enabled file carries a manifest store, serialised as JUMBF (the JPEG Universal Metadata Box Format, ISO 19566-5) and embedded in the file or held externally. The store contains one or more manifests, one per step in the content's history; the last one is the active manifest.
Each manifest has three parts. Assertions are the statements: an actions assertion listing what was done, ingredient assertions pointing at the manifests of input files, and a hard-binding assertion. A standard manifest must record either a c2pa.created or a c2pa.opened action. For generative output, the action carries a digital source type; the IPTC vocabulary term for trained-model output is http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia. The claim lists hashed references to every assertion, so no assertion can be swapped afterwards. The claim signature signs the claim as a COSE_Sign1 structure with ES256, PS256 or Ed25519, using an X.509 certificate; the specification permits no other credential type.
Hard bindings depend on format: c2pa.hash.data hashes byte ranges with an exclusion for the manifest itself, c2pa.hash.boxes hashes box-structured formats, c2pa.hash.bmff.v3 covers ISO BMFF media such as MP4 with variable block sizes for streaming, and c2pa.hash.collection.data covers multi-file assets. Finally, an RFC 3161 time-stamp from a time-stamping authority proves the signature existed while the certificate was valid. Without one, a manifest stops validating when its certificate expires.
Architecture: the two pipelines
The signing pipeline lives next to the model. It watermarks the output, writes the manifest, hashes the final bytes, signs, time-stamps and embeds, and also registers the manifest and its soft-binding keys in a repository so it can be recovered later. Order matters: watermarking changes the bytes, so it must happen before the hard-binding hash, and any later re-encode invalidates the hash.
The verification pipeline lives wherever content enters: an upload service, a moderation queue, a newsroom ingest tool. It extracts the manifest store or, if there is none, tries soft-binding recovery; checks structure, hashes and signatures; evaluates the signer against a trust list; then hands a structured result to a policy layer that decides labels and routing.
Building the signing side
Treat the signing key like a code-signing key. Keep it in a KMS or HSM that performs the signature, so the private key never reaches the application host; issue a certificate per service and environment so a compromise has a small blast radius; and check that the certificate chains to a root your verifiers will trust. The C2PA trust list expects claim-signing certificates to carry a dedicated C2PA extended key usage, so a generic TLS certificate is not a substitute.
For prototyping, the open-source c2patool CLI from the Content Authenticity Initiative signs a file from a JSON manifest definition and prints the manifest report of a file you give it. The manifest definition names the signing algorithm, the certificate chain and key, an optional time-stamp authority URL, and the assertions; check field names against the version you install.
# sign: the output must use the same file type as the input
c2patool generated.jpg -m manifest.json -o generated_signed.jpg
# inspect: prints the manifest store and validation results as JSON
c2patool generated_signed.jpgIn production you call an SDK from the generation service instead. The orchestration looks like this; sdk, kms and repo stand for your C2PA library binding, key service and manifest repository.
def publish_generated(image_bytes, request, model):
marked = watermark.embed(image_bytes, payload=request.id) # before hashing
manifest = {
"claim_generator": f"imgsvc/{SERVICE_VERSION}",
"assertions": [{
"label": "c2pa.actions.v2",
"data": {"actions": [{
"action": "c2pa.created",
"digitalSourceType": IPTC_TRAINED_ALGORITHMIC_MEDIA,
"softwareAgent": {"name": model.name, "version": model.version},
}]},
}],
}
signer = kms.signer(key_id=CLAIM_KEY, alg="ES256", chain=CERT_CHAIN)
signed = sdk.sign(marked, manifest, signer, tsa_url=TSA_URL) # hash + COSE + RFC 3161
report = sdk.verify(signed) # verify what you ship
if report.state != "trusted":
raise RuntimeError(f"self-check failed: {report.errors}")
repo.put(manifest_id=report.manifest_id, manifest=report.store_bytes,
soft_keys=[watermark.key(request.id), fingerprint(marked)])
audit.log(request.id, report.manifest_id, model.version)
return signedTwo habits matter. Verify every file you sign against the same trust list your users' verifiers will use, so an expired intermediate or a missing time-stamp fails in your pipeline rather than in theirs. And log the manifest identifier with the request in an audit log, so a disputed file can be traced to the request, the model version and the user.
Building the verification side
The specification defines three outcome states, and the distinction is the most important thing to get right in an interface.
| State | Meaning | What it does not mean |
|---|---|---|
| Well-formed | The manifest parses and follows the structural rules | Anyone signed it, or the bytes match |
| Valid | Well-formed, hashes match the content, signature verifies, credential not revoked | The signer is anyone you have heard of |
| Trusted | Valid, and the signing certificate chains to a root on the verifier's trust list | The claims are true, or the scene really happened |
A verifier works through those states in order and records every failure code instead of stopping at the first, because the policy layer needs to tell a re-encoded file (hash mismatch, signature fine) from a forged one (signature fails). It then recurses into ingredients: an edited image's active manifest may be trusted while an ingredient's manifest is missing or invalid, and the ingredient chain is where you learn that the edited image started life as model output.
def assess(file_bytes):
store = sdk.read_store(file_bytes)
if store is None:
store = recover(file_bytes) # soft binding -> repository
if store is None:
return Verdict(state="none", synthetic=None)
r = sdk.validate(store, file_bytes, trust_list=TRUST_LIST)
synthetic = any(a.digital_source_type in AI_SOURCE_TYPES
for m in r.manifests for a in m.actions)
return Verdict(state=r.state, synthetic=synthetic,
signer=r.signer_subject, errors=r.failure_codes,
recovered=store.recovered)The policy layer maps verdicts to actions: show a provenance label for trusted manifests, show the signer but no badge for valid ones, treat hash mismatches as "edited after signing", and never treat the absence of a manifest as evidence that content is authentic. Most content on the internet carries no manifest at all.
Worked example: a credential that survives a platform
A user generates a product image. The service watermarks it with a payload keyed to the request, signs a manifest stating c2pa.created with the trained-media source type, time-stamps it and stores the manifest in its repository. The user opens the image in an editor that supports C2PA; the editor adds a manifest of its own recording a crop and colour change, listing the original as an ingredient, and signs with the editor vendor's key.
The user then posts the image to a platform that re-encodes uploads and strips metadata. The bytes now match neither manifest and no manifest store is present. The platform's verifier finds no store, runs the watermark detector, decodes the payload, queries the repository, and recovers the generator's manifest. It cannot validate the hard binding against these bytes, so it reports "recovered by soft binding" rather than "valid", but it can show that a manifest exists for content matching this image and that the manifest declares it AI-generated. C2PA describes this combination of embedded manifest, watermark and repository as durable Content Credentials. The editor's manifest is lost unless the editor also registered it, which is a reminder that each signer in a chain has to publish its own record.
Failure modes and attacks
- Stripping. Metadata removal is the default behaviour of many pipelines, not an attack. Plan for soft-binding recovery from day one.
- Badge on the wrong state. An interface that shows a check mark for a merely well-formed manifest lets anyone mint one with a self-signed certificate. Badge only trusted results, and show the signer's name.
- Trusted is not true. Re-photographing a screen, or signing a staged scene with a genuine device, yields a trusted manifest for misleading content. Provenance tells you who vouches; forensic analysis, covered in AI forensics, is a separate discipline.
- Key compromise. A stolen claim-signing key lets an attacker sign anything as you. Keep keys in hardware, scope certificates narrowly, monitor signing volume per key, and have a revocation runbook.
- Missing time-stamps. Without RFC 3161 time-stamps, every manifest you ever signed becomes invalid when the certificate expires. C2PA update manifests can add time-stamps and revocation information later, but only while you still can.
- Watermark attacks. Regeneration through another model, heavy cropping or adversarial noise can remove a watermark, and an attacker who learns the scheme may forge one onto real content. Treat a detection as a lookup hint with a false-positive rate you have measured, never as proof.
Text, regulation and trade-offs
Everything above assumes media with room for metadata and signal. Text is harder: a manifest cannot travel inside a pasted paragraph, and statistical text watermarks weaken under paraphrase. Provenance for LLM text output relies on different tools, described in LLM output provenance architecture.
Regulation now pushes in this direction. Article 50 of the EU AI Act requires providers of generative systems to mark synthetic output in a machine-readable, detectable way, and its transparency obligations apply from 2 August 2026; reporting on the Digital Omnibus describes a later deadline for systems already on the market, so check the current text for your situation. See the EU AI Act guide for scope. The law states the outcome, not the mechanism; manifests plus watermarks is the common engineering answer.
| Mechanism | Strength | Weakness | Use it for |
|---|---|---|---|
| Signed manifest | Exact, offline, attributable | Stripped by many pipelines | Primary record of origin and edits |
| Watermark | Survives re-encoding and screenshots | Few bits, removable, probabilistic | Recovering the manifest, flagging |
| Fingerprint | Nothing added to the file | Needs a lookup service, collisions | Recovery when watermarking is not possible |
What to do next
- Decide which outputs you sign and which states your interface will badge; write it down before any code.
- Obtain a claim-signing certificate from a CA on the C2PA trust list, held in a KMS or HSM, and configure an RFC 3161 time-stamp authority.
- Prototype with c2patool on sample outputs, then move signing into the service with a verify-after-sign self-check.
- Add a watermark or fingerprint and a manifest repository so stripped files can be recovered.
- Build the verifier at ingestion with all three states, ingredient recursion and failure codes, and test it on stripped, re-encoded, self-signed and expired-certificate files.
- Write the key compromise and revocation runbook, and log every manifest identifier with its request.