If you operate an image, video or audio generator, you will be asked to mark its output as AI-generated in a way machines can read. In the EU that request became law: Article 50 of the AI Act applies from 2 August 2026 and requires providers of generative systems to mark output in a machine-readable, detectable form. C2PA Content Credentials are the format most of the industry has converged on for the metadata half of that job.

This article is written from the generator operator's side. The anatomy of a C2PA manifest, COSE signatures, hard bindings and the general signing and verification pipelines are covered in Content Authentication; read that first if a manifest is new to you. Here the questions are narrower and more practical: what exactly should a generated or AI-edited asset claim, where in a generation pipeline marking belongs, how credentials can survive platforms that strip metadata, what the EU code of practice expects, and where the approach is weak.

What a generated asset should claim

A C2PA manifest for generated content says two things that matter to anyone downstream. The actions assertion says what happened, and as of specification 2.2 a standard manifest must include either a c2pa.created or a c2pa.opened action. Each action can carry a digitalSourceType, a URI from the IPTC digital source type vocabulary that says what kind of process produced the media. Choosing that term correctly is the main decision an operator makes.

PipelineActionIPTC digital source typeNotes
Text to image, text to audioc2pa.createdtrainedAlgorithmicMediaCreated using generative AI
Inpainting or outpainting a user photoc2pa.opened, then c2pa.editedcompositeWithTrainedAlgorithmicMediaEdited using generative AI; user file is an ingredient
Collage with at least one generated elementc2pa.createdcompositeSyntheticComposite including generative AI elements
Denoise or sharpen without new contentc2pa.editedalgorithmicallyEnhancedOnly if the main content is unchanged
Procedural or fractal art, no training datac2pa.createdalgorithmicMediaPure algorithmic media

The full URI takes the form http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia. Do not confuse the IPTC media terms with C2PA's own http://c2pa.org/digitalsourcetype/trainedAlgorithmicData, which describes data rather than media. Generative upscalers are the hard case: they invent detail, so labelling them as merely enhanced understates what happened. Decide a policy once and encode it, rather than letting each team choose:

IPTC = "http://cv.iptc.org/newscodes/digitalsourcetype/"

def source_type(job):
    """Map a generation job to an IPTC digital source type. One policy, in code."""
    if job.kind in ("txt2img", "txt2audio", "txt2video"):
        return IPTC + "trainedAlgorithmicMedia"
    if job.kind in ("inpaint", "outpaint", "generative_fill"):
        return IPTC + "compositeWithTrainedAlgorithmicMedia"
    if job.kind == "upscale":
        # Generative upscalers synthesise detail; treat them as generative edits.
        return IPTC + ("compositeWithTrainedAlgorithmicMedia" if job.generative
                       else "algorithmicallyEnhanced")
    if job.kind == "collage" and job.has_generated_layer:
        return IPTC + "compositeSynthetic"
    raise ValueError(f"no source-type policy for {job.kind}")

Raising an error for unknown job kinds is deliberate. A new feature that ships without a policy should fail in testing, not silently emit unmarked output.

Where marking belongs in a generation pipeline

Marking a generator pipeline: watermark first, then hash, sign and registerRequestprompt + inputsModelraw pixels or audioWatermark embedpayload = asset IDManifest builderaction + source typeSignerkey in KMS or HSMTime-stamp authorityRFC 3161 tokenhash final bytesDeliver assetmanifest embeddedManifest repositorykeyed by asset IDFingerprint indexoptional, perceptual hashcopyVerifierembedded manifest valid? else read watermark or fingerprint, look up manifest, re-check signaturestripped copylookup
Watermark before hashing, because the hard binding covers the final bytes. The repository and index make a stripped copy recoverable.

Order matters. A C2PA hard binding is a hash over the asset's bytes, so anything that changes the bytes after signing invalidates the manifest. Embed the watermark, apply any final encoding, then build, hash and sign. Keep the signing key in a KMS or HSM, request an RFC 3161 time-stamp so signatures outlive the certificate, and write a copy of the manifest to a repository keyed by an asset identifier that the watermark also carries.

With the open source c2patool CLI, a minimal manifest definition for a text-to-image output looks like this; the tool signs the file named on the command line and writes the result to the path given with -o:

{
  "claim_generator_info": [{"name": "example-image-service", "version": "3.4.0"}],
  "assertions": [
    {
      "label": "c2pa.actions",
      "data": {
        "actions": [
          {
            "action": "c2pa.created",
            "digitalSourceType": "http://cv.iptc.org/newscodes/digitalsourcetype/trainedAlgorithmicMedia",
            "softwareAgent": {"name": "example-diffusion", "version": "2026-09"}
          }
        ]
      }
    }
  ]
}

Leave prompts out of the manifest by default. A prompt can contain personal data or a customer's trade secrets, and once signed and distributed it cannot be withdrawn. Record the model identifier and version, which is what a verifier actually needs.

Durable credentials: surviving metadata stripping

Embedded manifests are fragile. Many platforms re-encode uploads and drop metadata, and a screenshot removes everything. A stripped copy is indistinguishable from an asset that never had credentials. Durable content credentials address this by pairing the manifest with a soft binding: an invisible watermark or a perceptual fingerprint that survives re-encoding and points back to a stored manifest. C2PA defines soft binding assertions for this, and version 2.2 added a supplementary specification for a Soft Binding Resolution API, so a verifier can ask a repository for manifests matching a watermark payload or fingerprint.

The verifier's logic then has three main branches, refined into four verdicts below: an embedded manifest that validates; no embedded manifest but a watermark or fingerprint that resolves to a stored manifest, whose hash will no longer match the re-encoded bytes but whose signature and claims can still be checked; or nothing found, which proves nothing. Two cautions apply. A watermark or fingerprint match is a probabilistic link, so treat a recovered manifest as strong evidence about the asset's origin, not as a byte-exact binding. And watermark detection is a secret worth protecting: an attacker with a free detector can iterate on removal until it fails.

The verifier is worth writing out, because the branch a result came from changes what you may say about it. The functions below stand for whatever SDK, watermark detector and repository client you use; the structure, not the names, is the point:

def verify(asset_bytes):
    m = read_embedded_manifest(asset_bytes)          # None if stripped
    if m is not None:
        status = validate(m, asset_bytes)            # signature, chain, time-stamp, hard binding
        if status.ok:
            return Verdict("bound", m, status)       # claims apply to these exact bytes
        return Verdict("tampered_or_invalid", m, status)

    asset_id = detect_watermark(asset_bytes)         # None, or payload with a confidence score
    candidates = []
    if asset_id is not None and asset_id.confidence >= WM_THRESHOLD:
        candidates = repository.lookup(asset_id.value)
    if not candidates:
        candidates = fingerprint_index.nearest(perceptual_hash(asset_bytes), max_distance=FP_RADIUS)

    for m in candidates:
        status = validate_signature_only(m)          # bytes changed, so skip the hard binding
        if status.ok:
            return Verdict("recovered", m, status)   # probable origin, not byte-exact
    return Verdict("unknown", None, None)            # proves nothing either way

Report the four verdicts differently in any interface. 'Bound' supports a strong statement about the file in hand. 'Tampered or invalid' means the bytes or the signature do not match what was signed, which is worth investigating but is also what an innocent re-save by a non-C2PA editor produces. 'Recovered' should be phrased as 'this appears to derive from content generated by', with the match confidence. 'Unknown' must never be shown as 'authentic'. Tune the watermark threshold and fingerprint radius on your own re-encoded test set, measuring false matches against unrelated images as well as recall.

The EU Article 50 baseline

Article 50(2) requires providers of AI systems that generate synthetic audio, image, video or text to ensure outputs are marked in a machine-readable format and detectable as artificially generated or manipulated, as far as technically feasible. Article 50(4) separately requires deployers to disclose deepfakes. On 10 June 2026 the Commission published a voluntary Code of Practice on Transparency of AI-Generated Content to support compliance. As summarised by the IPTC and law firm commentaries, it sets a multi-layered baseline: digitally signed, time-stamped metadata plus imperceptible watermarking, with fingerprinting or logging as an optional addition that cannot replace them. Reporting on the code also describes a later deadline for systems already on the market and a date for interoperable detection; check the current text for the exact dates that apply to you rather than relying on a summary.

The code does not make C2PA mandatory, and the Act is technology neutral. In practice, signed C2PA manifests are the metadata layer most providers will choose, because browsers, platforms and newsroom tools already read them. The labelling duty for deployers, including deepfake disclosure, is discussed in Deepfakes.

Text is the weak spot

C2PA works best for files: images, audio, video and documents with a container that can hold a manifest. Generated text is the weak spot. A chat answer is copied as plain characters, so there is nothing to embed a manifest in and nothing to hash once it has been edited. You can sign text at the API boundary and keep the signature server-side, which lets you answer 'did our system produce this exact string?', but it says nothing once a user changes a word.

Statistical text watermarks fill part of the gap and fail under paraphrase and translation; their limits are covered in LLM Output Provenance. For text, combine a watermark where your model supports one, server-side logs for lookup, and visible disclosure in the interface. Do not promise customers that text marking survives editing.

Worked example: generative fill on a signed photo

A photo app offers generative fill. A user uploads a holiday photo taken on a phone that signs captures with C2PA, erases a stranger in the background, and shares the result. What should the output claim?

  1. The user's photo becomes an ingredient, with its own manifest carried forward, so a verifier sees that the base image was a camera capture.
  2. The app records c2pa.opened for the ingredient and c2pa.edited with source type compositeWithTrainedAlgorithmicMedia. The output is not labelled as wholly generated, because it is not.
  3. The watermark is embedded in the final image, then the manifest is signed and time-stamped.
  4. Privacy review. The phone's manifest may carry location or device details, and ingredient thumbnails embed a copy of the original, including the stranger who was removed. C2PA allows assertions in ingredient manifests to be redacted; use that, and do not embed an ingredient thumbnail of user uploads by default.

Failure modes

  • Signing before the last byte changes. A CDN resize or format conversion after signing invalidates every manifest. Sign at the last step, or have the transforming service add its own manifest with the signed file as an ingredient.
  • Wrong source type. Labelling an inpainted photo as fully generated misleads as much as labelling it as a capture. Encode the policy, as above.
  • Leaking user data. Prompts, ingredient thumbnails and upstream manifests can expose personal data in a signed, unretractable form.
  • Key compromise. A leaked signing key lets anyone sign fakes as you. Use per-service keys, short-lived certificates and a revocation runbook.
  • Absence read as authenticity. Most media will carry no credentials for years. Never treat a missing manifest as evidence that content is real; see AI and Disinformation for how attackers exploit that assumption.
  • Watermark removal and forgery. Regeneration, heavy cropping or adversarial noise can strip a watermark, and a known scheme can sometimes be copied onto real content.

Trade-offs

ChoiceGainCost
Embedded manifest onlySimple, standard, no service to runLost on most re-uploads
Manifest plus watermark and repositoryRecoverable after stripping; matches the EU baselineRepository to operate, detector to protect
Rich manifests with promptsMore context for verifiersPrivacy and confidentiality exposure
Generative upscaling labelled as generativeHonest and conservativeMore content labelled as AI, which some users dislike

What to do next

  1. List every pipeline that produces or edits media and assign each an action and IPTC digital source type in code.
  2. Move signing to the last byte-changing step, with keys in a KMS or HSM and RFC 3161 time-stamps.
  3. Strip prompts and user identifiers from manifests by default; redact upstream ingredient assertions that carry personal data.
  4. Add an invisible watermark whose payload is an asset ID, and store every manifest in a repository keyed by it.
  5. Build a verifier that returns the four verdicts above and test it on files after upload to the main social platforms.
  6. Read the Code of Practice on Transparency of AI-Generated Content and record which commitments you meet, with dates.
  7. For text output, keep server-side signatures and logs, and add visible disclosure in the interface.
  8. Write a key compromise and revocation runbook, and rehearse it once.
Key takeaway: For a generator operator, C2PA is mostly a set of decisions: which action and digital source type each pipeline emits, signing at the last byte-changing step, keeping personal data out of signed manifests, and pairing the manifest with a watermark and repository so credentials survive stripping. Treat text separately, and never read a missing credential as proof that content is real.