Generative models now produce images, video and voices that ordinary viewers cannot reliably tell apart from capture. The generator's provider can mark output at creation time, but most viewers never meet the generator: they meet the file on a social feed, a news site, a marketplace listing or a messaging app. Whoever publishes or hosts the file has to decide what label, if any, the viewer sees, and has to make that decision from signals that are often missing, stripped or ambiguous.

This article is written from that downstream side: the platform, publisher or enterprise that hosts and displays media made by others or by its own tools. It covers who has to label what under the EU AI Act, a label taxonomy that avoids the most common mistake, the signals available and how far to trust each, a label decision engine with code, the base-rate arithmetic that keeps classifiers in their place, label persistence and appeals, and failure modes. How a generator writes C2PA manifests and watermarks is covered in a sibling article; here we consume them. None of this is legal advice.

Who has to label what

The EU AI Act splits the duty by role, and its transparency rules in Article 50 have applied since 2 August 2026. Article 50(2) obliges providers of systems that generate synthetic audio, image, video or text to mark outputs in a machine-readable format, detectable as artificially generated or manipulated, as far as technically feasible. It exempts systems that perform an assistive function for standard editing or do not substantially alter the input. The AI Digital Omnibus gave generators already on the market before 2 August 2026 until 2 December 2026 to meet the marking duty; systems placed on the market later get no grace period.

Article 50(4) obliges deployers, the organisations using such systems, to disclose deepfakes: generated or manipulated image, audio or video that resembles real people, objects, places or events and would falsely appear authentic. For evidently artistic, satirical or fictional work, the disclosure can be limited to an appropriate notice that does not spoil the work. Deployers publishing AI-generated text to inform the public on matters of public interest must also disclose it, unless the text went through human review or editorial control and someone holds editorial responsibility for publication.

So a newsroom that publishes a generated illustration is a deployer with a disclosure duty, and a platform that hosts user uploads labels for policy and trust reasons even where the statute does not reach it directly. Platform policies already work this way. YouTube asks creators to disclose realistic altered or synthetic content and can apply a label itself, and Meta and TikTok read industry metadata such as C2PA to label automatically. Treat this summary as orientation; check the consolidated text for your role.

A taxonomy that says only what is proven

The most common labelling mistake is a single binary flag. In May 2024 Meta began applying a "Made with AI" label partly from C2PA and IPTC metadata. Photographers objected that ordinary photos retouched with AI-assisted editing tools were tagged as if generated, and on 1 July 2024 Meta renamed the label "AI info". The detection did not change; the wording did, because the label had claimed more than the evidence supported. Build the taxonomy so the label says only what the signals prove.

LabelMeaningTypical evidenceViewer sees
AI-generatedCreated wholly by a generative modelSigned manifest with trainedAlgorithmicMedia; vendor watermark; declarationProminent badge
AI-modifiedReal capture with generated regions or contentManifest with a composite source type; declarationBadge naming the edit
AI-assisted editStandard edits made with AI toolsManifest actions without a generative source typeDetails panel only
Verified captureSigned from a capture device, edits recordedValid capture manifest, intact chainOptional details
UnknownNo usable signalNothing, or stripped metadataNo label

Two rules follow. "Unknown" is the default and is never rendered as "authentic"; most real photos carry no credentials and most stripped AI images look exactly like them. And only a deepfake-style label, AI-generated or AI-modified depicting a real person or event, needs the prominent treatment that Article 50(4) has in mind.

The signals and how far to trust them

Six signal types are available at ingest. They differ in how strongly they prove generation and in how easily they disappear.

SignalStrength when presentHow it fails
Signed C2PA manifestStrong: cryptographic claim by the tool that made itStripped by re-encoding or screenshots; signer may be untrusted
Unsigned IPTC or XMP fieldModerate: a declaration anyone can editTrivially added or removed
Invisible watermarkStrong for the vendor that embeds and detects itEach detector covers one vendor; heavy edits weaken it
Fingerprint matchStrong for items already knownOnly covers media you or a partner have seen
Uploader declarationStrong positive; weak negativeBad actors simply do not declare
Classifier scoreWeak, probabilisticFalse positives on edited real photos; drifts with new generators

Signals are asymmetric: their presence is informative, their absence is not. A stripped screenshot of a generated image carries no manifest and no metadata, and possibly a damaged watermark. Design the engine so missing signals push towards unknown, not towards authentic.

The label decision engine

Label decision engine: independent signals in, one label and an evidence record outC2PA manifestsigned source typeUnsigned metadataIPTC, XMP fieldsWatermark detectorsper vendor, scoredFingerprint matchknown generated itemsUploader declarationtoggle at uploadClassifier scoreweakest, base-rate boundDecision engineprecedence + thresholdslabelreviewevidenceVisible labelbadge + details panelHuman review queueclassifier-only, appealsDecision logsignals, version, outcomeAbsence of every signal means unknown, never authentic. A classifier alone never auto-applies a definitive label.
Signals feed a precedence-ordered decision engine. Its outputs are a visible label, an optional human-review task and a logged evidence record.

The engine is a precedence list, not a weighted sum, because the signals are not comparable quantities: a valid signature from a trusted generator is a different kind of evidence from a 0.8 classifier score. Signed manifests come first, then a vendor watermark above a tuned threshold, then fingerprint matches, then the uploader's own declaration. Classifier scores never apply a label by themselves; above a high threshold, or a lower one when the media depicts a real person, they send the item to human review.

# label_engine.py - turn ingest signals into one label plus evidence
from dataclasses import dataclass, field

GENERATIVE = {"trainedAlgorithmicMedia"}
COMPOSITE = {"compositeWithTrainedAlgorithmicMedia"}
WM_THRESHOLD = 0.95          # per-vendor detector, tuned on re-encoded test sets
CLF_REVIEW = 0.90            # classifier only ever routes to review

@dataclass
class Signals:
    c2pa_valid: bool = False
    c2pa_trusted_signer: bool = False
    source_types: set = field(default_factory=set)   # IPTC URIs from manifest actions
    unsigned_source_type: str | None = None
    watermark_scores: dict = field(default_factory=dict)  # vendor -> score
    fingerprint_hit: bool = False
    declared_ai: bool = False
    classifier: float = 0.0
    depicts_real_person: bool = False

def decide(s: Signals, engine_version="label-v7"):
    evidence = []
    terms = {t.rsplit("/", 1)[-1] for t in s.source_types}   # full URI to term
    if s.c2pa_valid and s.c2pa_trusted_signer and terms & GENERATIVE:
        label = "ai_generated"; evidence.append("c2pa:generative")
    elif s.c2pa_valid and s.c2pa_trusted_signer and terms & COMPOSITE:
        label = "ai_modified"; evidence.append("c2pa:composite")
    elif any(v >= WM_THRESHOLD for v in s.watermark_scores.values()):
        label = "ai_generated"; evidence.append(f"watermark:{max(s.watermark_scores, key=s.watermark_scores.get)}")
    elif s.fingerprint_hit:
        label = "ai_generated"; evidence.append("fingerprint")
    elif s.declared_ai:
        label = "ai_generated"; evidence.append("declared")
    elif s.c2pa_valid and s.c2pa_trusted_signer:
        label = "verified_or_assisted"; evidence.append("c2pa:no_generative_action")
    else:
        label = "unknown"
        if (s.unsigned_source_type or "").rsplit("/", 1)[-1] in GENERATIVE:
            evidence.append("unsigned_metadata:generative")   # shown in details, not a badge
    review = label == "unknown" and (s.classifier >= CLF_REVIEW or s.depicts_real_person and s.classifier >= 0.7)
    prominent = label in ("ai_generated", "ai_modified") and s.depicts_real_person
    return {"label": label, "prominent": prominent, "review": review,
            "evidence": evidence, "classifier": s.classifier, "engine": engine_version}

The output carries the evidence and the engine version, and both are logged. When a label is disputed, support staff can see exactly which signal produced it, and when you change a threshold you can recompute which past decisions would have changed. Map the IPTC source type terms from the vocabulary you actually receive, and test against real manifests from the generators your users rely on.

Worked example: classifier base rates

Why not let a good classifier label automatically? Base rates. Suppose a platform ingests 1,000,000 images a day, 3% of which are generated, so 30,000 synthetic and 970,000 real. A classifier with 90% recall and a 2% false positive rate flags 27,000 synthetic images and 19,400 real ones. Precision is 27,000 divided by 46,400, about 58%: more than four in ten labels land on real photos, and a disproportionate share of those are heavily edited professional work, the same failure that forced Meta's rename.

Raising the threshold so the false positive rate falls to 0.2% might drop recall to 60%. Now 18,000 true and 1,940 false flags give precision near 90%, and the daily volume of about 20,000 is small enough to put through review. The detector is then a triage tool that decides what humans look at, which is the role its error rates support. Re-measure both rates every time a major new generator ships, because recall on unseen generators is usually much lower than on the training distribution.

Keeping labels attached

A label is only as durable as the pipeline that carries it. Three engineering habits matter. First, preserve credentials through your own processing: many image and video pipelines strip metadata during resizing and transcoding, destroying the manifest you just verified. Either keep the manifest attached or store it and re-attach a manifest of your own that records the transformation. Second, store the decision with a perceptual hash of the item, so a re-upload or re-share of the same media inherits the label even after its metadata is gone. Third, carry the label through shares, embeds and API responses. A badge that disappears when the post is embedded on another site protects nobody.

In the interface, keep the badge short and put the evidence one tap away: which signal, from which tool, and what it does not prove. "Contains AI-generated content, according to content credentials from the creating app" is accurate. "Fake" is not a label; it is an accusation the evidence rarely supports.

Declarations and appeals

Creators need two paths. At upload, a declaration toggle that the engine trusts as a positive signal, with clear guidance on when realistic content needs it. After labelling, an appeal that shows the creator the evidence class and lets them submit originals or a capture manifest. Route appeals for classifier-originated reviews to people who did not make the first call, and track the overturn rate per signal; a signal whose labels are often overturned is miscalibrated. Label appeals should never remove a label backed by a valid generative manifest; they can correct its category, for example from AI-generated to AI-assisted edit when the manifest shows only standard edits.

Failure modes

FailureSymptomFix
Binary label from any AI actionEdited real photos tagged as generatedSeparate generated, modified and assisted edits
Absence read as authenticStripped fakes get a trust badgeDefault to unknown; never show authentic without a capture manifest
Classifier auto-labelsComplaints from photographers and news outletsClassifier routes to review only
Own pipeline strips manifestsVerified uploads lose credentials on displayPreserve or re-sign through transcoding
Untrusted signer acceptedForged manifests claim provenanceValidate against a maintained trust list
Label lost on re-shareSame fake recirculates unlabelledPerceptual-hash index of past decisions

Trade-offs

ChoiceGainCost
Precedence rules over scoresExplainable, auditable labelsMisses weak combined evidence
Classifier to review onlyFew false accusationsReview staffing; some fakes unlabelled
Fine-grained taxonomyLabels claim only what is provenMore UI and policy to explain
Hash-indexed label memoryLabels survive re-uploadsIndex storage; hash collisions to tune

What to do next

  1. Decide your role for each media flow: provider, deployer or host, and which Article 50 duties attach to it.
  2. Adopt a five-state taxonomy with unknown as the default and no authentic badge without a capture manifest.
  3. Verify C2PA manifests at ingest against a trust list, and run the watermark detectors you have access to.
  4. Implement the decision engine as precedence rules, logging evidence and engine version with every decision.
  5. Measure your classifier's precision at your real base rate and limit it to routing review.
  6. Audit your media pipeline for metadata stripping, add a perceptual-hash label index, and publish an appeal path.

Keep learning: C2PA for AI-generated content for the generator side of marking, content authentication for signing and verification pipelines, deepfakes for impersonation controls, and the EU AI Act for roles and timelines.

Key takeaway: Labelling AI-generated media is a decision under uncertainty. Use a taxonomy that separates generated, modified and assisted edits. Trust signed and watermarked evidence first, keep classifiers to triage, treat missing signals as unknown, and log the evidence behind every label so it can survive re-uploads and appeals.