Generative models now produce images, video and voices that ordinary viewers cannot reliably tell apart from capture. The generator's provider can mark output at creation time, but most viewers never meet the generator: they meet the file on a social feed, a news site, a marketplace listing or a messaging app. Whoever publishes or hosts the file has to decide what label, if any, the viewer sees, and has to make that decision from signals that are often missing, stripped or ambiguous.
This article is written from that downstream side: the platform, publisher or enterprise that hosts and displays media made by others or by its own tools. It covers who has to label what under the EU AI Act, a label taxonomy that avoids the most common mistake, the signals available and how far to trust each, a label decision engine with code, the base-rate arithmetic that keeps classifiers in their place, label persistence and appeals, and failure modes. How a generator writes C2PA manifests and watermarks is covered in a sibling article; here we consume them. None of this is legal advice.
Who has to label what
The EU AI Act splits the duty by role, and its transparency rules in Article 50 have applied since 2 August 2026. Article 50(2) obliges providers of systems that generate synthetic audio, image, video or text to mark outputs in a machine-readable format, detectable as artificially generated or manipulated, as far as technically feasible. It exempts systems that perform an assistive function for standard editing or do not substantially alter the input. The AI Digital Omnibus gave generators already on the market before 2 August 2026 until 2 December 2026 to meet the marking duty; systems placed on the market later get no grace period.
Article 50(4) obliges deployers, the organisations using such systems, to disclose deepfakes: generated or manipulated image, audio or video that resembles real people, objects, places or events and would falsely appear authentic. For evidently artistic, satirical or fictional work, the disclosure can be limited to an appropriate notice that does not spoil the work. Deployers publishing AI-generated text to inform the public on matters of public interest must also disclose it, unless the text went through human review or editorial control and someone holds editorial responsibility for publication.
So a newsroom that publishes a generated illustration is a deployer with a disclosure duty, and a platform that hosts user uploads labels for policy and trust reasons even where the statute does not reach it directly. Platform policies already work this way. YouTube asks creators to disclose realistic altered or synthetic content and can apply a label itself, and Meta and TikTok read industry metadata such as C2PA to label automatically. Treat this summary as orientation; check the consolidated text for your role.
A taxonomy that says only what is proven
The most common labelling mistake is a single binary flag. In May 2024 Meta began applying a "Made with AI" label partly from C2PA and IPTC metadata. Photographers objected that ordinary photos retouched with AI-assisted editing tools were tagged as if generated, and on 1 July 2024 Meta renamed the label "AI info". The detection did not change; the wording did, because the label had claimed more than the evidence supported. Build the taxonomy so the label says only what the signals prove.
| Label | Meaning | Typical evidence | Viewer sees |
|---|---|---|---|
| AI-generated | Created wholly by a generative model | Signed manifest with trainedAlgorithmicMedia; vendor watermark; declaration | Prominent badge |
| AI-modified | Real capture with generated regions or content | Manifest with a composite source type; declaration | Badge naming the edit |
| AI-assisted edit | Standard edits made with AI tools | Manifest actions without a generative source type | Details panel only |
| Verified capture | Signed from a capture device, edits recorded | Valid capture manifest, intact chain | Optional details |
| Unknown | No usable signal | Nothing, or stripped metadata | No label |
Two rules follow. "Unknown" is the default and is never rendered as "authentic"; most real photos carry no credentials and most stripped AI images look exactly like them. And only a deepfake-style label, AI-generated or AI-modified depicting a real person or event, needs the prominent treatment that Article 50(4) has in mind.
The signals and how far to trust them
Six signal types are available at ingest. They differ in how strongly they prove generation and in how easily they disappear.
| Signal | Strength when present | How it fails |
|---|---|---|
| Signed C2PA manifest | Strong: cryptographic claim by the tool that made it | Stripped by re-encoding or screenshots; signer may be untrusted |
| Unsigned IPTC or XMP field | Moderate: a declaration anyone can edit | Trivially added or removed |
| Invisible watermark | Strong for the vendor that embeds and detects it | Each detector covers one vendor; heavy edits weaken it |
| Fingerprint match | Strong for items already known | Only covers media you or a partner have seen |
| Uploader declaration | Strong positive; weak negative | Bad actors simply do not declare |
| Classifier score | Weak, probabilistic | False positives on edited real photos; drifts with new generators |
Signals are asymmetric: their presence is informative, their absence is not. A stripped screenshot of a generated image carries no manifest and no metadata, and possibly a damaged watermark. Design the engine so missing signals push towards unknown, not towards authentic.
The label decision engine
The engine is a precedence list, not a weighted sum, because the signals are not comparable quantities: a valid signature from a trusted generator is a different kind of evidence from a 0.8 classifier score. Signed manifests come first, then a vendor watermark above a tuned threshold, then fingerprint matches, then the uploader's own declaration. Classifier scores never apply a label by themselves; above a high threshold, or a lower one when the media depicts a real person, they send the item to human review.
# label_engine.py - turn ingest signals into one label plus evidence
from dataclasses import dataclass, field
GENERATIVE = {"trainedAlgorithmicMedia"}
COMPOSITE = {"compositeWithTrainedAlgorithmicMedia"}
WM_THRESHOLD = 0.95 # per-vendor detector, tuned on re-encoded test sets
CLF_REVIEW = 0.90 # classifier only ever routes to review
@dataclass
class Signals:
c2pa_valid: bool = False
c2pa_trusted_signer: bool = False
source_types: set = field(default_factory=set) # IPTC URIs from manifest actions
unsigned_source_type: str | None = None
watermark_scores: dict = field(default_factory=dict) # vendor -> score
fingerprint_hit: bool = False
declared_ai: bool = False
classifier: float = 0.0
depicts_real_person: bool = False
def decide(s: Signals, engine_version="label-v7"):
evidence = []
terms = {t.rsplit("/", 1)[-1] for t in s.source_types} # full URI to term
if s.c2pa_valid and s.c2pa_trusted_signer and terms & GENERATIVE:
label = "ai_generated"; evidence.append("c2pa:generative")
elif s.c2pa_valid and s.c2pa_trusted_signer and terms & COMPOSITE:
label = "ai_modified"; evidence.append("c2pa:composite")
elif any(v >= WM_THRESHOLD for v in s.watermark_scores.values()):
label = "ai_generated"; evidence.append(f"watermark:{max(s.watermark_scores, key=s.watermark_scores.get)}")
elif s.fingerprint_hit:
label = "ai_generated"; evidence.append("fingerprint")
elif s.declared_ai:
label = "ai_generated"; evidence.append("declared")
elif s.c2pa_valid and s.c2pa_trusted_signer:
label = "verified_or_assisted"; evidence.append("c2pa:no_generative_action")
else:
label = "unknown"
if (s.unsigned_source_type or "").rsplit("/", 1)[-1] in GENERATIVE:
evidence.append("unsigned_metadata:generative") # shown in details, not a badge
review = label == "unknown" and (s.classifier >= CLF_REVIEW or s.depicts_real_person and s.classifier >= 0.7)
prominent = label in ("ai_generated", "ai_modified") and s.depicts_real_person
return {"label": label, "prominent": prominent, "review": review,
"evidence": evidence, "classifier": s.classifier, "engine": engine_version}The output carries the evidence and the engine version, and both are logged. When a label is disputed, support staff can see exactly which signal produced it, and when you change a threshold you can recompute which past decisions would have changed. Map the IPTC source type terms from the vocabulary you actually receive, and test against real manifests from the generators your users rely on.
Worked example: classifier base rates
Why not let a good classifier label automatically? Base rates. Suppose a platform ingests 1,000,000 images a day, 3% of which are generated, so 30,000 synthetic and 970,000 real. A classifier with 90% recall and a 2% false positive rate flags 27,000 synthetic images and 19,400 real ones. Precision is 27,000 divided by 46,400, about 58%: more than four in ten labels land on real photos, and a disproportionate share of those are heavily edited professional work, the same failure that forced Meta's rename.
Raising the threshold so the false positive rate falls to 0.2% might drop recall to 60%. Now 18,000 true and 1,940 false flags give precision near 90%, and the daily volume of about 20,000 is small enough to put through review. The detector is then a triage tool that decides what humans look at, which is the role its error rates support. Re-measure both rates every time a major new generator ships, because recall on unseen generators is usually much lower than on the training distribution.
Keeping labels attached
A label is only as durable as the pipeline that carries it. Three engineering habits matter. First, preserve credentials through your own processing: many image and video pipelines strip metadata during resizing and transcoding, destroying the manifest you just verified. Either keep the manifest attached or store it and re-attach a manifest of your own that records the transformation. Second, store the decision with a perceptual hash of the item, so a re-upload or re-share of the same media inherits the label even after its metadata is gone. Third, carry the label through shares, embeds and API responses. A badge that disappears when the post is embedded on another site protects nobody.
In the interface, keep the badge short and put the evidence one tap away: which signal, from which tool, and what it does not prove. "Contains AI-generated content, according to content credentials from the creating app" is accurate. "Fake" is not a label; it is an accusation the evidence rarely supports.
Declarations and appeals
Creators need two paths. At upload, a declaration toggle that the engine trusts as a positive signal, with clear guidance on when realistic content needs it. After labelling, an appeal that shows the creator the evidence class and lets them submit originals or a capture manifest. Route appeals for classifier-originated reviews to people who did not make the first call, and track the overturn rate per signal; a signal whose labels are often overturned is miscalibrated. Label appeals should never remove a label backed by a valid generative manifest; they can correct its category, for example from AI-generated to AI-assisted edit when the manifest shows only standard edits.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Binary label from any AI action | Edited real photos tagged as generated | Separate generated, modified and assisted edits |
| Absence read as authentic | Stripped fakes get a trust badge | Default to unknown; never show authentic without a capture manifest |
| Classifier auto-labels | Complaints from photographers and news outlets | Classifier routes to review only |
| Own pipeline strips manifests | Verified uploads lose credentials on display | Preserve or re-sign through transcoding |
| Untrusted signer accepted | Forged manifests claim provenance | Validate against a maintained trust list |
| Label lost on re-share | Same fake recirculates unlabelled | Perceptual-hash index of past decisions |
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Precedence rules over scores | Explainable, auditable labels | Misses weak combined evidence |
| Classifier to review only | Few false accusations | Review staffing; some fakes unlabelled |
| Fine-grained taxonomy | Labels claim only what is proven | More UI and policy to explain |
| Hash-indexed label memory | Labels survive re-uploads | Index storage; hash collisions to tune |
What to do next
- Decide your role for each media flow: provider, deployer or host, and which Article 50 duties attach to it.
- Adopt a five-state taxonomy with unknown as the default and no authentic badge without a capture manifest.
- Verify C2PA manifests at ingest against a trust list, and run the watermark detectors you have access to.
- Implement the decision engine as precedence rules, logging evidence and engine version with every decision.
- Measure your classifier's precision at your real base rate and limit it to routing review.
- Audit your media pipeline for metadata stripping, add a perceptual-hash label index, and publish an appeal path.
Keep learning: C2PA for AI-generated content for the generator side of marking, content authentication for signing and verification pipelines, deepfakes for impersonation controls, and the EU AI Act for roles and timelines.