Most writing on AI and copyright is about training: whether copying works into a training set is lawful, and what the model memorised. That side is covered in copyright and AI training. This article is about the other end of the pipe. When your team ships text, images or code that a model produced, two questions follow. Can anyone own it? And could shipping it infringe someone else's rights, for example by reproducing licensed source code?

Both questions have engineering answers, because both turn on evidence that only your systems can keep: who wrote which part, which model produced what, and what the output was checked against. This is a practical guide to that evidence, not legal advice; the law is still moving, and anything consequential needs counsel in each jurisdiction where you sell.

The human authorship rule, as of October 2026

The position as of October 2026, for the points that drive engineering decisions:

SourceWhat it says
US Copyright Office, Part 2 report on copyrightability (January 2025)Copyright requires human authorship. Given current generally available technology, prompts alone do not provide sufficient human control to make the user the author of the output. Human expression that is perceptible in the output, creative selection and arrangement, and creative modifications can be protected.
Thaler v. Perlmutter (US)Courts upheld the refusal to register a work listing an AI system as its sole author. The Supreme Court denied certiorari on 2 March 2026, leaving the human-authorship requirement in place.
US registration practiceApplicants must disclose AI-generated material that is more than de minimis and claim only the human contribution.
China, Beijing Internet Court (2023)Granted protection to an AI-generated image where the user's prompting and parameter choices were found to reflect intellectual effort. Other courts have reached different results.
UK, CDPA section 9(3)Protects computer-generated works with no human author, naming the person who made the arrangements as author. The government consulted on its future in 2024 and 2025.
EU AI Act, Article 50Transparency duties from 2 August 2026, including machine-readable marking of synthetic output by providers. The marking duty in Article 50(2) applies from 2 December 2026 for systems already on the market before 2 August 2026.

The common thread across most jurisdictions is that protection follows human creative control. The more of the expression a person decided, the stronger the claim; the more the model decided, the weaker it is.

What is protectable in a mixed work

Few real works are purely human or purely generated. A product page might have human-written copy, a generated hero image, a human-designed layout and generated alt text. Under the US approach the protectable parts are the human ones: the copy, the layout as a selection and arrangement, and any substantial human edits to the image. The raw image is not protected, which means a competitor could copy it.

That has practical consequences. You cannot rely on copyright to stop others reusing unedited generated assets, so if exclusivity matters, use trade marks, contracts or trade secrets, or invest in human modification. Registration filings must describe the AI material accurately, and an inaccurate application can undermine the registration. Contracts in which you warrant that deliverables are original and owned may be untrue for generated parts. And an open-source licence you attach to generated code assumes rights you may not hold, though the licence still works as a statement of terms for the human-authored parts.

An output-side pipeline

An output-side copyright pipeline: evidence in, checks in the middle, labels outGeneration gatewaymodel, prompt, output hashHuman editsdiffs, selection, authorshipContribution ledgerwho made which partLicence scancode fingerprintsSimilarity checktext and imagesTerms checkvendor and licenceRelease recordlabels, disclosuresTakedown loopclaims, removalscomplaintsThe ledger answers two questions later: what can we claim, and where did this come from.
Evidence flows from the generation gateway and the editing tools into one ledger; checks run before release, and complaints feed back into removals.

The architecture has three layers. Capture records every generation and every human edit. Checks run before anything ships: licence fingerprints for code, similarity for text and images, and a terms check. Release writes a record with the labels and disclosures attached, and a takedown loop handles claims afterwards. Keep it lightweight; most of the value comes from capture, which is cheap if it is built into the tools people already use.

The contribution ledger

The contribution ledger records, per asset, which spans came from which source. Route all model calls through a gateway that logs the model, version, prompt and an output hash; most teams already have one for cost and safety, as described in LLM data governance. Then record human edits as diffs against the generated version.

CREATE TABLE generation (
  gen_id       uuid PRIMARY KEY,
  model        text NOT NULL,          -- provider/model/version
  prompt_hash  bytea NOT NULL,         -- keep full prompts where policy allows
  output_hash  bytea NOT NULL,
  created_by   text NOT NULL,
  created_at   timestamptz NOT NULL
);
CREATE TABLE contribution (
  contribution_id bigserial PRIMARY KEY,
  asset_id     text NOT NULL,          -- file path, document id or image id
  asset_rev    text NOT NULL,          -- commit or revision
  source       text NOT NULL,          -- 'human:<user>' or 'gen:<gen_id>'
  span_start   int, span_end int,      -- character or line span; NULL for whole asset
  note         text                    -- e.g. 'selected 3 of 12 variants, recoloured'
);
CREATE INDEX contribution_asset ON contribution (asset_id, asset_rev);

For text and code, a diff gives a rough measure of how much of the final version survived from the generated draft:

import difflib

def generated_share(generated: str, final: str) -> float:
    """Fraction of the final text's characters that match the generated draft."""
    sm = difflib.SequenceMatcher(None, generated, final, autojunk=False)
    kept = sum(block.size for block in sm.get_matching_blocks())
    return kept / max(1, len(final))

Treat this number as evidence, not as a legal test. No court protects a work because 41 percent was human; what matters is whether the human contribution was creative expression. The ledger's job is to let you show what the person actually did: the drafts, the selections, the rewrites, and when.

Generated code and open-source licences

Generated code raises a sharper risk. A model can reproduce snippets of licensed code from its training data, and if a substantial copyleft snippet lands in proprietary code, the obligations of that licence, or an infringement claim, may follow. The risk is highest for long, distinctive blocks such as well-known algorithms copied with their original comments, and lowest for short idiomatic lines that anyone would write.

Controls stack. First, if your assistant offers a setting that blocks suggestions matching public code, turn it on for proprietary repositories. Second, scan generated code before merge against an index of known open-source code. Commercial snippet scanners do this at scale; the core technique, winnowing, is small enough to understand and to prototype:

import hashlib, re

def tokens(source: str) -> list[str]:
    """Normalise code: drop comments and whitespace differences, keep identifiers and symbols."""
    source = re.sub(r"//[^\n]*|#[^\n]*|/\*.*?\*/", " ", source, flags=re.S)
    return re.findall(r"[A-Za-z_]\w*|\d+|\S", source)

def fingerprints(source: str, k: int = 12, window: int = 8) -> set[int]:
    """Winnowing: hash every k-token shingle, keep the minimum hash in each window."""
    toks = tokens(source)
    hashes = [int.from_bytes(hashlib.blake2b(" ".join(toks[i:i + k]).encode(), digest_size=8).digest(), "big")
              for i in range(len(toks) - k + 1)]
    return {min(hashes[i:i + window]) for i in range(max(0, len(hashes) - window + 1))}

def overlap(candidate: str, known: set[int]) -> float:
    fp = fingerprints(candidate)
    return len(fp & known) / len(fp) if fp else 0.0

Fingerprints survive reformatting and comment changes because the tokens are normalised; in a quick test, a re-indented copy with an added comment scored 1.0 against the original, a copy with one identifier renamed throughout scored 0.78, and unrelated code scored about 0.01. Index the fingerprints of the licensed corpora you care about, keyed to their licence, and fail the merge when overlap with a copyleft source exceeds a threshold you set from your own false-positive rate. Third, run a licence detector over dependencies and vendored files, because copied files often arrive with their headers intact.

For text and images the equivalent check is regurgitation and near-duplicate detection, covered in the training-side article.

Contracts, terms and indemnities

Your rights in generated output also depend on contracts. Most commercial model providers' terms assign or disclaim rights in outputs to the customer, to the extent any rights exist; they cannot grant protection the law does not give. Several vendors offer copyright indemnities on paid tiers, typically conditioned on using their safety filters and not deliberately prompting for infringing material. The conditions change, so read the current terms for each provider, record which tier and settings each team uses, and keep the gateway logs that would prove you met the conditions if a claim arrives.

Open-weight models add their own licences, some of which restrict use cases or require attribution. Record the model licence in the generation table alongside the model name, because the obligation travels with the output in some licences and not in others.

Labelling and provenance

Labelling is now partly a legal duty. Under Article 50 of the EU AI Act, providers of systems that generate synthetic audio, images, video or text must mark outputs in a machine-readable, detectable way, and deployers must disclose deepfakes and certain AI-generated text published to inform the public. Content credentials, signed provenance manifests embedded in media files, are one technical approach to marking; provenance and signing covers the mechanics and the EU AI Act guide covers the roles and timetable.

The same metadata serves copyright. A release record that says which parts are generated, by which model, with which human contribution, is what a registration filing, a customer's due diligence or a takedown response needs. Build it once.

Worked example: a guide and a feature

A company publishes an illustrated developer guide and ships a feature in the same sprint. For the guide, writers draft with a model, then rewrite; an illustrator generates twelve candidate diagrams per chapter, selects three, and redraws labels and layout. The ledger shows the generated drafts, the diffs (most chapters retain under a fifth of the generated text), and the selections and redraws. The registration application claims the text, the selection and arrangement of illustrations and the redrawn elements, and disclaims the raw generated images. The release record carries machine-readable marks on the images.

For the feature, an engineer accepts a 60-line generated function. The pre-merge scan finds 0.9 overlap with a function in a copyleft library. The merge is blocked; the engineer reimplements from the algorithm's description, the new version scores near zero, and the ledger records both the block and the rewrite. Six months later a customer's audit asks whether the product contains copyleft code. The answer is a query, not a scramble.

Failure modes

  • No capture at generation time. Reconstructing who wrote what after a dispute is guesswork. Log at the gateway.
  • Claiming ownership of raw output. Registration or warranties that ignore AI material invite challenges. Disclose and disclaim.
  • Scanning only dependencies. Licence tools that check manifests miss snippets pasted into first-party files. Scan diffs.
  • Thresholds set once and forgotten. A scanner that blocks too much gets bypassed; one that blocks nothing is theatre. Review false positives quarterly.
  • Assuming the indemnity applies. Disabled filters or an unsupported tier can void it. Record settings per team.
  • Labels stripped in the pipeline. Image resizing and CMS uploads often drop embedded metadata. Test the published file, not the generated one.

Trade-offs

Capture everything and you store prompts that may contain personal or confidential data; hash them where policy forbids retention, and accept that hashes prove less. Strict pre-merge scanning catches licence problems but slows developers and produces false positives on boilerplate; tune k and the threshold, and allowlist permissive licences that only need attribution. Investing in human modification strengthens ownership but costs the time the model was meant to save; reserve it for assets where exclusivity matters, such as brand imagery, and accept unprotected output for disposable material.

What to do next

  1. Route every model call through a gateway that logs model, version, prompt hash, output hash and user.
  2. Add a contribution table and record human edits and selections for assets you intend to protect.
  3. Enable public-code matching filters where your assistant offers them, for proprietary repositories.
  4. Add a pre-merge snippet scan with a licence-keyed fingerprint index, and review its false positives each quarter.
  5. Read each provider's current output and indemnity terms, and record tier and filter settings per team.
  6. Disclose and disclaim AI-generated material in registrations and in warranties to customers.
  7. Mark synthetic media in a machine-readable way, and verify the marks survive your publishing pipeline.
  8. Read copyright and AI training for the input side, and the EU AI Act guide for labelling duties.
Key takeaway: Copyright in most jurisdictions follows human creative control, so raw model output is weakly protected or not at all, while human text, selection, arrangement and substantial edits can be. Keep evidence at the source: log every generation, record human edits, scan generated code for licensed snippets before merge, track vendor terms and filter settings, and attach machine-readable labels that survive publishing. This is engineering hygiene, not legal advice.