Generative music systems turn a text prompt into a finished track with vocals in seconds, and they have moved the music industry's oldest disputes, who owns a song and who may sound like whom, into engineering territory. A team that builds or hosts one of these systems now needs to answer questions that used to belong only to lawyers: which recordings were in the training set and under what licence, whether an output reproduces a protected melody or lyric, whether a voice in a generated track belongs to a real singer who did not agree, and whether the uploads reaching streaming services are music or a royalty-fraud operation.

This article treats those questions as security and governance controls. It summarises where the disputes stood as of October 2026, then builds the controls a music-AI pipeline needs: a training provenance ledger, a voice consent registry, an output gate that checks fingerprints, melodies, lyrics and voices, disclosure metadata, and fraud monitoring on distribution. Nothing here is legal advice.

Where things stand, as of October 2026

In June 2024 the three major record companies, coordinated by the RIAA, sued the two best-known music generators, Suno and Udio, in US federal court, alleging that both trained on copyrighted recordings without permission. The cases have since split into settlements and continuing litigation. Universal Music Group settled with Udio in October 2025, and Warner Music Group settled with both Udio and Suno in November 2025; the settlements were reported to include licensing arrangements and new models trained on licensed catalogue. Sony Music had not settled, and reporting through 2026 described Universal's claims against Suno as still unresolved. Check the current docket before relying on any of this.

Europe has produced the first judgments. On 11 November 2025 the Munich Regional Court ruled in GEMA v OpenAI that a language model that memorised and reproduced song lyrics infringed, both in the model and in its outputs. On 31 July 2026 the same court ruled for GEMA against Suno, finding that Suno was not entitled to use the musical works GEMA represents and ordering it to stop the infringing uses, disclose information and pay damages; that ruling was first-instance and Suno said it was considering an appeal. Separately, Tennessee's ELVIS Act, in force since July 2024, added a person's voice to the state's right-of-publicity protection, and the EU AI Act's transparency duties for synthetic audio apply from August 2026, subject to any changes the Commission makes to its timetable.

The engineering implication is consistent across all of it: you will be asked to prove what went in, what came out and who consented, and you can only prove what you recorded at the time.

Five harms, five controls

Five distinct harms, each needing its own control:

HarmWhere it happensPrimary control
Unlicensed training dataData acquisitionProvenance ledger and manifest-gated training
Output reproduces a protected workGenerationFingerprint and melody similarity gate
Lyrics memorised and regurgitatedGeneration (text side)Lyric n-gram filter
Unauthorised voice cloneGeneration and uploadConsent registry and speaker-similarity check
Streaming fraud with synthetic tracksDistributionUpload and stream anomaly detection

A model trained only on licensed data can still emit a near copy of a song it learned, so training and output controls are complements, not substitutes.

Controls along a generative music pipelineLicensed sourcescatalogs, stems, opt-insProvenance ledgerrights, hash, termsTrainingmanifest-gatedGenerationprompt filtersOutput gatefingerprint, melody, lyrics, voiceRelease metadataAI disclosure, watermark, C2PADistributionDSPs via distributorFraud monitoringstream patterns, upload velocityVoice consent registrywho may be cloned, for whatEach box is a place where evidence is produced; the ledger and registry are what you show a court or a platform.
Controls sit at acquisition, training, generation, release and distribution. The ledger and the consent registry are the records you will be asked to produce.

A provenance ledger for audio

A training set for music is a set of audio files plus rights in at least two layers: the sound recording, usually owned by a label, and the underlying composition and lyrics, usually administered by publishers and collecting societies such as GEMA. A licence for one is not a licence for the other, which is exactly the distinction the GEMA cases turned on. The ledger records both, per file, and training jobs read only from manifests built from it.

from dataclasses import dataclass
from hashlib import sha256

@dataclass(frozen=True)
class AudioRight:
    content_sha256: str          # hash of the decoded PCM, not the container
    isrc: str | None             # recording identifier
    iswc: str | None             # composition identifier
    recording_licence: str       # contract id, or "owned", or "none"
    composition_licence: str
    allowed_uses: frozenset      # {"train", "fine_tune", "eval"}
    expires: str | None

def build_manifest(ledger, use="train"):
    ok, rejected = [], []
    for r in ledger:
        if (use in r.allowed_uses and r.recording_licence != "none"
                and r.composition_licence != "none"):
            ok.append(r.content_sha256)
        else:
            rejected.append(r)
    return ok, rejected      # store both, signed, with the training run id

Hashing decoded audio rather than the file means a re-encoded MP3 of a licensed WAV still matches, while a different master of the same song does not. If you use web-scraped audio at all, the training-data copyright guide covers acquisition hygiene and opt-out signals.

Voice consent and similarity

Voice is the sharpest line. Streaming services have written it into policy: Spotify's 2025 impersonation policy allows vocal impersonation only when the impersonated artist has authorised it. A system that offers voice cloning, or accepts uploads with vocals, needs a registry of who consented to what, and a check that links each generated voice to an entry.

def voice_gate(track_audio, requested_voice_id, registry, protected_index,
               threshold=0.75):
    emb = speaker_embedding(isolate_vocals(track_audio))   # e.g. ECAPA-style model
    if requested_voice_id:
        consent = registry.get(requested_voice_id)
        if not consent or not consent.covers(use="release", territory="all"):
            return "block: no consent for requested voice"
    nearest, score = protected_index.nearest(emb)            # cosine similarity
    if score >= threshold and nearest != requested_voice_id:
        return f"review: resembles protected voice {nearest} ({score:.2f})"
    return "pass"

The consent record must be specific: which uses, which territories, which period, revocable how, and paid how. The similarity check catches the case where nobody asked for a clone but the output sounds like a famous singer anyway, through a style prompt or a fine-tune. Speaker embeddings are not identity proof: thresholds trade false accusations against misses, singing voices are harder than speech, and heavy effects defeat them. Route matches to human review rather than auto-publishing a verdict. The fraud side of voice cloning, impersonation for payment or access, is covered in the deepfakes article.

The output gate

An output gate checks each generated track against protected works before it is returned or released. Three signals cover most of the risk, and each catches something the others miss.

  • Audio fingerprinting finds near-identical audio: a model reproducing a recording, or a user uploading someone else's track as generated. Chromaprint, the open-source fingerprinter behind AcoustID, and commercial content-ID services work this way. Fingerprints are robust to encoding but not to re-performance, so they miss a melody sung by a new voice.
  • Melody similarity catches the re-performance. Extract a pitch track, segment it into notes (Spotify's open-source Basic Pitch transcribes audio to MIDI), convert to intervals so the comparison ignores key, index protected works by interval n-grams, then align candidates with dynamic time warping that also weighs rhythm.
  • Lyric overlap catches memorised text, the GEMA v OpenAI fact pattern. Index protected lyrics by word n-grams and flag generated lyrics with long exact runs.
def melody_ngrams(notes, n=6):
    """notes: list of (onset_s, midi_pitch, duration_s), sorted by onset."""
    iv = [b[1] - a[1] for a, b in zip(notes, notes[1:])]
    dur = [round(math.log2(max(b[2], 1e-3) / max(a[2], 1e-3))) for a, b in zip(notes, notes[1:])]
    return {tuple(zip(iv[i:i+n], dur[i:i+n])) for i in range(len(iv) - n + 1)}

def melody_gate(gen_notes, index, min_shared=4, max_score=0.35):
    cands = index.candidates(melody_ngrams(gen_notes))      # works sharing n-grams
    for work_id, shared in cands.most_common(5):
        if shared >= min_shared and dtw_contour_score(gen_notes, index.notes(work_id)) > max_score:
            return f"review: melody close to {work_id}"
    return "pass"
Melody similarity: compare interval contours, not raw pitchesAudiogenerated trackPitch trackf0 per frameNote eventsonset, MIDI pitchIntervals+2 +2 -4 +5 ...n-gram index of protected workscandidate matchesAlignment scoreDTW on contour + rhythmIntervals are transposition-invariant; quantised durations make the score tempo-tolerant.
Pitch to notes to intervals makes melody matching key-independent; an n-gram index narrows candidates before an expensive alignment.

Short melodic fragments are shared by thousands of songs, so tune thresholds on known pairs: true copies, independent songs and ordinary genre similarity. The index for candidate retrieval is the same near-neighbour problem solved by locality-sensitive hashing.

Disclosure: metadata, watermarks and credentials

Disclosure has three layers, and they fail differently. Metadata travels with the release through distribution: DDEX, the music industry's standards body for metadata exchange, has developed a way for labels and distributors to state where AI was used in a track, for example vocals, instrumentation or post-production, and Spotify said in 2025 it would display those disclosures in its app. Metadata is easy to strip and depends on honest submitters. Watermarks are embedded in the audio itself: Meta's open-source AudioSeal and Google's SynthID, used on its Lyria music model, are examples. They survive ordinary encoding but can be weakened by heavy processing, pitch shifting or re-recording, and only the vendor's detector can read them. Content credentials under the C2PA standard attach a signed provenance manifest to the file, which proves origin when present but proves nothing when absent.

Use all three, because they cover each other's gaps. Platforms also run detectors that need no cooperation from the generator. Deezer says it receives nearly 75,000 AI-generated tracks a day, more than 44% of its daily deliveries, and that it tagged over 13.4 million in 2025; it reports its detector's accuracy as 99.8%. Those are the vendor's own figures, and a detector tuned on today's generators can lose accuracy on the next model release, so treat detection as one signal rather than the policy. The broader output-labelling picture is in the AI output copyright guide.

Streaming fraud

Synthetic music has made streaming fraud cheap. The pattern is to upload large numbers of generated tracks under many artist names and stream them with bot accounts, each track below the volume that would attract attention, collecting royalties from a pool shared with real artists. In 2024 US federal prosecutors charged a North Carolina man with running such a scheme, alleging that it used AI-generated songs and bots to collect more than $10 million in royalties. Deezer reported that in 2025 fully AI-generated tracks drew only 1 to 3% of its streams, but about 85% of those streams were fraudulent.

If you operate a distributor or a platform, the signals are behavioural and do not depend on detecting AI at all: upload velocity per account, many artist names sharing payment details, catalogues of near-identical durations, plays concentrated in a small set of accounts that listen around the clock, and listening sessions that play each track just past the royalty threshold. Combine those with AI detection and fingerprint clustering, demonetise rather than delete while investigating, and keep the evidence.

Worked example: licensed fan remixes

Consider a music app adding a feature that lets fans make AI remixes of songs from a label that has licensed its catalogue for that purpose, with an optional guest vocal by one of three artists who signed voice agreements. The controls map directly:

  1. The ledger lists the licensed recordings and compositions with allowed_uses set to remix generation, and the generation service refuses source tracks not in the manifest.
  2. The consent registry holds the three voice agreements, scoped to this app, non-commercial fan use and a two-year term. Requests for any other voice are blocked; the speaker-similarity check sends outputs that resemble an unlisted famous singer to review.
  3. The output gate runs fingerprint matching against the wider catalogue, not just the licensed one, to catch a remix that drifts into a different song; melody similarity runs on any user-supplied melody; lyric overlap runs on any generated words.
  4. Each export carries a watermark, a C2PA manifest and an AI-use disclosure, and the terms forbid uploading remixes to streaming services, which fingerprinting at distributors can enforce.
  5. Every decision is logged with the manifest hash, consent record version and gate scores, so a complaint about a specific remix can be answered with evidence within a day.

Failure modes

Ways these controls fail in practice:

  • One-layer licensing. Licensing recordings but not compositions, or the reverse; the missing layer is the one that sues.
  • Consent without scope. A voice agreement that does not state uses, territories and term invites disputes when the voice appears somewhere unexpected.
  • Gates only on the generator. User uploads, fine-tunes and imported stems bypass a check that runs only on model output.
  • Thresholds never calibrated. A melody gate tuned on nothing either blocks everything or nothing; calibrate on labelled pairs and re-check after model updates.
  • Detector as policy. Blocking on an AI-detector score alone punishes human artists when the detector errs and misses the next generator it was not trained on.
  • No retention. Without logs of what was generated, with which inputs and checks, you cannot answer a takedown or a court order.

Trade-offs

Every control costs something. Output similarity checks add seconds of latency and compute per track and produce false positives that frustrate creators; running them asynchronously before export rather than before preview keeps the product responsive. Strict consent rules exclude playful uses that most artists would tolerate; scoped, cheap licences are better than blanket bans. Licensed-only training data shrinks the corpus and, today, the quality of the model. Watermarking every output may conflict with professional users who want clean stems, which is a contract question, not a technical one. Prompt-injection style attacks on audio pipelines, such as hidden instructions in uploaded audio, are a separate risk covered in the audio injection article.

What to do next

  1. Inventory every audio source in your training data and record recording and composition rights separately for each.
  2. Make training jobs read only from signed manifests built from that ledger, and store the manifest hash with each checkpoint.
  3. Create a voice consent registry with use, territory, term and revocation fields before offering any voice feature.
  4. Add fingerprint, melody and lyric checks to the export path and calibrate their thresholds on labelled pairs.
  5. Attach a watermark, a C2PA manifest and DDEX AI disclosure to every released track.
  6. If you distribute music, add upload-velocity and stream-pattern anomaly detection.
  7. Re-check the status of the label cases, the GEMA appeals and EU AI Act guidance each quarter.
Key takeaway: Generative music is governed by evidence. Record recording and composition rights for every training file, gate training on signed manifests, require scoped consent for any real voice, check outputs for fingerprints, melodies and lyrics, disclose AI use in metadata, watermarks and credentials together, and watch distribution for fraud. The courts are still deciding the rules; your logs decide whether you can show you followed them.