When a team picks an open-weight model, the conversation is usually about benchmarks, context length and how many GPUs it needs. The license gets a glance at the model page, a word like "open" or "commercial use allowed" is noted, and work begins. Months later someone asks whether the fine-tuned checkpoint can ship to a customer, and by then the answer is expensive.

This article treats model licenses as an engineering input, the same way you treat a dependency's license in a software bill of materials. It explains the three families of terms you will meet, what each clause means for the things engineers actually do (fine-tune, merge, quantize, distil, serve, redistribute), how to build a license inventory and a CI gate that blocks a non-compliant model before it ships, and the failure modes that catch careful teams. It is not legal advice: the aim is that you know which questions to ask, can answer most of them from the documents, and can recognise when you need a lawyer. Statements about specific licenses are as of October 2026; vendors revise terms between releases, so always read the LICENSE file shipped with the exact checkpoint you use.

Weights are not source code

Software licenses were written for source code. A model release is a bundle of different things: the weights (large arrays of numbers produced by training), the configuration and tokenizer files, the inference code, sometimes the training code, and documentation such as a model card. Each part can carry different terms. It is common for the inference code to be under Apache 2.0 while the weights are under a custom agreement, so "the repository is Apache" tells you nothing about the file you care about.

Weights also raise a question code does not: whether they are protected by copyright at all is unsettled in most jurisdictions. Vendors therefore write their terms as contracts. You accept them by clicking through a gated download or by using the materials, and they bind you whether or not copyright would. Behave as though the terms apply; your customers' procurement teams will.

Finally, a model license usually points to other documents. An acceptable use policy (AUP) lists forbidden uses and is incorporated by reference, so it is part of the license even though it lives at a different URL and can be updated. Trademark and attribution rules may be separate again. Your inventory needs to capture all of them, with the versions you accepted.

Three families of terms

Almost every license you will meet falls into one of three families. Knowing the family tells you which questions matter.

FamilyExamples (as of Oct 2026)Core obligationsWhat to watch
Permissive OSI licensesApache 2.0 (e.g. the gpt-oss models, many Qwen3 and Mistral checkpoints), MIT (e.g. DeepSeek-R1)Keep the license text and notices; Apache also requires marking changed files and carries a patent grantCheck each checkpoint: one vendor can ship different sizes under different terms
Vendor community licensesLlama community licenses (3.1, 3.3, 4)Copy of the agreement on redistribution, attribution text, naming rule for derivatives, AUP compliance, a user-count thresholdThe threshold and AUP are business restrictions, so the license is not open source by the OSI definition
Use-restricted licensesGemma Terms of Use, OpenRAIL-M and other RAIL variantsUse restrictions that must be passed to everyone downstream as enforceable termsFlow-down to your own customers, broad definitions of derivative, vendor rights to restrict use

A fourth bucket, research-only or non-commercial terms such as CC BY-NC, is easy to violate when a prototype quietly becomes a product.

The clauses that decide what you can do

Whatever the family, the same handful of clauses decide what you can do. Read every license for these and record the answer in your inventory.

  • Scope of grant. Is commercial use allowed? Is the grant worldwide? Some licenses grant rights to the materials but exclude certain modalities or regions in the AUP; record exclusions explicitly instead of assuming the headline grant covers you.
  • User or revenue thresholds. The Llama community licenses require a separate license from Meta if your products had more than 700 million monthly active users at the relevant date. Few companies hit that, but acquirers and large partners ask.
  • Acceptable use. AUPs forbid categories such as weapons development, unlicensed professional practice or deceptive content. They apply to your users' behaviour too, so the obligation is operational: you need terms of service and monitoring that cover the listed categories.
  • Derivative definition. This is the clause that matters most for engineering. Gemma's terms define model derivatives to include models trained by transferring patterns from Gemma's weights or outputs, including by distillation. Llama 2 forbade using outputs to improve other large language models; Llama 3.1 and later instead allow it but require the resulting model's name to begin with "Llama".
  • Redistribution conditions. Typically a copy of the agreement, retained notices, and attribution such as displaying "Built with Llama" in product documentation. RAIL-style licenses require you to include the use restrictions in your own license to recipients.
  • Termination and remote restriction. Apache 2.0 terminates the patent license if you sue over patents in the work. Gemma's terms let Google restrict usage it believes violates the agreement. Termination clauses tell you how fragile a dependency is.
  • Output ownership. Most vendor terms disclaim rights in outputs, but read it: outputs are what your customers receive.

How terms follow an artifact

Where license obligations attach as weights move through your stackUpstream checkpointLICENSE + AUP + NOTICEModel inventoryrepo, revision, license idPolicy gateallowed for this use?Approved basepinned by hashFine-tune / LoRAderivative: terms followQuantize / mergederivative: terms followDistil from outputscheck output clausesHosted API onlyno weights leave youDistribute weightscopy of terms, notices, naming, AUP flow-downServe to usersAUP enforcement, user thresholds, attributionEvery arrow is a point where the gate must re-run: a new artifact carries the most restrictive terms of its inputs.
A model enters through an inventory and a policy gate, then every transformation produces a new artifact that inherits terms. Distribution and serving each add their own obligations.

The diagram is the mental model to keep. Terms attach to an artifact and follow it through every transformation. A LoRA adapter trained on a Llama base, a 4-bit GGUF conversion, a merge of two checkpoints and a small model distilled from a large one's outputs are all new artifacts, and each carries the union of its inputs' obligations. A merge of an Apache 2.0 model and a Gemma model is not Apache; it carries Gemma's use restrictions, because you cannot satisfy Gemma's flow-down requirement while redistributing under terms that omit them.

Do not assume serving is exempt. Permissive licenses ask little of a hosted API, but the Llama 3.3 license attaches its copy, attribution and naming clauses to anything you "distribute or make available", including a product or service. If weights leave your control, your customer also becomes a licensee who must be told the terms.

A model inventory record

The control that prevents most problems is boring: a model inventory with one row per artifact, pinned to an exact revision. The record below is the minimum. Store it next to the weights in your registry and refuse to load anything that has no record.

# model_inventory/support-assistant-v3.yaml
artifact: registry.internal/models/support-assistant-v3
sha256: 9f2c...e41a                      # hash of the weight files you ship
parents:
  - repo: meta-llama/Llama-3.3-70B-Instruct
    revision: 6f6073b4                     # exact commit, not "main"
    license_id: llama3.3
    license_file_sha256: 3c1e...77b0         # the text you actually accepted
    aup_url: recorded in license file
    accepted_by: platform-team, 2026-09-14
transformations: [lora_finetune, merge_adapter, awq_4bit]
derivative_obligations:
  name_prefix: "Llama"
  attribution: "Built with Llama"
  ship_license_copy: true
  aup_flow_down: true
uses:
  serving: approved
  redistribution: approved_with_obligations
  distillation_target: not_reviewed

Three details make this work. Hash the license file you accepted, because the same URL can serve a new version later and you need to show which text bound you. Record parents, not just the immediate base, so a merge or distillation carries every upstream term. And record uses separately: approval to serve a model internally is not approval to ship it to customers.

A license gate in CI

Hugging Face model cards carry a YAML header with a license field, and the hub exposes it as a license:<id> tag. Values include standard identifiers like apache-2.0 and mit, model-specific ones like llama3.3 and gemma, and other with a license_name and license_link. The header is written by whoever uploaded the repository, so it is a hint to check, never proof. The gate below reads it, compares it with the inventory record and with a policy table, and fails the build on any mismatch.

from huggingface_hub import HfApi

POLICY = {  # license id -> uses your legal team has approved
    "apache-2.0": {"serving", "redistribution", "distillation_target"},
    "mit":        {"serving", "redistribution", "distillation_target"},
    "llama3.3":   {"serving", "redistribution"},   # obligations recorded per artifact
    "gemma":      {"serving"},                       # flow-down not yet reviewed
}

def hub_license(repo, revision):
    info = HfApi().model_info(repo, revision=revision)
    tags = [t.split(":", 1)[1] for t in (info.tags or []) if t.startswith("license:")]
    return tags[0] if tags else "unknown"

def check(record, intended_use):
    errors = []
    for parent in record["parents"]:
        declared = hub_license(parent["repo"], parent["revision"])
        if declared != parent["license_id"]:
            errors.append(f"{parent['repo']}: hub says {declared}, inventory says {parent['license_id']}")
        if intended_use not in POLICY.get(parent["license_id"], set()):
            errors.append(f"{parent['repo']}: {intended_use} not approved for {parent['license_id']}")
    return errors

# in CI: fail if check(load_yaml(path), "redistribution") returns anything

Run the gate when a model is first registered, whenever a parent revision changes, and before any release that changes how the artifact is used. Treat unknown and other as failures that require a human to read the linked text and add an explicit policy entry.

Worked example: three bases, one product

A company builds a customer-support assistant. It will serve the model from its own cloud, and two large customers want the fine-tuned model on their own hardware. The team shortlists three bases of similar quality: an Apache 2.0 checkpoint, Llama 3.3 70B Instruct, and a Gemma 3 model.

For the Apache 2.0 base, serving needs nothing beyond keeping notices. Shipping weights to customers needs the license text, any NOTICE file, and a statement that the files were modified. The company may license its fine-tune to customers on its own commercial terms. The main risk is mislabelling, so confirm this exact checkpoint, not just the family, is Apache.

For Llama 3.3, even the hosted service needs AUP terms, "Built with Llama" in documentation and a fine-tune name starting with "Llama", so Acme-Support-v3 becomes something like Llama-Acme-Support-v3. Shipping weights adds a copy of the agreement for each customer, who becomes a licensee bound by the AUP. Procurement will ask about the 700 million user clause; neither customer is near it.

For Gemma, shipping weights means the company's own customer agreement must include Gemma's use restrictions as enforceable terms and give recipients notice of the Gemma terms. That is a contract change, not a file in the tarball, and the sales team must be involved. The team also notes that if it later distils a smaller model from the Gemma fine-tune's outputs, that model is a derivative under Gemma's definition.

The team chooses the Apache base for the on-premises SKU and records the Llama option as approved for serving only.

Outputs, synthetic data and distillation

Teams increasingly generate synthetic training data with a large model and train a small one on it. This is where license terms are least intuitive, because no weights are copied. Three questions decide it. Does the teacher's license define derivatives to include models trained on its outputs? Gemma's does. Does it restrict using outputs to build competing models? Llama 2 did; Llama 3.1 and later permit it with the naming rule. Do the provider's terms for a hosted API restrict training on outputs? Many commercial API terms do, and those are separate from any open-weight license the same vendor publishes.

Record the teacher as a parent of the student in the inventory, with transformation distillation, so the gate applies the teacher's terms. If you mix outputs from several teachers in one dataset, keep per-example provenance; otherwise you cannot later remove one teacher's contribution when a term changes or a review rejects it.

Failure modes

Failure modeHow it happensControl
Mislabelled re-uploadA community quantization lists MIT while the original was a community licenseResolve to the original repo and revision; licenses follow the parent
Family assumptionOne size of a family is Apache, another uses a custom licenseCheck per checkpoint, never per family
Silent term changeThe AUP at a URL is revised after you accepted itHash accepted texts; review diffs on a schedule
Merge contaminationA merge recipe pulls in a use-restricted modelInventory every merge input as a parent
Distillation blind spotA student trained on teacher outputs is treated as cleanRecord teachers as parents; check output clauses
Contract gapWeights ship to customers without flowing down required termsGate redistribution separately from serving
Lost acceptance recordNobody can show which terms were accepted, when, by whomStore acceptance in the inventory with the hash

Trade-offs

Permissive licenses buy freedom and simplicity at the cost of choice: the strongest model for your task may not be permissively licensed. Community licenses are workable for most companies, but naming, attribution and AUP flow-down add product and contract work, and the threshold clause matters in acquisitions. Use-restricted licenses carry the heaviest downstream burden because your customers must accept the restrictions too. A reasonable default: prefer permissive bases for anything you redistribute, and always benchmark one permissive alternative so the choice stays reversible.

What to do next

  1. List every model artifact you serve or ship, with repository, exact revision and weight hash.
  2. Download and hash the license and AUP text for each parent; store them with the record.
  3. Write a policy table mapping license identifiers to approved uses, reviewed with counsel.
  4. Add the hub-tag versus inventory check to CI and fail on unknown, other or any mismatch.
  5. Record merges, adapters and distillation teachers as parents, and gate redistribution separately from serving.
  6. Extend your model cards with license provenance, following the model cards guide, and track weights in the same program as other dependencies, as in the AI supply chain program and SBOMs and software supply chain.
  7. For training data and output ownership questions, continue with copyright and AI training and AI and copyright.
Key takeaway: A model license attaches to an artifact and follows it through every fine-tune, merge, quantization and distillation. Classify each checkpoint into permissive, community or use-restricted terms, read the clauses on derivatives, thresholds, AUPs and redistribution, and keep an inventory that pins revisions and hashes the accepted text. A CI gate that compares hub metadata with that inventory, and treats serving and redistribution as separate approvals, turns licensing from a late legal surprise into a routine build check.