A contract for an AI service looks like any other software agreement until you read the definitions. The thing you are buying changes underneath you when the vendor ships a new model snapshot. The inputs you send are often your most sensitive data. The outputs can carry someone else's intellectual property. Failure is statistical rather than binary, so a service can be up and still be wrong. Each of those properties is settled, or left unsettled, by a handful of clauses. Engineers usually meet them only after something has gone wrong.

This article reads an AI contract the way an engineer reads an interface specification, and ends with an obligations registry that turns signed terms into gateway checks. Vendor lifecycle management (inventory, tiering and exit plans) is covered in Third-Party LLM Risk, and writing your own product terms is covered in AI Terms of Service. Nothing here is legal advice. The aim is to help you ask counsel the right questions and then build what the answers require.

The document stack and precedence

An AI agreement is rarely one document. It is a stack, and the order of precedence decides which promise wins when two documents disagree. A typical stack has a master agreement or customer terms, an order form with prices and commitments, a data processing addendum, an acceptable use policy, a service level agreement, service-specific terms for the AI products, and several policies incorporated by reference through a URL.

The URL-incorporated documents are the engineering risk: a usage or retention page the vendor can edit unilaterally changes your obligations without a signature, unless the master agreement says changes cannot materially reduce what you negotiated.

DocumentWhat it usually settlesEngineering question
Master agreementLiability caps, indemnities, precedenceWhich document wins on data use?
Order formPrice, committed spend, term, named servicesAre the AI services named, or only the platform?
Data processing addendumProcessor role, sub-processors, regions, breach noticeDo prompts and outputs count as personal data here?
Service-specific AI termsTraining use, retention, abuse monitoring, output ownershipDoes this override the DPA, or the other way round?
Acceptable use policyProhibited uses, rate abuse, safety rulesCan a violation by one end user suspend the whole account?
Service level agreementAvailability target, exclusions, creditsWhich failures count as downtime?

Definitions decide what is protected

A clause saying the vendor will not train on Customer Data protects only what that definition contains. Walk every artifact your integration produces through the definitions.

ArtifactOften defined asWatch for
Prompts and uploaded filesCustomer Data or InputFiles sent to a separate storage API under different terms
Model outputsOutput, owned by the customerOwnership assigned only "as between the parties" and only to the extent the vendor has rights
Embeddings and vector indexesSometimes not mentionedDerived data that is neither Input nor Output
Fine-tuned weightsCustomer model or vendor propertyWeights usable only inside the vendor platform, so there is no exit
Thumbs up or down, correctionsFeedbackA broad, perpetual licence that can include the prompt the feedback refers to
Logs and telemetryService Data or Usage DataA category the vendor may use to improve services

The feedback row deserves attention. Many agreements grant the vendor a broad licence to Feedback. If your interface sends a thumbs-down with the conversation attached, that conversation may have just left the no-training protection. The fix can be contractual (define Feedback to exclude Customer Data) or technical (send ratings without content). Ideally do both.

Training, retention and human review

A no-training clause is necessary but not sufficient. Three related provisions decide what actually happens to a prompt after the response is returned.

  • Retention for abuse monitoring. Many providers keep inputs and outputs for a limited period to detect misuse, even when they do not train on them. Some offer reduced or zero retention for approved customers or endpoints. Get the retention period and the conditions into the signed documents rather than relying on a web page.
  • Human review. Ask whether vendor staff can read flagged content, under what controls and in which countries; your processing records need to show it.
  • Region and sub-processors. Where inference runs, where logs are stored, and which hosting companies are involved are often separate answers. The DPA should list sub-processors and give you notice and an objection right before new ones are added.

These protections often apply only to particular endpoints, tiers or account settings, so a call to an excluded beta endpoint is outside the contract while using the same model. The gateway has to enforce that boundary.

IP indemnities are conditional

Since 2023 the large providers have offered copyright indemnities for generated output. Microsoft's Customer Copyright Commitment covers commercial Copilot customers, on condition that the customer has not disabled or evaded the content filters and safety systems built into the product. Google's generative AI indemnity covers training data and generated output for named Cloud services, but excludes customers who intentionally create or use output to infringe. OpenAI's Copyright Shield covers business customers of ChatGPT Enterprise and the API. The details differ by provider and change over time, so read the current text in your own agreement. Who owns the output in the first place is covered in AI and Copyright.

The common pattern is that the indemnity is conditional, and the conditions are technical facts about your deployment. Typical conditions are listed below.

  • Safety filters and citation or attribution features stay enabled with default settings.
  • You did not deliberately prompt for infringing material or ignore the vendor's warnings.
  • You did not modify the output in a way that created the infringement.
  • The claim is about the service as supplied, not about combining it with your own content.
  • You give prompt notice, let the vendor control the defence and cooperate with it.

Each condition is something you may have to prove months later. A filter turned down to cut false positives, or a fine-tuned deployment, can void the indemnity exactly when you need it. Treat conditions as gateway invariants: log filter configuration, model and endpoint with every call, and alert when a supposedly covered deployment falls outside them. Also check whether the indemnity sits inside or outside the general liability cap. An indemnity limited to twelve months of fees on a small contract is worth much less than its headline.

SLAs for an API that can be up and still fail

Standard SLAs promise monthly availability and pay service credits when it is missed. For an LLM API, the hard part is deciding what counts as unavailable. A request can fail with a 5xx, be throttled with a 429 below your contracted limit, time out, be refused by a content filter, or stream half a response and stop. Most SLAs count only server errors and exclude throttling and beta features. Few commit to latency at all.

Before you sign, replay a month of your own logs against the contract's definition and your own definition. The gap between them is the risk you are carrying yourself.

from collections import Counter

def availability(events, counts_as_failure):
    """events: dicts with status, ttft_ms, within_quota, truncated."""
    total = len(events)
    failed = sum(1 for e in events if counts_as_failure(e))
    return 100.0 * (total - failed) / total

contract = lambda e: e["status"] >= 500                 # typical SLA wording
ours = lambda e: (e["status"] >= 500
                  or (e["status"] == 429 and e["within_quota"])  # throttled below our limit
                  or e["ttft_ms"] > 10_000                       # user gave up
                  or e["truncated"])                             # stream died mid-answer

def credit(pct, tiers=((99.9, 0), (99.0, 10), (95.0, 25), (0.0, 50))):
    for floor, credit_pct in tiers:
        if pct >= floor:
            return credit_pct

Take an illustrative month of 2.1 million calls. Suppose 1,050 return a 5xx, 12,000 are throttled with a 429 during a regional capacity shortage while the account is well inside its quota, and 1,400 more stall or are truncated. The contract definition gives 99.95 percent. Ours gives 99.31 percent. Under the contract, that month was compliant and earned no credit. The useful negotiation targets are to count 429s below the contracted limit as errors, to treat a stalled stream as a failure, and to make credits a share of the affected service's fees rather than a token amount.

Model change, deprecation and evaluation rights

A model vendor that ships a new snapshot behind the same alias is usually not in breach. The relevant clauses:

  • Version pinning. Whether dated model versions exist, and whether an alias such as "latest" can move without notice.
  • Deprecation notice. How many days separate the retirement announcement from shutdown, and whether a successor is guaranteed at the same price.
  • Material change. Whether a change that degrades your use case lets you terminate without paying for unused committed spend.
  • Evaluation rights. Whether you may benchmark the service and publish or share the results internally. Some terms restrict benchmarking.

Notice only helps if you use it: pin dated versions, evaluate the successor as soon as it is announced, and record when you saw the notice, because termination rights often run from it.

Back-to-back flow-down to your customers

If you sell an AI feature, customers will send you their own AI addendum. The rule is back-to-back: never promise more than your upstream contract gives you unless you knowingly carry the risk.

Your customer asks forYour upstream vendor givesGap and fix
Zero retention of promptsLimited retention for abuse monitoringGet zero retention upstream or disclose the retention period
99.9% monthly, 429s counted99.9%, 5xx onlyAdd a second provider or lower your promise
Uncapped IP indemnityIndemnity inside a 12-month fee capCap yours to match, or price the risk
EU-only processingRegion choice for inference, global logsConfirm log location in the DPA
90 days' notice of model changeDeprecation notice shorter than thatPromise notice of your own changes only

The usual mistake is sales accepting a customer's AI addendum without anyone checking the upstream column.

An obligations registry the gateway enforces

Contracts live in a document system, and enforcement lives in an API gateway. Connect them with a small, version-controlled obligations registry. Each entry records what was signed, the date it took effect, and the runtime conditions it depends on. The gateway refuses calls that would leave the contract.

Every clause you sign should map to a control you run and evidence you keepContract clauseRuntime controlEvidence retainedNo training on Customer DataApproved tier onlyEndpoint + tier in logIP indemnity conditionsFilters on, never offFilter config per callRegion and sub-processorsRegion pin in gatewayRegion per requestSLA: what counts as an errorClassify every failureStatus, latency, causeModel change noticePinned version + evalsVersion + eval scoresA clause with no control is a promise you cannot keep; a control with no evidence is a claim you cannot prove.
Clause, control and evidence: the three columns every AI contract obligation needs.
# obligations.yaml -- reviewed by legal, owned by platform engineering
vendors:
  model_vendor_a:
    agreement: MSA-2026-014 + AI terms v3 (signed 2026-03-02)
    endpoints_covered: [chat-eu, embed-eu]          # zero-retention applies only here
    regions: [eu-west]
    data_classes_allowed: [public, internal, confidential]
    indemnity:
      requires_default_filters: true
      excluded: [fine_tuned, beta]
    pinned_models: [model-x-2026-05-14]
class ContractViolation(Exception):
    pass

def admit(call, reg):
    v = reg["vendors"][call.vendor]
    if call.endpoint not in v["endpoints_covered"]:
        raise ContractViolation(f"{call.endpoint} not covered by {v['agreement']}")
    if call.region not in v["regions"]:
        raise ContractViolation(f"region {call.region} outside contract")
    if call.data_class not in v["data_classes_allowed"]:
        raise ContractViolation(f"{call.data_class} data not permitted")
    if call.model not in v["pinned_models"]:
        raise ContractViolation(f"unpinned model {call.model}")
    indemnified = (not v["indemnity"]["requires_default_filters"] or call.filters == "default") \
        and call.deployment_kind not in v["indemnity"]["excluded"]
    return {"agreement": v["agreement"], "indemnified": indemnified}   # logged with the call

The returned record is logged beside the request, so a later claim can show which agreement governed the call and whether the indemnity conditions held. On renewal, the diff to this file is the change review.

Worked example: two contracts in one week

A legal-research startup is about to sign an enterprise agreement with a model vendor and, the same week a law-firm customer sends its AI addendum. The definitions review finds that the "report a bad answer" button sends the full conversation as broadly licensed Feedback, so the button now sends only a rating and a ticket ID. The indemnity review finds that citation extraction uses a fine-tuned model the indemnity excludes, so that path is recorded as not indemnified. The SLA replay shows the gap above, so they decline to count 429s as downtime until they add a fallback provider. The law firm wants 90 days' notice of model changes the vendor does not give, so they offer notice of their own changes plus a right to re-run the firm's evaluation set before any switch. Every issue surfaced before signature.

Failure modes

  • Consumer-account drift. Developers prototype on personal or free-tier accounts under different terms, and that code reaches production; an employee AI usage policy with tool tiers is the organisational half of the fix.
  • Policy by URL. A retention or usage page changes, nobody is watching it, and the obligations you thought were fixed have moved. Snapshot incorporated pages on signature and diff them on a schedule.
  • Indemnity voided by tuning. A team lowers a filter threshold or moves to a fine-tuned or beta model, and the coverage the business relies on silently stops applying.
  • Feedback leakage. Rating and correction flows send full conversations under a broad Feedback licence.
  • Unmatched flow-down. Sales accepts a customer clause that no upstream contract supports, and the gap is discovered in an incident.

Trade-offs

Custom terms need leverage; smaller buyers often get standard paper, so invest in the technical controls you fully own. Zero retention may cost more or limit features. Strict pinning delays improvements until your evaluation clears them. A second provider closes SLA and flow-down gaps but doubles the contract surface and evaluation work.

What to do next

  1. List every document in each AI vendor's stack, mark which are incorporated by URL, and snapshot them.
  2. Walk prompts, outputs, embeddings, fine-tuned weights, feedback and logs through the definitions, and record where each lands.
  3. Write down the indemnity conditions, check every production deployment against them, and log the evidence per call.
  4. Replay a month of logs under the contract's availability definition and your own, and take the gap to the next negotiation.
  5. Pin dated model versions and set reminders for deprecation dates and notice-based termination windows.
  6. Build the back-to-back table before accepting any customer AI addendum.
  7. Create the obligations registry and make the gateway refuse calls that fall outside it.
Key takeaway: An AI contract is defined by its definitions, its conditions and its exclusions. Find where your prompts, outputs, feedback and derived data fall in the definitions. Treat every indemnity condition as a runtime invariant, measure the SLA against your own failure definition, and never promise customers more than your vendor promised you. Then put the signed terms in a registry your gateway enforces, so the contract and the system cannot drift apart.