Almost every LLM application depends on someone else's model. The prompts you send carry customer data to a vendor, the answers you receive shape what your product does, and the model behind the API can change, be deprecated or become unavailable on the vendor's schedule rather than yours. Third-party risk management, the discipline procurement and security teams already apply to SaaS, still applies, but the standard questionnaire misses most of what is new.

This page is about managing that vendor relationship over its whole life: finding every LLM vendor you actually use, tiering them by what they touch, asking the due-diligence questions that are specific to models, writing the contract terms that matter, detecting silent model change, accounting for the vendor's own suppliers, and planning the exit before you need it. The runtime controls at the API boundary, such as brokers, scoped credentials and treating responses as untrusted, are covered in third-party API security, which this page assumes.

Why an LLM vendor is a different kind of vendor

A traditional SaaS vendor sells software that changes in visible releases and stores the data you choose to put in it. An LLM vendor differs in four ways that matter for risk:

  • Data flows in prompts. Whatever your application assembles into context, including retrieved documents, tool results and conversation history, leaves your boundary. Users and developers often do not know what that includes.
  • Behaviour is the product, and it moves. A model alias can point to a new version, safety filters can be retuned, and output quality can shift with no change in your code. Your tests passed against a model you may no longer be calling.
  • Data use is a policy, not a property. Whether prompts are retained, reviewed by people for abuse or used for training depends on the account type, product tier and settings, and those terms change.
  • Deep supply chains. The vendor runs on a cloud provider, may route to another model provider, and may call search or tools on your behalf. Each is a party your data reaches.

Frameworks have caught up. The NIST AI Risk Management Framework 1.0 (January 2023) has a GOVERN category specifically for third-party AI risk, ISO/IEC 42001:2023 defines an AI management system that covers suppliers, and the EU AI Act has placed obligations on general-purpose model providers since 2 August 2025. None replaces doing the work below; they give it a vocabulary auditors recognise.

Inventory: find every model you depend on

1. Inventoryevery model in use2. Tierdata x impact3. Due diligenceevidence, not answers4. Contractdata use, notice, exit5. Integratebroker, pinning6. Monitorcanary evals, notices7. Re-assesson trigger or yearly8. Exittested fallbackre-tierThe loop never ends at the contract: model versions, sub-processors and terms change underneath you.
The third-party LLM risk lifecycle. Monitoring feeds back into tiering whenever a model, a data flow or a term changes.

You cannot assess vendors you do not know about, and LLM usage spreads faster than procurement. Build the inventory from evidence, not from a survey:

  • Egress and DNS logs for known model API hostnames, which find direct calls from services and laptops.
  • Dependency manifests for model SDKs and framework integrations, which find code that can call a model even if it is not yet doing so.
  • Expense and card data for AI subscriptions, which find teams buying tools directly.
  • SaaS vendors' own release notes: an existing CRM, ticketing or document tool that adds an AI feature has just added a model provider to your supply chain.

Record each use as one row: the vendor, the specific model or product, the feature that uses it, the data classes it receives, whether it can take actions, the owner, and the contract it runs under. A consumer account used with customer data is a finding in itself.

Tier the use, not the vendor

Not every vendor deserves the same scrutiny. Tier on two axes, the most sensitive data the use receives and the most consequential thing its output can do, and let the higher of the two set the tier:

TierTypical useMinimum assurance
1, criticalRegulated or customer personal data in prompts, or output that triggers payments, account changes or decisions about peopleFull due diligence, independent attestation, negotiated contract, canary evals, tested exit plan, yearly reassessment
2, elevatedInternal confidential data, or output a human reviews before it takes effectQuestionnaire with evidence, enterprise terms with no training on your data, version pinning, two-yearly reassessment
3, lowPublic or synthetic data, output used only as a draftStandard terms reviewed, inventory entry, owner named

Tier the use, not the vendor. The same provider can be tier 3 for a marketing draft tool and tier 1 for the support copilot that reads tickets, and the stronger controls must cover the riskier use.

Due diligence questions that are specific to models

Generic questionnaires ask about encryption, access control and penetration tests, and you should still ask. For a model vendor, add questions whose answers change your risk, and ask for evidence wherever it exists:

AreaQuestionEvidence to request
Training useAre our inputs or outputs used to train or improve any model, and how is that turned off?Contract or data processing terms, and the account setting
RetentionWhat is kept, for how long, where, and who can read it, including abuse-monitoring logs?Written retention schedule; eligibility for reduced or zero retention
LocationWhich regions process and store our data, including failover?Region list; residency commitments in contract
Sub-processorsWho else receives our data, and how are we told about changes?Sub-processor list and notification mechanism
VersioningHow are model versions named, how long is a version available, how much notice before retirement?Published deprecation policy and history
Safety testingHow are models evaluated and red-teamed before release?System cards or evaluation summaries
AssuranceWhich independent attestations cover the service we will use?SOC 2 Type 2, ISO/IEC 27001, ISO/IEC 42001 certificates and scope statements
IncidentsHow fast are customers notified of a security incident affecting their data?Contracted notification window

Read the scope section of every attestation. A SOC 2 report for a vendor's console does not cover its inference API, and a certificate for one region says nothing about another. How to treat a provider's report inside your own audit is in SOC 2 for LLM applications.

Contract terms that matter

The answers above only bind the vendor if they are in the contract. Data protection terms and processor roles are covered in GDPR for LLM applications; on top of those, negotiate or at least confirm these for tier 1 and 2 uses:

ClauseWhat good looks like
No training on customer dataExplicit, covering inputs and outputs, not just a dashboard setting
RetentionStated maximum, deletion on termination, any abuse-review exception named
Sub-processor changesAdvance notice with a right to object or terminate
Model lifecycleMinimum notice before a pinned version is retired or changed
Incident notificationA fixed window and a named contact
Audit rightsAccess to current attestations each year, and to remediation of findings
Output termsWho owns outputs, and any indemnity for third-party claims about them
ExitData return or deletion, and certification that it happened

Monitoring: terms, attestations and silent model change

A signed contract describes the vendor on one day. Three things drift afterwards. Terms and sub-processors change, so subscribe to the vendor's notice channels and route them to the inventory owner. Attestations expire, so record renewal dates. And the model changes, which is the drift most programmes miss.

Pin explicit model versions in production configuration instead of aliases that move. Then run a canary evaluation on a schedule and on every vendor announcement: a fixed set of prompts that represent your use, scored by checks that do not depend on the model being tested. A sudden change in the pass rate, refusal rate or output length tells you behaviour moved before your users do.

import json, statistics

CANARIES = json.load(open("canaries.json"))   # [{"id", "prompt", "must_contain", "max_words"}]
BASELINE = json.load(open("baseline.json"))   # {"pass_rate": 0.96, "median_words": 140}

def run_canaries(call_model, model_version):
    passed, lengths, refusals = 0, [], 0
    for case in CANARIES:
        out = call_model(model_version, case["prompt"])
        words = len(out.split())
        lengths.append(words)
        if out.lower().startswith(("i can't", "i cannot", "i'm unable")):
            refusals += 1
        if all(s.lower() in out.lower() for s in case["must_contain"]) \
                and words <= case["max_words"]:
            passed += 1
    rate = passed / len(CANARIES)
    alerts = []
    if rate < BASELINE["pass_rate"] - 0.05:
        alerts.append(f"pass rate {rate:.2f} below baseline")
    if abs(statistics.median(lengths) - BASELINE["median_words"]) > 0.3 * BASELINE["median_words"]:
        alerts.append("median output length moved more than 30%")
    if refusals > 0.05 * len(CANARIES):
        alerts.append(f"{refusals} refusals")
    return {"model": model_version, "pass_rate": rate, "alerts": alerts}

Keep the canary set small enough to run cheaply and specific enough to matter: twenty to fifty prompts drawn from real traffic with personal data removed. The same harness evaluates a candidate replacement model, which makes it the core of the exit plan as well.

Worked example: a ticket summariser

A company wants to summarise support tickets with a hosted model. Tickets contain names, emails and order details, and summaries are read by agents who act on them. Data is customer personal data and output is reviewed by a person, so the use is tier 1 on the data axis even though it is tier 2 on impact.

Due diligence finds that the vendor's enterprise terms exclude training on customer data, that standard retention includes an abuse-monitoring window, and that the attestation covers the inference API in the needed region. The team asks for reduced retention, which the vendor offers for this account type, and confirms it in the order form. Sub-processors include the vendor's cloud host and a content-safety service; both are added to the record.

Integration pins a dated model version and sends tickets through the company's broker, which removes card numbers and replaces emails with tokens before the prompt leaves. Monitoring runs forty canary tickets nightly. Four months later the vendor announces the pinned version will retire in ninety days. The canary harness scores the successor at the same pass rate but with summaries 40 percent longer, so the team tightens the length instruction, re-runs the canaries and switches with a week to spare. That is the process working: a vendor change that became a scheduled task instead of an incident.

Fourth parties and concentration

Your vendor's suppliers are your fourth parties. A model provider hosted on one cloud inherits that cloud's regional outages; an application vendor that wraps another company's model adds that company's terms and retention to yours; a tool or search integration sends parts of your prompt to yet another service. Ask for the chain, record it, and check it against the data classes the use receives.

Concentration risk is the aggregate view. If every tier 1 use depends on one model provider, one outage or one policy change stops all of them at once. You need not use several providers everywhere, but you should know where you depend on one, and for critical uses, have a second option that has been tested.

The exit plan

An exit plan written during an outage is not a plan. For each tier 1 use, write down the alternative model or degraded mode, keep the prompt and tool contracts portable across providers behind your broker, and run the canary set against the alternative at least quarterly so you know its pass rate today. Decide in advance what triggers a switch: retirement notice, an unacceptable term change, a security incident or a sustained outage.

Some uses can degrade instead of switching: a support copilot can fall back to showing the raw ticket, a classifier to a rules-based default. Write the degraded behaviour down, test it, and include it in LLM incident response playbooks so on-call engineers know it exists.

Failure modes

Programmes fail in predictable ways:

  • Inventory by survey. Teams forget, and AI features inside existing SaaS are never reported. Use egress and spend data.
  • Vendor-level tiering. One low-risk assessment covers a later high-risk use of the same vendor. Tier each use.
  • Answers without evidence. A questionnaire saying "we do not train on your data" is not a contract term.
  • Unpinned models. An alias silently moves and quality or safety behaviour changes with no record of when.
  • Ignored notices. Sub-processor and deprecation emails land in a shared inbox nobody owns.
  • Untested exits. The fallback provider's credentials expired months ago and its prompts were never evaluated.

What to do next

Start with the uses that would hurt most:

  1. Build the inventory from egress logs, dependency manifests and spend data, one row per use.
  2. Tier each use on data sensitivity and output impact; take the higher of the two.
  3. For every tier 1 and 2 use, collect evidence on training use, retention, regions, sub-processors, versioning and attestation scope.
  4. Get the contract clauses into writing: no training, retention, sub-processor notice, model lifecycle notice, incident window, exit.
  5. Pin model versions and run a nightly canary evaluation with alerts on pass rate, length and refusals.
  6. Assign an owner to every vendor notice channel and record attestation renewal dates.
  7. Map fourth parties and mark single points of dependency across tier 1 uses.
  8. Write and test an exit or degraded mode for each tier 1 use, and add data flows to your data governance inventory.
Key takeaway: Treat every LLM vendor as a supply chain whose product, terms and suppliers change after you sign. Inventory uses from evidence, tier each use by the data it receives and what its output can do, ask model-specific questions about training use, retention, regions, sub-processors and versioning, and put the answers in the contract. Pin model versions, run canary evaluations to catch silent change, map fourth parties, and keep a tested exit or degraded mode for every critical use.