Customers buying an AI product ask questions they never asked about a database: will my data train your model, who else sees my prompts, how long do you keep them, which model provider is behind this, and what happens when it gets something badly wrong. Most companies answer with a trust centre page and a data processing addendum. The answers are usually honest when written. The trouble is that nobody connects them to the code, so six months later a debug logger keeps prompts for ninety days, a new model provider is added without notice, and the promise on the website is quietly false.
This page treats customer trust as an engineering problem: every public promise becomes a registered commitment, bound to a control that enforces it, a test that proves it in CI, and a surface where the customer can check it themselves. It is written for the engineers and security leads who have to make the trust page true. The company in the examples is hypothetical; nothing here quotes any vendor's actual terms.
Architecture at a glance
From first principles: what trust is made of
Trust is a prediction a customer makes about your future behaviour with limited information. They cannot read your code, so they rely on three kinds of signal: what you say (policies and contracts), what someone independent says about you (audits and certifications), and what they can observe themselves (settings, logs, behaviour over time). The first is cheap to produce and easy to falsify by accident. The second is valuable but periodic and sampled. The third is the strongest, because it is continuous and the customer controls it.
AI products strain all three. Model behaviour is probabilistic, so an output cannot be guaranteed the way a query result can. Data flows multiply: prompts, retrieved documents, tool calls, logs, evaluation sets, fine-tuning sets and the model provider's own systems. And the supply chain changes faster than contracts do, because swapping a model is a one-line change. The engineering goal is therefore narrow and concrete: make each statement you publish about data and model behaviour mechanically true, keep it true as the code changes, and let customers verify as much of it as possible.
What customers actually ask, as commitments
Security questionnaires for AI products converge on a small set of themes. Write each answer as a commitment you can enforce, not as prose:
| Customer question | Commitment shape | Enforcing control |
|---|---|---|
| Do you train on my data? | Customer content is excluded from training unless the tenant opts in. | Per-record data-use flag checked by every dataset builder. |
| How long do you keep prompts and outputs? | Retained N days for abuse review, then deleted. | TTL on every store that holds content, including logs. |
| Where is my data processed? | Processing stays in the tenant's chosen region. | Region-pinned routing; model endpoints listed per region. |
| Which model providers see it? | Only listed subprocessors, with notice before changes. | Provider allowlist in the gateway, tied to the published list. |
| Does a human read my conversations? | Only for flagged abuse cases, by named roles, logged. | Access gated by case id; every view written to an audit log. |
| Will you tell us when the model changes? | Notice before a model change that affects outputs. | Model version on every response; change gate with notice period. |
| What happens in an incident? | Notification within a contracted time. | Incident runbook with a customer-notice step and clock. |
Two answers deserve care. "We do not train on your data" must cover more than the base model: evaluation sets, prompt libraries, classifier training and fine-tunes built from support tickets all count in a customer's eyes. And region commitments fail most often through supporting services, such as a logging pipeline or an embedding API, rather than through the main model call.
The commitment registry
Keep commitments as data in the repository, next to the code that enforces them. Each entry names the published wording, an owner, the control, the test and the customer-visible proof. Code review then catches the change that breaks a promise, because the test that guards it fails.
# commitments.yaml (hypothetical company "Northwind AI")
- id: C-TRAIN-01
published: "Customer content is not used to train or fine-tune models unless the workspace admin opts in."
owner: ml-platform
control: data_use flag on every content record; dataset builders call allow_training()
test: tests/commitments/test_training_exclusion.py
customer_proof: admin console shows the opt-in state and the date it last changed
- id: C-RET-01
published: "Prompts and outputs are deleted 30 days after creation."
owner: platform-infra
control: ttl_days=30 on conversation store, request log index, trace store, abuse queue
test: nightly probe writes a canary record and asserts it is gone at day 31
customer_proof: deletion receipts available through the audit log export
- id: C-SUB-01
published: "Model providers are listed on the subprocessor page; 30 days notice before additions."
owner: security
control: gateway provider allowlist generated from subprocessors.yaml
test: CI fails if the gateway config names a provider absent from subprocessors.yaml
customer_proof: subprocessor page with change history and email subscription
Enforcing data-use promises in code
The training promise is the one most often broken by accident, because data reaches training through side doors. The fix is a flag that travels with the content and a single function every dataset builder must call. Here is the shape in Python:
from dataclasses import dataclass
from enum import Enum
class DataUse(Enum):
SERVICE_ONLY = "service_only" # default for every tenant
TRAINING_OK = "training_ok" # set only by an explicit admin opt-in
@dataclass(frozen=True)
class ContentRecord:
tenant_id: str
record_id: str
text: str
data_use: DataUse
class TrainingExclusionError(Exception):
pass
def allow_training(records, opt_in_registry, audit_log):
# Fail closed: a missing flag or a tenant without a current opt-in excludes the record.
kept, dropped = [], 0
for r in records:
tenant_ok = opt_in_registry.is_opted_in(r.tenant_id)
if r.data_use is DataUse.TRAINING_OK and tenant_ok:
kept.append(r)
else:
dropped += 1
audit_log.write("dataset_filter", kept=len(kept), dropped=dropped)
return kept
def build_dataset(source, opt_in_registry, audit_log):
records = source.load()
if not all(isinstance(r, ContentRecord) for r in records):
raise TrainingExclusionError("untyped records cannot prove their data-use flag")
return allow_training(records, opt_in_registry, audit_log)The important properties are that the default is the restrictive value, that the check consults the current opt-in state rather than a value copied at ingestion (so a withdrawn opt-in takes effect at the next build), and that untyped data cannot enter at all. The matching contract test seeds two tenants, one opted in and one not, runs every registered dataset builder against both and asserts that no record from the second appears. Register new builders in the same list the test iterates, so an unregistered builder is a review question.
The other controls follow the same pattern. Retention is a TTL on every store, and the nightly probe checks the stores you forgot, such as trace backends and exported analytics tables. Region commitments are enforced in the gateway: a request tagged with a tenant region can only be routed to endpoints listed for that region. Human access to content goes through a case-scoped tool that writes who looked, at what and why.
Making commitments verifiable
Customers trust what they can check. Build these surfaces into the product rather than the sales deck:
- Settings that show state, not just toggles. The admin console shows the training opt-in, who changed it and when, the processing region, and the retention period that applies.
- A customer-exportable audit log. Admin changes, human access to content (with case reference), deletion receipts and data exports. If staff read a conversation, the customer can see that it happened.
- Model provenance on responses. Return the model identifier and version with each API response, and publish a change log, so a customer can correlate a change in behaviour with a model change.
- A subprocessor list with history. Generated from the same file the gateway reads, with a subscription for changes.
- Independent attestation. A SOC 2 Type 2 report covers whether controls operated over a period; ISO/IEC 42001 certifies an AI management system. Neither proves a specific promise unless the promise is in scope, so put the registry's controls into the audit scope explicitly.
- Disclosure to end users. If people talk to your system, tell them it is AI. In the EU, Article 50 of the AI Act sets transparency duties for chatbots and generated content, applying from 2 August 2026, with a later amendment moving part of it for systems already on the market; confirm the current dates with counsel.
Worked example: a retention promise drifts
Northwind AI sells a support copilot. Its trust page promises 30-day retention (C-RET-01) and no training on customer content without opt-in (C-TRAIN-01). In March an engineer adds verbose request logging to debug latency, writing full prompts to a new log index with the platform default retention of 90 days. Without a registry, nobody notices: the conversation store still deletes at 30 days, and that is what the last audit sampled.
With the registry, three things happen. The pull request adds a new store, and the reviewer's checklist asks which commitments touch it, since the store holds content. If that is missed, the nightly retention probe catches it: it writes a canary conversation tagged canary-ret-0301, and on day 31 searches every store in the content inventory for the tag. The new index is in the inventory because the logging library registers its sinks, and the probe finds the canary still present. The commitment's owner gets a page, the index TTL is set to 30 days, and the 31-to-90-day backlog is purged.
Then comes the part that builds trust rather than just preserving it. The incident runbook asks whether a customer commitment was breached. It was, for about five weeks, though no data left the company. Northwind's DPA requires notice of incidents affecting customer data, so customers get a short factual note: what was retained, for how long, who could access it, what was deleted and when, and the control added. Customers routinely forgive a caught, disclosed and fixed drift. They do not forgive one they discover themselves.
Change is the main threat
Most trust failures in AI products come from change, not from attackers. Put a gate in front of these changes, and have it ask which commitments the change touches:
- A new model or provider: subprocessor list, region list, notice period, and an evaluation run showing behaviour on the customer-relevant test set.
- A new data store or sink: content inventory, retention TTL, region, access controls.
- A new dataset builder or evaluation pipeline: registration with the training-exclusion test.
- A new human-review workflow: role definition, case-scoped access, audit logging.
- A change to the published wording: registry entry updated in the same pull request, with legal review.
The gate is a CI check plus a named reviewer, not a committee. Its output is a diff of the registry, which is also the change log you publish.
Failure modes
- Promise wider than the control. "We never store prompts" while traces capture them. Write promises from the control outwards, never from marketing inwards.
- Side-door training. Support tickets, thumbs-down feedback or red-team transcripts flow into fine-tuning or classifier sets without the data-use check.
- Region leak through helpers. Main inference is regional, but embeddings, moderation or logging call a global endpoint.
- Silent model swap. A provider alias moves to a new model version and output behaviour changes without notice. Pin versions and record them on every response.
- Attestation scope gap. The SOC 2 report covers the platform but not the AI feature, and customers read it as covering both.
- Opt-out that does not reach backups or caches. Withdrawal stops new use but not data already copied. State plainly what withdrawal does and when.
- Undisclosed incidents. A breach of a commitment is handled as an internal bug. Make "was a commitment breached?" a mandatory question in every incident review.
Trade-offs
Narrow, verifiable promises win more durable trust than broad ones, but they read as weaker on a comparison grid, so sales will push for broader wording; the registry gives you a principled answer, because anything published must have a control and a test. Customer-visible audit logs expose staff activity and make internal mistakes visible, which is uncomfortable and is exactly the point. Strict region pinning limits which models you can offer in each region and can delay launches. Attestations are expensive and periodic; contract tests are cheap and continuous, and you need both because customers weigh independent assurance and engineers need fast feedback. And default-off training opt-in costs you data you might want, which is the trade most enterprise buyers now expect you to make.
What to do next
- Collect every published statement about data and models: trust page, DPA, docs, sales answers. Turn each into a registry entry with an owner.
- Mark entries that have no enforcing control or no test. Either build them or narrow the wording.
- Build a content inventory of every store and sink that can hold customer content, and put a retention TTL and a region on each.
- Add a fail-closed data-use flag and route every dataset builder through one check; add the two-tenant contract test.
- Add a nightly canary probe for retention and deletion across the whole inventory.
- Generate the gateway's provider allowlist from the published subprocessor list.
- Ship customer-visible proof: console settings with history, an audit log export and model version on responses.
- Add "which commitments does this touch?" to the change gate and "was a commitment breached?" to incident review.
- Keep learning: SOC 2 for LLM systems, audit preparation, data governance and deletion, tenant isolation and incident response.