Designing a payment-capable LLM application so that the model never sees a card number is half the job. The companion article PCI DSS for LLM systems covers that half: hosted card capture, tokens in and tokens out, an inbound guard and the requirements that bite hardest. This article covers the other half, which is what an assessor actually tests: can you show where every AI component sits in your PCI DSS scope, who is responsible for each control when a model vendor is involved, how you would know if card data reached a trace store or vector index anyway, and what you do when it has.
It covers scope classification, a scoping inventory as code, a sweep for stray card numbers, a 12.10.7 playbook per AI store, and the PCI SSC AI Principles as controls. Requirement references are to PCI DSS v4.0.1. This is engineering guidance; scope is agreed with your assessor or acquirer, and the standard is the authority.
Scope follows data and connectivity
PCI DSS scope is defined by data and connectivity, not by technology. System components that store, process or transmit cardholder data or sensitive authentication data form the cardholder data environment (CDE). Components that connect to the CDE, or that could affect its security, such as identity providers, deployment pipelines and anything holding credentials to CDE systems, are connected-to or security-impacting and are in scope for the requirements relevant to them. Only components with no access to account data and no route into or influence over the CDE are out of scope.
AI systems blur the categories. A component's category depends on what customers actually do: a chat front end joins the CDE the first time someone types a card number into it. And AI pipelines copy text widely, so one unguarded hop can pull many stores into the CDE at once. The diagram shows a typical classification for an agent built on the tokenised design.
A scoping inventory as code
Requirement 12.5.2 requires you to document and confirm scope at least once every 12 months and on significant change, covering data flows, the locations of account data, and the system components in or connected to the CDE. Service providers must do it at least every six months under 12.5.2.1, a requirement that became mandatory on 31 March 2025. The cheapest way to meet them for a fast-moving AI estate is a scoping inventory kept as code, so a change in the AI stack produces a diff someone has to review.
from dataclasses import dataclass
@dataclass
class Component:
name: str
receives_unguarded_text: bool # sits before the inbound PAN guard
handles_pan_by_design: bool # e.g. payment tool service
pan_found_last_sweep: bool # result of the discovery sweep
connects_to_cde: bool # network path or API credentials into CDE
can_affect_cde_security: bool # deploys, configures or authenticates CDE parts
days_since_clean_sweep: int | None = None
def classify(c: Component, sweep_max_age_days: int = 31) -> tuple[str, list[str]]:
evidence = []
if c.receives_unguarded_text or c.handles_pan_by_design or c.pan_found_last_sweep:
if c.pan_found_last_sweep:
evidence.append("open 12.10.7 incident: retrieve, delete or migrate into CDE")
return "CDE", evidence + ["full PCI DSS requirements apply"]
if c.connects_to_cde or c.can_affect_cde_security:
return "connected-to/security-impacting", ["requirements relevant to the connection"]
if c.days_since_clean_sweep is None or c.days_since_clean_sweep > sweep_max_age_days:
return "UNPROVEN", ["no recent clean sweep: cannot argue out of scope"]
return "out of scope", [f"clean sweep {c.days_since_clean_sweep} days ago", "data-flow diagram"]
estate = [
Component("chat-edge", True, False, False, True, False),
Component("pan-guard", True, False, False, True, False),
Component("agent-runtime", False, False, False, True, False),
Component("trace-backend", False, False, False, False, False, 6),
Component("vector-memory", False, False, False, False, False, None),
]
for comp in estate:
print(comp.name, *classify(comp))The important output is UNPROVEN. Saying that a store downstream of the guard cannot contain card data is a design claim; an assessor will ask how you know. The sweep history is the evidence, and a store without one is a store you cannot defend as out of scope. Treat these AI changes as significant changes that trigger a re-run: adding a tool that touches payment systems, adding or switching a model provider, enabling content capture in traces, adding conversation memory or retrieval over transcripts, and starting a fine-tuning pipeline. Segmentation that keeps AI infrastructure out of scope needs its own penetration testing.
AI vendors as service providers
Model providers, hosted vector databases, observability vendors and evaluation platforms are third-party service providers if they receive account data or could affect its security. Requirement 12.8 asks for a list, written agreements, due diligence, annual monitoring and, under 12.8.5, a record of which requirements each party manages. Keep it as a responsibility matrix per vendor.
| Control area | Your responsibility | Vendor responsibility (if it receives PAN) | Evidence |
|---|---|---|---|
| Data minimisation | Guard and tokenise before the call | None | Guard tests, sweep results |
| Retention of prompts and outputs | Choose retention settings, contract terms | Honour retention and deletion terms | Contract clause, configuration screenshot |
| Encryption in transit | TLS configuration on your client | TLS on the endpoint | Configuration, scan output |
| Access to stored content | Your staff and service accounts | Vendor staff and abuse-review access | Vendor attestation, access reviews |
| Incident notification | Your 12.10 plan | Notify you within the agreed time | Contract clause |
| Compliance status | Monitor at least annually | Provide attestation of compliance | Attestation on file, review date |
The simplest matrix is one where the vendor column says none: if sweeps prove the provider never receives account data, most rows fall away, which shrinks vendor management as well as the CDE.
Finding card data where it should not be
Discovery means scanning the places card data could have leaked to. For AI systems, scan exported trace attributes, application logs, conversation memory, vector store payloads, semantic caches, evaluation datasets and fine-tuning corpora. Use a Luhn check to cut false positives and report only the last four digits, so the sweep report does not itself become a store of card data.
import json, re, sys
from pathlib import Path
CANDIDATE = re.compile(r"(?<!\d)(?:\d[ -]?){13,19}(?!\d)")
SAD_HINT = re.compile(r"\b(cvv2?|cvc2?|cid|security code|card verification)\b", re.I)
def luhn_ok(digits: str) -> bool:
total, alt = 0, False
for ch in reversed(digits):
d = int(ch)
if alt:
d = d * 2 - 9 if d > 4 else d * 2
total, alt = total + d, not alt
return total % 10 == 0
def scan_text(text: str):
for m in CANDIDATE.finditer(text):
digits = re.sub(r"\D", "", m.group())
if 13 <= len(digits) <= 19 and luhn_ok(digits):
window = text[max(0, m.start() - 60): m.end() + 60]
yield {"last4": digits[-4:], "len": len(digits),
"possible_sad": bool(SAD_HINT.search(window))}
def sweep_jsonl(store: str, path: Path):
"""Each line: {"id": ..., "text": ...} exported from a trace, vector or eval store."""
hits = []
with path.open(encoding="utf-8") as fh:
for line in fh:
rec = json.loads(line)
for hit in scan_text(rec.get("text", "")):
hits.append({"store": store, "record": rec["id"], **hit})
return hits
if __name__ == "__main__":
found = sweep_jsonl(sys.argv[1], Path(sys.argv[2]))
print(json.dumps(found, indent=2))
sys.exit(1 if found else 0) # non-zero exit makes a scheduled job page someoneRun it on a schedule and keep the reports: a dated empty report is the evidence that keeps a store out of scope. Review Luhn-passing order numbers rather than suppressing them by pattern; a real card number looks the same.
The 12.10.7 response for AI data stores
Requirement 12.10.7, mandatory since 31 March 2025, requires incident response procedures that start when stored PAN is found anywhere it is not expected. The procedures must decide what happens to the data (retrieval, secure deletion or migration into the defined CDE), identify whether sensitive authentication data is stored with it, find where it came from and how it got there, and fix the leak or process gap. AI stores make the deletion step harder than it is for a file share.
| Store | Disposition | Watch out for |
|---|---|---|
| Trace backend | Delete affected traces if the backend supports it; otherwise restrict access until retention expires and record the decision | Copies in exported datasets and dashboards |
| Log pipeline and SIEM | Delete or mask at every hop | Archive buckets and backups |
| Vector store or memory | Delete the vectors and their payloads, then confirm the deletion is physical | Tombstones until compaction, snapshots; embeddings derived from the text are themselves derived data |
| Semantic cache | Purge entries keyed on the affected text | Replicas and warm standby caches |
| Evaluation datasets | Delete rows from every dataset version | Notebooks and local copies made by engineers |
| Fine-tuned model | Retire the model and retrain on cleaned data | Weights cannot be cleaned; memorised data can be reproduced |
| Model provider | Apply the contract's deletion terms; treat as a provider incident | Abuse-monitoring retention you do not control |
If the sweep flags possible sensitive authentication data, such as a CVV next to the number, escalate: requirement 3.3.1 forbids keeping it after authorisation at all, so there is no option to migrate it into the CDE; it has to be deleted.
The PCI SSC AI Principles as controls
On 11 September 2025 the PCI SSC published AI Principles: Securing the Use of AI in Payment Environments. It does not change any requirement; its starting point is that AI systems must comply with the PCI SSC standards that apply to them. It then sorts guidance into what AI systems should not be, should be and may be. Paraphrased, the most useful items for engineers, with a control and evidence for each:
| Principle (paraphrased) | Control | Evidence |
|---|---|---|
| Not trusted with secrets such as keys and credentials | Agent never sees API keys; tools hold them server-side | Tool code review, secret scanning of prompts |
| Not given access beyond what the use case needs | Per-agent tool allow-list and scoped data access | Configuration in version control |
| Given limited, use-case-specific, revocable credentials | Separate service account per agent and tool | Identity inventory, revocation test |
| Actions logged, with a named human responsible | Audit log of every payment tool call; owner per agent | Log samples, ownership register |
| Can be disabled easily | Feature flag or gateway kill switch per agent | Tested switch-off record |
| Validated before and throughout deployment | Evaluation suite including payment edge cases, run on each change | Evaluation reports |
| Protected against malicious input and output | Prompt-injection screening, output checks before tools run | Test results |
| Treated as a potential malicious insider in threat models and incident exercises | Tabletop where the agent is compromised | Exercise record |
| May use protected payment data such as tokens | Tokens and last four digits only | Data-flow diagram |
Worked example: card numbers in a trace store
A subscription company runs a support agent on the tokenised design. The trace backend is classified out of scope, backed by weekly sweeps. One Monday the sweep exits non-zero: three spans contain a Luhn-valid 16-digit number with the same last four digits, all in the tool response of get_billing_profile, and the possible-SAD flag is false.
The 12.10.7 procedure runs in order. Disposition: the backend supports deletion by trace ID, so the three traces are deleted, and the team checks that no dashboard export or dataset was built from them since the last clean sweep. Sensitive authentication data: none found, confirmed by reviewing the spans before deletion. Source: the billing team's API had added a full card number field for an internal tool two days earlier, and the agent's tool returned the whole response object. Remediation: the tool now maps the API response through an explicit allow-list of fields, a contract test fails the build if any field matching a card pattern appears, and the sweep runs daily for the next month.
For those two days the trace backend held PAN and was part of the CDE. The incident record, deletion evidence and resumed clean sweeps are what let the team argue it is out of scope again.
Failure modes and related reading
| Failure | Consequence | Control |
|---|---|---|
| Store declared out of scope with no sweep history | Assessor cannot accept the claim | Scheduled sweeps with retained reports |
| Tool returns whole upstream objects | Card data appears after an unrelated API change | Field allow-lists and contract tests |
| Vector deletes are logical only | Data remains until compaction and in snapshots | Confirm physical deletion; include backups |
| Fine-tune built from unswept transcripts | Card data in weights that cannot be cleaned | Sweep training sets before every run |
| Sweep report lists full numbers | The report becomes a new CDE store | Report last four only |
| AI change not treated as significant | Scope document silently wrong | Change checklist that names AI triggers |
Related reading: PII handling for LLM applications for detection beyond card numbers, audit logging for LLM systems for the tool-call log the AI Principles expect, preparing for an AI audit for assembling evidence, and AP2 and PCI scope reduction for agent payments that use mandates instead of card numbers.
What to do next
- List every AI component, including traces, memory, vector stores, caches, evaluation sets, fine-tuning data and model providers, and classify each with the inventory logic above.
- Run the discovery sweep on a schedule, keep dated reports, and make a non-empty result page someone.
- Write the 12.10.7 procedure per store type, including physical deletion for vector stores and retirement for fine-tuned models.
- Build a responsibility matrix for each AI vendor and check you hold a current attestation for any that receive account data.
- Add the AI triggers to your significant-change checklist so scope is re-confirmed when tools, providers or data stores change.
- Map the PCI SSC AI Principles to controls and collect one piece of evidence for each before your next assessment.