Designing a payment-capable LLM application so that the model never sees a card number is half the job. The companion article PCI DSS for LLM systems covers that half: hosted card capture, tokens in and tokens out, an inbound guard and the requirements that bite hardest. This article covers the other half, which is what an assessor actually tests: can you show where every AI component sits in your PCI DSS scope, who is responsible for each control when a model vendor is involved, how you would know if card data reached a trace store or vector index anyway, and what you do when it has.

It covers scope classification, a scoping inventory as code, a sweep for stray card numbers, a 12.10.7 playbook per AI store, and the PCI SSC AI Principles as controls. Requirement references are to PCI DSS v4.0.1. This is engineering guidance; scope is agreed with your assessor or acquirer, and the standard is the authority.

Scope follows data and connectivity

PCI DSS scope is defined by data and connectivity, not by technology. System components that store, process or transmit cardholder data or sensitive authentication data form the cardholder data environment (CDE). Components that connect to the CDE, or that could affect its security, such as identity providers, deployment pipelines and anything holding credentials to CDE systems, are connected-to or security-impacting and are in scope for the requirements relevant to them. Only components with no access to account data and no route into or influence over the CDE are out of scope.

AI systems blur the categories. A component's category depends on what customers actually do: a chat front end joins the CDE the first time someone types a card number into it. And AI pipelines copy text widely, so one unguarded hop can pull many stores into the CDE at once. The diagram shows a typical classification for an agent built on the tokenised design.

CDE: stores, processes or transmits PANConnected-to / security-impactingOut of scope, with evidenceChat and voice edgebefore the guardInbound PAN guarddetects and tokenisesPayment tool servicecalls tokenisation APIAny store found with PANuntil purged or migratedAgent runtimeinvokes payment toolsLLM gatewayholds tool credentialsCI/CD for the agentdeploys code that chargesIdentity provideradmins of the aboveModel providerreceives guarded text onlyTrace and log storesdownstream of the guardVector store, memoryclean discovery sweepsEval and fine-tune datasampled after the guardtokensguarded textA store moves left the moment a sweep finds PAN in it; the evidence that keeps it right is the sweep history.
A typical scope classification for a tokenised LLM payment agent. The right-hand column is only out of scope while sweeps keep proving it.

A scoping inventory as code

Requirement 12.5.2 requires you to document and confirm scope at least once every 12 months and on significant change, covering data flows, the locations of account data, and the system components in or connected to the CDE. Service providers must do it at least every six months under 12.5.2.1, a requirement that became mandatory on 31 March 2025. The cheapest way to meet them for a fast-moving AI estate is a scoping inventory kept as code, so a change in the AI stack produces a diff someone has to review.

from dataclasses import dataclass

@dataclass
class Component:
    name: str
    receives_unguarded_text: bool    # sits before the inbound PAN guard
    handles_pan_by_design: bool      # e.g. payment tool service
    pan_found_last_sweep: bool       # result of the discovery sweep
    connects_to_cde: bool            # network path or API credentials into CDE
    can_affect_cde_security: bool    # deploys, configures or authenticates CDE parts
    days_since_clean_sweep: int | None = None

def classify(c: Component, sweep_max_age_days: int = 31) -> tuple[str, list[str]]:
    evidence = []
    if c.receives_unguarded_text or c.handles_pan_by_design or c.pan_found_last_sweep:
        if c.pan_found_last_sweep:
            evidence.append("open 12.10.7 incident: retrieve, delete or migrate into CDE")
        return "CDE", evidence + ["full PCI DSS requirements apply"]
    if c.connects_to_cde or c.can_affect_cde_security:
        return "connected-to/security-impacting", ["requirements relevant to the connection"]
    if c.days_since_clean_sweep is None or c.days_since_clean_sweep > sweep_max_age_days:
        return "UNPROVEN", ["no recent clean sweep: cannot argue out of scope"]
    return "out of scope", [f"clean sweep {c.days_since_clean_sweep} days ago", "data-flow diagram"]

estate = [
    Component("chat-edge", True, False, False, True, False),
    Component("pan-guard", True, False, False, True, False),
    Component("agent-runtime", False, False, False, True, False),
    Component("trace-backend", False, False, False, False, False, 6),
    Component("vector-memory", False, False, False, False, False, None),
]
for comp in estate:
    print(comp.name, *classify(comp))

The important output is UNPROVEN. Saying that a store downstream of the guard cannot contain card data is a design claim; an assessor will ask how you know. The sweep history is the evidence, and a store without one is a store you cannot defend as out of scope. Treat these AI changes as significant changes that trigger a re-run: adding a tool that touches payment systems, adding or switching a model provider, enabling content capture in traces, adding conversation memory or retrieval over transcripts, and starting a fine-tuning pipeline. Segmentation that keeps AI infrastructure out of scope needs its own penetration testing.

AI vendors as service providers

Model providers, hosted vector databases, observability vendors and evaluation platforms are third-party service providers if they receive account data or could affect its security. Requirement 12.8 asks for a list, written agreements, due diligence, annual monitoring and, under 12.8.5, a record of which requirements each party manages. Keep it as a responsibility matrix per vendor.

Control areaYour responsibilityVendor responsibility (if it receives PAN)Evidence
Data minimisationGuard and tokenise before the callNoneGuard tests, sweep results
Retention of prompts and outputsChoose retention settings, contract termsHonour retention and deletion termsContract clause, configuration screenshot
Encryption in transitTLS configuration on your clientTLS on the endpointConfiguration, scan output
Access to stored contentYour staff and service accountsVendor staff and abuse-review accessVendor attestation, access reviews
Incident notificationYour 12.10 planNotify you within the agreed timeContract clause
Compliance statusMonitor at least annuallyProvide attestation of complianceAttestation on file, review date

The simplest matrix is one where the vendor column says none: if sweeps prove the provider never receives account data, most rows fall away, which shrinks vendor management as well as the CDE.

Finding card data where it should not be

Discovery means scanning the places card data could have leaked to. For AI systems, scan exported trace attributes, application logs, conversation memory, vector store payloads, semantic caches, evaluation datasets and fine-tuning corpora. Use a Luhn check to cut false positives and report only the last four digits, so the sweep report does not itself become a store of card data.

import json, re, sys
from pathlib import Path

CANDIDATE = re.compile(r"(?<!\d)(?:\d[ -]?){13,19}(?!\d)")
SAD_HINT = re.compile(r"\b(cvv2?|cvc2?|cid|security code|card verification)\b", re.I)

def luhn_ok(digits: str) -> bool:
    total, alt = 0, False
    for ch in reversed(digits):
        d = int(ch)
        if alt:
            d = d * 2 - 9 if d > 4 else d * 2
        total, alt = total + d, not alt
    return total % 10 == 0

def scan_text(text: str):
    for m in CANDIDATE.finditer(text):
        digits = re.sub(r"\D", "", m.group())
        if 13 <= len(digits) <= 19 and luhn_ok(digits):
            window = text[max(0, m.start() - 60): m.end() + 60]
            yield {"last4": digits[-4:], "len": len(digits),
                   "possible_sad": bool(SAD_HINT.search(window))}

def sweep_jsonl(store: str, path: Path):
    """Each line: {"id": ..., "text": ...} exported from a trace, vector or eval store."""
    hits = []
    with path.open(encoding="utf-8") as fh:
        for line in fh:
            rec = json.loads(line)
            for hit in scan_text(rec.get("text", "")):
                hits.append({"store": store, "record": rec["id"], **hit})
    return hits

if __name__ == "__main__":
    found = sweep_jsonl(sys.argv[1], Path(sys.argv[2]))
    print(json.dumps(found, indent=2))
    sys.exit(1 if found else 0)   # non-zero exit makes a scheduled job page someone

Run it on a schedule and keep the reports: a dated empty report is the evidence that keeps a store out of scope. Review Luhn-passing order numbers rather than suppressing them by pattern; a real card number looks the same.

The 12.10.7 response for AI data stores

Requirement 12.10.7, mandatory since 31 March 2025, requires incident response procedures that start when stored PAN is found anywhere it is not expected. The procedures must decide what happens to the data (retrieval, secure deletion or migration into the defined CDE), identify whether sensitive authentication data is stored with it, find where it came from and how it got there, and fix the leak or process gap. AI stores make the deletion step harder than it is for a file share.

StoreDispositionWatch out for
Trace backendDelete affected traces if the backend supports it; otherwise restrict access until retention expires and record the decisionCopies in exported datasets and dashboards
Log pipeline and SIEMDelete or mask at every hopArchive buckets and backups
Vector store or memoryDelete the vectors and their payloads, then confirm the deletion is physicalTombstones until compaction, snapshots; embeddings derived from the text are themselves derived data
Semantic cachePurge entries keyed on the affected textReplicas and warm standby caches
Evaluation datasetsDelete rows from every dataset versionNotebooks and local copies made by engineers
Fine-tuned modelRetire the model and retrain on cleaned dataWeights cannot be cleaned; memorised data can be reproduced
Model providerApply the contract's deletion terms; treat as a provider incidentAbuse-monitoring retention you do not control

If the sweep flags possible sensitive authentication data, such as a CVV next to the number, escalate: requirement 3.3.1 forbids keeping it after authorisation at all, so there is no option to migrate it into the CDE; it has to be deleted.

The PCI SSC AI Principles as controls

On 11 September 2025 the PCI SSC published AI Principles: Securing the Use of AI in Payment Environments. It does not change any requirement; its starting point is that AI systems must comply with the PCI SSC standards that apply to them. It then sorts guidance into what AI systems should not be, should be and may be. Paraphrased, the most useful items for engineers, with a control and evidence for each:

Principle (paraphrased)ControlEvidence
Not trusted with secrets such as keys and credentialsAgent never sees API keys; tools hold them server-sideTool code review, secret scanning of prompts
Not given access beyond what the use case needsPer-agent tool allow-list and scoped data accessConfiguration in version control
Given limited, use-case-specific, revocable credentialsSeparate service account per agent and toolIdentity inventory, revocation test
Actions logged, with a named human responsibleAudit log of every payment tool call; owner per agentLog samples, ownership register
Can be disabled easilyFeature flag or gateway kill switch per agentTested switch-off record
Validated before and throughout deploymentEvaluation suite including payment edge cases, run on each changeEvaluation reports
Protected against malicious input and outputPrompt-injection screening, output checks before tools runTest results
Treated as a potential malicious insider in threat models and incident exercisesTabletop where the agent is compromisedExercise record
May use protected payment data such as tokensTokens and last four digits onlyData-flow diagram

Worked example: card numbers in a trace store

A subscription company runs a support agent on the tokenised design. The trace backend is classified out of scope, backed by weekly sweeps. One Monday the sweep exits non-zero: three spans contain a Luhn-valid 16-digit number with the same last four digits, all in the tool response of get_billing_profile, and the possible-SAD flag is false.

The 12.10.7 procedure runs in order. Disposition: the backend supports deletion by trace ID, so the three traces are deleted, and the team checks that no dashboard export or dataset was built from them since the last clean sweep. Sensitive authentication data: none found, confirmed by reviewing the spans before deletion. Source: the billing team's API had added a full card number field for an internal tool two days earlier, and the agent's tool returned the whole response object. Remediation: the tool now maps the API response through an explicit allow-list of fields, a contract test fails the build if any field matching a card pattern appears, and the sweep runs daily for the next month.

For those two days the trace backend held PAN and was part of the CDE. The incident record, deletion evidence and resumed clean sweeps are what let the team argue it is out of scope again.

Failure modes and related reading

FailureConsequenceControl
Store declared out of scope with no sweep historyAssessor cannot accept the claimScheduled sweeps with retained reports
Tool returns whole upstream objectsCard data appears after an unrelated API changeField allow-lists and contract tests
Vector deletes are logical onlyData remains until compaction and in snapshotsConfirm physical deletion; include backups
Fine-tune built from unswept transcriptsCard data in weights that cannot be cleanedSweep training sets before every run
Sweep report lists full numbersThe report becomes a new CDE storeReport last four only
AI change not treated as significantScope document silently wrongChange checklist that names AI triggers

Related reading: PII handling for LLM applications for detection beyond card numbers, audit logging for LLM systems for the tool-call log the AI Principles expect, preparing for an AI audit for assembling evidence, and AP2 and PCI scope reduction for agent payments that use mandates instead of card numbers.

What to do next

  1. List every AI component, including traces, memory, vector stores, caches, evaluation sets, fine-tuning data and model providers, and classify each with the inventory logic above.
  2. Run the discovery sweep on a schedule, keep dated reports, and make a non-empty result page someone.
  3. Write the 12.10.7 procedure per store type, including physical deletion for vector stores and retirement for fine-tuned models.
  4. Build a responsibility matrix for each AI vendor and check you hold a current attestation for any that receive account data.
  5. Add the AI triggers to your significant-change checklist so scope is re-confirmed when tools, providers or data stores change.
  6. Map the PCI SSC AI Principles to controls and collect one piece of evidence for each before your next assessment.
Key takeaway: Under PCI DSS an AI component's scope is set by what data actually reaches it and what it connects to, not by its design intent. Keep a scoping inventory as code, treat AI changes as significant changes, and back every out-of-scope claim for traces, memory, vector stores and evaluation data with dated discovery sweeps. When card data turns up, follow requirement 12.10.7: decide disposition, check for sensitive authentication data, find the source and fix it, remembering that vector deletes may be logical and fine-tuned weights cannot be cleaned.