Japan's Act on the Protection of Personal Information (APPI) is the binding law most likely to touch an LLM product that serves people in Japan. The broader policy picture, including the AI Promotion Act and the government guidelines, is covered in Japan's approach to AI. This article goes one level down: where APPI's obligations attach in an LLM architecture, how to decide whether a call to a model API is a provision of personal data to a foreign third party, how prompt logs, retrieval stores and fine-tuning sets fit the Act, what changes with the amendment passed in July 2026, and what to build so you can show compliance rather than assert it.
This is engineering guidance, not legal advice. Article numbers below follow the Act as renumbered in April 2022. Where the regulator's position depends on facts, the article says so; confirm specifics with counsel and the Personal Information Protection Commission (PPC) guidelines and Q&A before you rely on them.
The definitions that decide the obligations
APPI uses layered definitions, and the obligations depend on the layer. Personal information is information about a living individual that can identify them, including through easy combination with other information, or that contains an individual identification code such as a My Number or passport number. Personal data is personal information that forms part of a personal information database, meaning a collection organised so specific individuals can be searched. Special care-required personal information covers categories such as race, creed, social status, medical history, criminal record and having been a victim of crime.
The layering matters for LLM systems because the same text moves between layers. A name typed into a chat box is personal information. Once the prompt is written to a searchable log keyed by user ID, it is personal data, and the provisions on security, third-party provision and disclosure requests apply to that log. A retrieval corpus of customer tickets is a personal information database by design. Two further categories are useful tools rather than burdens: pseudonymously processed information (Article 41), which may be used internally for purposes beyond the original ones after the new purpose is published but may not generally be given to third parties, and anonymously processed information (Article 43), which is processed so it cannot be restored and carries its own publication duties.
Mapping APPI onto an LLM architecture
Walk the diagram left to right. Purpose (Articles 17 and 21): specify the purpose of use as concretely as you can and notify or publish it. "Improve our services" is weak cover for training a model on customer chats; a purpose that names AI model development is stronger. Using data beyond the purpose needs consent (Article 18). Acquisition (Article 20): acquiring special care-required information generally needs prior consent, and a user pasting a medical history into a chat assistant is acquisition. Security (Article 23) and supervision of employees and trustees (Articles 24 and 25) cover the gateway, the vendor and the logs. Retention (Article 22): make efforts to delete personal data once it is no longer needed, which is a direct argument for short prompt-log retention.
Breach (Article 26): leaks of special care-required data, data whose misuse could cause financial damage, leaks caused with an unlawful purpose such as an attack, and leaks affecting more than 1,000 people must be reported to the PPC and notified to the individuals. PPC guidelines expect a preliminary report promptly, generally within three to five days of becoming aware, and a final report within 30 days, or 60 days where an unlawful purpose is involved. A prompt-log bucket made public is a reportable event.
Model APIs: no handling, entrustment, provision, and abroad
The central question for most teams is what happens when a prompt containing personal data is sent to a model API. There are three possible readings. If the provider does not handle the data, for example because the contract forbids access and the service is designed so the provider does not use the content, the PPC has long treated comparable cloud services as not being a provision at all, though your security obligations remain. If the provider processes the data only to perform your service, that is entrustment (Article 27(5)), which does not need consent but requires you to supervise the trustee. If the provider uses the inputs for its own purposes, such as training its models, it is a provision to a third party and generally needs the individual's consent.
The foreign dimension is separate and easy to miss. Article 28 restricts provision to a third party in a foreign country, and the entrustment route does not by itself satisfy it. You need the individual's consent given with information about the country, its data protection system and the recipient's measures; or a recipient that maintains a system equivalent to APPI's standards, backed by contract, with ongoing monitoring; or a destination the PPC recognises as equivalent, which covers the EEA and the UK. Even in the no-handling case, security control measures include understanding the legal environment of the country where data is stored and disclosing it.
The PPC applied this logic to generative AI directly. In June 2023 it cautioned businesses that entering personal information into prompts must stay within the purpose of use, and that if the provider uses prompt data for machine learning, handing it over without consent may breach the Act; it told businesses to confirm the provider does not do so. It also cautioned OpenAI about collecting special care-required information.
Enforcing it: a gateway in front of every model call
The control that makes the vendor analysis true is a gateway between your app and every model call. It enforces which vendors may receive which data classes, strips identifiers you do not need, and writes an audit record that can be shown to a regulator. A simplified version:
import re, hashlib, time
# Japanese identifiers the model never needs. Tune and test these; regexes miss things.
PATTERNS = {
"my_number": re.compile(r"(?<!\d)\d{4}[ -]?\d{4}[ -]?\d{4}(?!\d)"),
"phone": re.compile(r"(?<!\d)0\d{1,4}-\d{1,4}-\d{4}(?!\d)"),
"email": re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+"),
}
VENDORS = { # filled in from the vendor review, not by the developer at call time
"jp-hosted": {"handles": False, "abroad": False, "trains": False},
"us-api-zdr": {"handles": True, "abroad": True, "trains": False,
"art28_basis": "equivalent-system contract 2026-03"},
}
def gate(prompt, vendor, purpose, allowed_purposes, sensitive_detected):
v = VENDORS[vendor]
if purpose not in allowed_purposes:
raise PermissionError("purpose not notified to users")
if v.get("trains"):
raise PermissionError("vendor trains on inputs: third-party provision")
if v["abroad"] and v["handles"] and not v.get("art28_basis"):
raise PermissionError("no Article 28 basis recorded for this vendor")
if sensitive_detected:
raise PermissionError("special care-required data: route to consented flow")
redacted = prompt
for name, rx in PATTERNS.items():
redacted = rx.sub(f"[{name}]", redacted)
audit = {"ts": time.time(), "vendor": vendor, "purpose": purpose,
"sha256": hashlib.sha256(redacted.encode()).hexdigest()}
return redacted, auditNotice what the code does not do: it does not decide the legal basis. The vendor table is the output of a review recorded elsewhere, and the gate only refuses calls that contradict it. Detection of special care-required content needs a classifier, not regexes, and should fail closed. For broader PII handling patterns see PII leakage in LLM systems.
Worked example: an insurer's claims assistant
A Japanese insurer builds a claims assistant. Agents paste claim notes, which often include diagnoses, into a chat tool backed by a model API hosted in the United States; a retrieval store holds past claims; prompts are logged for 180 days; and the data science team wants to fine-tune on the logs. Walk it through.
- Sensitive data. Diagnoses are special care-required information. The insurer already holds them with consent for claims handling, so using them to assess this claim is within purpose; the assistant must not widen that use.
- The API call. The vendor contract forbids training and retention beyond 30 days for abuse monitoring, so the vendor handles the data for the insurer: entrustment. Because the vendor is abroad, the insurer records an Article 28 basis, here an equivalent-system contract and periodic checks, and makes the country and the vendor's measures available to individuals on request.
- The logs. 180 days is hard to justify. The team cuts it to 30, redacts identifiers at write time and restricts access, reducing both breach exposure and Article 22 risk.
- Fine-tuning. Training a model on claim notes is a new purpose. The team pseudonymises the logs, publishes the new purpose, keeps the set internal, and excludes diagnosis text unless its consent covers it.
The July 2026 amendment
The Diet passed amendments to APPI on 10 July 2026, promulgated on 17 July 2026 as Act No. 56 of 2026. Criminal penalty changes start on 17 January 2027; the main reforms take effect on a date set by Cabinet Order, by 17 July 2028. Enforcement rules and guidelines were still to be published at the time of writing, so treat details as provisional. The changes most relevant to LLM teams:
- Statistical processing, including AI development. A consent exception for providing personal data, and acquiring publicly available special care-required information, when used exclusively for statistical analysis, which includes AI development comparable to statistics. Conditions include public disclosure and a written agreement with the recipient.
- Administrative surcharges. The PPC may impose surcharges for specified serious violations, with thresholds including more than 1,000 affected individuals.
- Children. Parental consent and notice for under-16s, and easier requests to stop use or delete.
- Biometrics. A category of specified biometric personal information, such as face data, with transparency duties and no opt-out route to third parties.
- Breach notice and processors. Possible relief from notifying individuals for low-risk breaches, though reporting to the PPC remains, and relief for processors whose contracts specify prescribed terms.
Failure modes
- Treating a no-training setting as the whole analysis. It answers the third-party question but not Article 28, retention or security.
- Logs nobody owns. Prompt and output logs become the largest personal data store in the system, with no retention, access control or disclosure process.
- Purpose drift. Support chats reused for model training under a purpose that never mentioned it.
- Sensitive data by accident. Users volunteer health or criminal history; without detection and routing, you acquire it without consent.
- Unanswerable requests. A disclosure or deletion request arrives and nobody can find that person's data in vector stores and logs. See erasure in LLM systems for techniques.
Trade-offs
| Design choice | APPI benefit | Cost |
|---|---|---|
| Model hosted in Japan | Avoids Article 28 for that hop | Model choice, price, latency |
| Foreign API, no training, short retention | Entrustment plus Article 28 basis | Contract work and monitoring |
| Redact before every call | Less personal data leaves at all | Lower answer quality on names and IDs |
| Consent screen for AI features | Clear basis, including abroad | Friction; consent must be informed |
| Pseudonymised training sets | Internal reuse for new purposes | Cannot share with third parties |
What to do next
- Draw your data flow like the diagram above and mark every store and every hop that leaves Japan.
- For each model vendor, record handling, training, retention, location and the Article 28 basis, with evidence.
- Rewrite the published purpose of use so AI processing you actually do is named.
- Put a gateway in front of model calls that enforces the vendor table, redacts and writes an audit record.
- Set prompt-log retention in days, not "indefinitely", and test that deletion reaches logs, caches and indexes.
- Rehearse a breach: who reports to the PPC within days, and what the 30-day report will contain.
- Track the 2026 amendment's enforcement rules and plan for the under-16 and biometric duties if they apply.