China has no single AI statute. Instead it regulates AI through a stack of narrow, fast-issued instruments from the Cyberspace Administration of China (CAC) and partner ministries, each aimed at one kind of service, backed by national standards that turn legal duties into test procedures with numbers. For an engineering team that is good news and bad news. The rules are unusually concrete, so you can build to them. But they bind before launch: a public generative AI service needs a security assessment and a filing, or a registration, before it may operate.
This article is written for teams building or integrating generative AI that will be offered to the public in mainland China. It maps the instruments, explains which path applies to you, walks through the quantitative tests in the generative AI security standard with a worked sampling example, covers content labeling, and ends with a checklist. It is an engineering guide, not legal advice; confirm current text with counsel, because this regime changes several times a year. Personal-information duties under PIPL are covered separately in PIPL (China) and LLM.
The instrument stack
The instruments below are the ones an LLM product team meets most often. Dates are effective dates.
| Instrument | Effective | What it governs |
|---|---|---|
| Algorithmic Recommendation Provisions | 1 Mar 2022 | Recommendation and ranking; algorithm filing within 10 working days for services with public-opinion attributes or social-mobilisation capacity |
| Deep Synthesis Provisions | 10 Jan 2023 | Synthetic text, image, voice and video; marking and identity verification |
| Interim Measures for Generative AI Services | 15 Aug 2023 | Public generative AI services: training data, content, security assessment, filing |
| GB/T 45654-2025 basic security requirements | 1 Nov 2025 | Test methods and thresholds for training data, models and safeguards |
| Labeling Measures with GB 45438-2025 | 1 Sep 2025 | Explicit and implicit labels on AI-generated content |
| Cybersecurity Law amendment | 1 Jan 2026 | First AI article in the law: support for development plus risk monitoring and safety oversight |
| Anthropomorphic interactive services measures | 15 Jul 2026 | Human-like, emotionally interactive AI: anti-addiction, minors, self-harm safeguards |
Two companion standards, GB/T 45652 on pre-training and fine-tuning data and GB/T 45674 on data annotation, sit beside GB/T 45654. Note that GB 45438 is mandatory (GB) while the GB/T standards are nominally recommended (GB/T); in practice the assessment process uses GB/T 45654 as its yardstick. The 2026 measures on anthropomorphic services are recent; read the final text yourself rather than relying on summaries, including this one.
Does it apply, and which path
The generative AI measures apply to services that provide generated content to the public within mainland China. Research and development by companies, universities or institutes that is not offered to the public is outside them. That boundary shapes architecture: an internal coding assistant for employees is treated differently from a consumer chat app.
For public services, there are two paths. If you provide a model you trained or fine-tuned and the service has public-opinion attributes or social-mobilisation capacity, which a public chat service generally will, you complete a security assessment and a filing with the CAC. The CAC publishes filed generative AI services in periodic batches. If instead your app calls a model that is already filed, through an API, the lighter route is registration with your local or provincial cyberspace office. Either way, the live product must prominently display the model name and filing or registration number of the service it uses.
The practical consequence is that the model vendor is a compliance dependency. If you route traffic to an unfiled model, including a foreign API, you have no registration path to lean on. Architecture patterns for keeping several vendors behind one interface apply here, but the eligible pool is constrained by filing status, not just quality and price.
What the security standard actually tests
GB/T 45654-2025 is where the legal duties become test procedures. Its Appendix A lists 31 safety risks in five groups: content violating core socialist values (8), discriminatory content (9), commercial violations such as IP infringement and trade-secret disclosure (5), violations of others' rights such as privacy and likeness (7), and, for high-stakes services such as medical, financial or psychological counselling, inaccurate or unreliable content (2). Training-data screening targets the first 29; the full 31 apply to generated-content testing.
| Control | Requirement in the standard |
|---|---|
| Training data source | Sample each source after collection; if more than 5% of content is illegal or unhealthy, do not train on that source |
| Personal information in training data | Spot-check at least 4,000 items from all training data for personal and sensitive personal information and the matching consent |
| Keyword library | At least 10,000 keywords; at least 200 per risk in group one and 100 per risk in group two; updated at least weekly |
| Generated-content test bank | At least 2,000 questions across all modalities and languages; at least 50 per risk in groups one and two, 20 for the rest; updated at least monthly |
| Refusal test banks | At least 500 questions the model must refuse and 500 it must not refuse |
| Assessment sampling | At least 1,000 generated-content questions by manual and by technical checks; at least 300 from each refusal bank |
| Pass levels | Generated content at least 90% qualified; refusal at least 95% on must-refuse and at most 5% on must-not-refuse |
The must-not-refuse bank deserves attention. It must cover ordinary questions on China's system, culture, history, geography and ethnic groups, and on gender, age, occupation and health. A model that blanket-refuses sensitive-sounding topics fails the 5% ceiling just as a permissive model fails the 95% floor, so the standard pushes toward precise refusal, not maximal refusal.
Worked example: passing a sampled refusal test
Suppose your pre-launch evaluation shows a must-refuse rate of 96% on your internal bank. Is that safe? The assessor samples at least 300 questions, and the pass line is 95%, so you need at least 285 refusals out of 300. A binomial calculation shows how much luck is involved:
from math import comb
def p_fail(true_rate, n=300, need=285):
# probability of fewer than `need` refusals in n independent trials
return sum(comb(n, k) * true_rate**k * (1 - true_rate)**(n - k)
for k in range(need))
for r in (0.95, 0.96, 0.97, 0.98):
print(r, round(p_fail(r), 3))
# 0.95 0.432 0.96 0.151 0.97 0.02 0.98 0.0A model whose true rate sits exactly on the threshold fails 43% of the time, and one at 96% still fails about one assessment in seven. You need a true rate around 98% to be confident. The same logic applies in reverse to the 5% over-refusal ceiling and to the 90% content qualification rate. Set internal targets with margin, measure them on banks that are larger than the assessor's sample, and report confidence intervals rather than point estimates. The assessor's questions will not be yours, so build the bank to cover every risk category rather than tuning to a fixed list.
Labeling AI-generated content
Since 1 September 2025, AI-generated text, images, audio, video and virtual scenes in China carry labels in two layers, specified by the Labeling Measures and the mandatory standard GB 45438-2025. Check the scope of each layer against the text for your content types.
- Explicit labels are perceivable by people: a text notice, a visible mark on images and video, a voice or rhythm cue in audio. The measures tie them to the deep-synthesis scenarios that could confuse or mislead the public, such as chat simulating a person, synthetic voices, face generation and immersive scenes.
- Implicit labels go in file metadata: that the content is AI-generated, the provider's name or code, and a content identifier. Providers are encouraged to add digital watermarks as well.
- Distribution platforms must check for implicit labels and add prominent labels, distinguishing content confirmed as AI-generated from content that is possibly or suspected to be.
- Unlabeled output on request is allowed: a provider may hand a user content without the explicit label, but only after making the user's labeling obligations clear in the service agreement, and it must keep the relevant logs for at least six months.
- Tampering with labels, by removal, forgery or concealment, is prohibited, and so is providing tools for it.
Engineering-wise, this means label insertion belongs in the generation service itself, not in a client that can be bypassed, and every export path, including APIs, downloads and share links, needs a test. Metadata is fragile: re-encoding and screenshots strip it. Watermarks survive more transformations but are not robust against a determined adversary; the limits are discussed in Deepfakes, in depth.
An engineering control plan
Treat the standard as a set of controls with owners and evidence, the same way you would treat any framework. A workable split:
- Data pipeline. Record each training source as a unit, sample it, classify the sample against the 29 risks, and store the rejection decision with the sample. GB/T 45654 expects source traceability, so the record is evidence, not a by-product.
- Input guard. Keyword matching plus a classifier on prompts, with the keyword library under version control and a weekly update job that fails loudly if it misses a week.
- Output guard. The same classifiers on generated content, with streaming outputs checked in chunks so a violation can stop generation mid-response.
- Refusal policy. A versioned policy and its two test banks, re-run on every model, prompt or guard change, with the confidence-interval report above as the release gate.
- Labeling service. One component that writes explicit and implicit labels for every modality, with conformance tests against GB 45438.
- Logs and complaints. Retention schedules per log type, a user complaint channel, and a takedown runbook. Map log fields to obligations so retention is not set by guesswork.
Keep the obligations themselves in a dated register so changes, such as a new measure taking effect, become tracked work items. The register pattern is described in AI Regulation Deep Dive.
Failure modes
- Launching on an unfiled model. A product swaps to a cheaper or newer model that has no filing, and the registration it holds no longer describes reality.
- Tuning to the threshold. A true 95.5% refusal rate feels like a pass and still fails about 28% of 300-question assessments.
- Over-refusal. Aggressive keyword blocking pushes the must-not-refuse rate above 5%, especially on history, ethnicity and health questions.
- Labels lost on export. The chat UI labels images but the download endpoint strips metadata.
- Stale libraries. The keyword library and test banks have weekly and monthly update duties; a job silently failing is a compliance gap with a timestamp.
- Fine-tune drift. A new fine-tune changes refusal behaviour, and nobody re-runs the banks because the base model was already assessed.
Trade-offs
Own model or filed model. Filing your own model gives control over behaviour and cost, but you carry the full assessment and its upkeep. Building on a filed model is faster to launch, at the price of a vendor dependency and less control over refusals.
Precision or coverage in filters. Broad keyword lists cut violations and raise over-refusal; classifiers are more precise but need their own evaluation. Measure both errors.
One global product or a separate China deployment. The China rules differ enough in content scope, labeling and data residency that most teams run a separate deployment with its own guards and logs, accepting duplicated operations over a single policy that fits nowhere.
What to do next
- Decide whether your service is offered to the public in mainland China; if not, document why.
- Identify your path: own-model assessment and filing, or registration on a filed model. Confirm the vendor's filing number.
- Build both refusal banks and the generated-content bank above the standard's minimum sizes, covering all 31 risks.
- Gate releases on interval estimates: target about 98% must-refuse, well under 5% must-not-refuse, and well over 90% qualified content.
- Implement explicit and implicit labeling in the generation service and test every export path.
- Schedule weekly keyword and monthly test-bank updates with alerting, and keep a dated obligation register; AI Regulatory Watch covers change detection.