Sensitive Data Protection is the current name for what Google Cloud launched as Cloud DLP. The rename widened the product from an inspection API into a family that also profiles data estates continuously, but the API endpoint is still dlp.googleapis.com and the client libraries are still dlp_v2, so code and documentation use both names. This article explains how detection actually works, the three ways to run it and when each fits, how de-identification differs from redaction, and how to build a reversible tokenisation flow without leaking the data you were trying to protect.
Pricing, quotas and request size limits change often and differ between the content methods and storage scans, so this article gives none; check the current pricing and quotas pages before sizing a deployment.
How detection works
Everything starts with an infoType, a named detector such as EMAIL_ADDRESS, CREDIT_CARD_NUMBER or US_SOCIAL_SECURITY_NUMBER. Built-in infoTypes combine patterns, checksums (a card number must pass the Luhn check), and context. Each finding comes back with a likelihood from VERY_UNLIKELY to VERY_LIKELY; the min_likelihood setting drops anything below it. Lowering it finds more and produces more false positives; there is no setting that does both.
You extend detection in three ways. Custom infoTypes can be a regular expression (an internal employee ID), a word list (project code names), or a stored dictionary built from a large term list in Cloud Storage or BigQuery. Hotword rules adjust likelihood when a nearby term appears, so a nine-digit number next to "SSN" is raised and one next to "order" is not. Exclusion rules remove findings, for example an email detector that should ignore your own support address. Rules belong in an inspect template, a stored, versioned configuration that every job and request references by name, so a tuning change lands everywhere at once.
Three ways to run it
Content methods (inspect_content, deidentify_content, reidentify_content) are synchronous calls on data you send in the request: a string, a table of rows, or an image. Use them inline in pipelines, for example a Dataflow job or a Cloud Run service that tokenises records before they reach the warehouse.
Inspection jobs read data where it sits in Cloud Storage, BigQuery or Datastore and write results through actions: save findings to a BigQuery table, publish to Pub/Sub, or publish a summary to Security Command Center. A job trigger reruns the job on a schedule. Jobs support sampling (rows_limit and sample_method for BigQuery, bytes_limit_per_file and files_limit_percent for Cloud Storage), which is how you keep a scan of a petabyte bucket affordable. Hybrid jobs let you push content from outside Google Cloud into a job and get the same findings output.
Discovery answers "where is sensitive data at all?" across an organisation, folder or project. A scan configuration names the scope, resource type, inspect templates and reprofiling schedule (daily, weekly or monthly, or on schema or template change). Supported sources include BigQuery, Cloud SQL, Cloud Storage and Vertex AI, plus Amazon S3 and Azure Blob Storage. The output is a data profile per project, table, column or file store, carrying predicted infoTypes and calculated data-risk and sensitivity levels, which can be published to BigQuery, Security Command Center, Knowledge Catalog and Pub/Sub.
Inspecting content from code
A minimal inspection with the Python client. Note the include_quote flag; it is discussed under failure modes.
from google.cloud import dlp_v2
dlp = dlp_v2.DlpServiceClient()
parent = "projects/my-project/locations/global"
inspect_config = {
"info_types": [{"name": "EMAIL_ADDRESS"}, {"name": "PHONE_NUMBER"}],
"custom_info_types": [{
"info_type": {"name": "EMPLOYEE_ID"},
"regex": {"pattern": r"EMP-\d{6}"},
"likelihood": "LIKELY",
}],
"min_likelihood": "POSSIBLE",
"include_quote": False, # never store the matched text by default
"limits": {"max_findings_per_request": 100},
}
resp = dlp.inspect_content(request={
"parent": parent,
"inspect_config": inspect_config,
"item": {"value": "Ticket from jo@example.com, EMP-004211, call 555-0100"},
})
for f in resp.result.findings:
loc = f.location.byte_range
print(f.info_type.name, f.likelihood.name, loc.start, loc.end)
De-identification: remove, generalise, pseudonymise or tokenise
Redaction removes a value; de-identification replaces it with something still useful. The transformations fall into three groups, and the choice decides whether the data can ever be joined or restored.
| Group | Transformations | Reversible? | Typical use |
|---|---|---|---|
| Remove or replace | redactConfig, replaceConfig, replaceWithInfoTypeConfig, replaceDictionaryConfig | No | Free text shown to humans or models |
| Generalise | characterMaskConfig, fixedSizeBucketingConfig, bucketingConfig, dateShiftConfig, timePartConfig | No | Analytics that need shape, not identity |
| Pseudonymise | cryptoHashConfig | No (one-way) | Stable join key, never restored |
| Tokenise | cryptoDeterministicConfig (AES-SIV), cryptoReplaceFfxFpeConfig (FFX) | Yes, with the key | Joinable tokens that an authorised service can restore |
Deterministic encryption and format-preserving encryption both produce the same token for the same input under the same key, so tokenised columns still join and group. Google's guidance is to use deterministic encryption unless you must preserve the input's alphabet and length (for example a legacy schema that demands sixteen digits), because FFX is much slower and limits alphabet size and input length. Date shifting keeps intervals within one record consistent when a context field such as a patient ID is supplied, which preserves time-series analysis without real dates.
Reversible tokens with a KMS-wrapped key
Reversible tokenisation needs a key, and the key must not sit in code. The pattern is a data key wrapped by Cloud KMS: generate a 256-bit key once, encrypt it with a KMS key, store only the wrapped bytes, and let the service unwrap it inside each request. The caller's service account needs decrypt permission on the KMS key; anyone without it holds only ciphertext.
crypto_key = {"kms_wrapped": {
"wrapped_key": wrapped_key_bytes, # AES-256 key encrypted by KMS, loaded from config
"crypto_key_name": "projects/p/locations/global/keyRings/dlp/cryptoKeys/tokeniser",
}}
def tokenise(text: str) -> str:
resp = dlp.deidentify_content(request={
"parent": parent,
"inspect_config": {"info_types": [{"name": "EMAIL_ADDRESS"}]},
"deidentify_config": {"info_type_transformations": {"transformations": [{
"info_types": [{"name": "EMAIL_ADDRESS"}],
"primitive_transformation": {"crypto_deterministic_config": {
"crypto_key": crypto_key,
"surrogate_info_type": {"name": "EMAIL_TOKEN"},
}},
}]}},
"item": {"value": text},
})
return resp.item.value # "... EMAIL_TOKEN(52):AfQx... ..."
def restore(text: str) -> str:
resp = dlp.reidentify_content(request={
"parent": parent,
"inspect_config": {"custom_info_types": [
{"info_type": {"name": "EMAIL_TOKEN"}, "surrogate_type": {}}]},
"reidentify_config": {"info_type_transformations": {"transformations": [{
"info_types": [{"name": "EMAIL_TOKEN"}],
"primitive_transformation": {"crypto_deterministic_config": {
"crypto_key": crypto_key,
"surrogate_info_type": {"name": "EMAIL_TOKEN"},
}},
}]}},
"item": {"value": text},
})
return resp.item.valueThe surrogate infoType is what makes restoration possible in free text: the output marks each token with its name and length, and the re-identify call finds tokens by that marker. In structured data you can skip the surrogate and apply the transformation to named columns with record transformations instead.
Worked example: support tickets for analytics and an LLM index
A support organisation wants ticket text in BigQuery for analytics and for a retrieval index used by an LLM assistant, without exposing customer emails, phone numbers or card numbers. The design:
- Tickets arrive on Pub/Sub. A Cloud Run service calls
deidentify_contentwith a stored de-identify template: emails and phone numbers are tokenised with deterministic encryption, card numbers are replaced with their infoType name, and free-text names are masked. - The service writes only transformed text to BigQuery and the retrieval index. Analysts can still count tickets per customer, because the same email always yields the same token.
- A separate re-identification service, with its own service account holding KMS decrypt, restores a single ticket for an agent handling an escalation. Every call is logged with the requester and ticket ID.
- A weekly job trigger inspects the warehouse table itself with the same inspect template and saves findings (without quotes) to a findings table. Any finding means something bypassed the pipeline.
- Discovery profiles the whole project monthly and publishes to Security Command Center, so a new table with high sensitivity shows up as a finding rather than a surprise.
The data flow keeps raw values in exactly two places: the Pub/Sub message in flight and the re-identification response. Retention on the topic and the absence of request logging in the Cloud Run service are therefore part of the control, not details.
Operating it: identity, templates and monitoring
Running the service well is mostly about identity and change control. Inspection jobs and discovery read data as the project's Sensitive Data Protection service agent, a Google-managed service account of the form service-PROJECT_NUMBER@dlp-api.iam.gserviceaccount.com. It needs read access to every bucket, dataset or instance in scope, and write access to the findings dataset. When a job reports success with nothing scanned, a missing grant to this agent is the usual cause.
Separate who configures from who calls. Give the small team that owns templates and scan configurations the administrator role, give pipeline service accounts only the user role they need to call content methods, and keep re-identification on its own account. If your data sits inside a VPC Service Controls perimeter, add the DLP API to the same perimeter so tokenisation calls do not become an exfiltration path.
Treat templates as code. Keep their JSON in version control, apply them through CI, and reference them by full resource name from jobs and services so a rollback is one deploy. Monitor three signals: job state changes (failed jobs often mean revoked access), finding counts per job compared with the previous run (a sudden drop is as suspicious as a spike), and new high-sensitivity profiles from discovery.
# Version-controlled de-identify template, applied from CI
dlp.create_deidentify_template(request={
"parent": "projects/my-project/locations/global",
"template_id": "tickets-v3",
"deidentify_template": {"display_name": "Ticket tokeniser v3",
"deidentify_config": TICKET_DEID_CONFIG},
})
Failure modes
- Findings that leak. With
include_quoteenabled, the findings table stores the matched text, creating a new, well-indexed copy of the sensitive data. Leave it off except for short tuning runs on restricted datasets. - Silent misses. A high
min_likelihood, a missing locale-specific infoType, or a format your regex does not cover all produce zero findings, which looks like success. Seed test records you know are sensitive and assert they are caught. - Lost keys. Destroying or disabling the KMS key version makes every token permanent. Treat it as you would a database encryption key: restricted access, alerts on state changes, documented recovery.
- Unbounded scans. An inspection job over a whole bucket without sampling can cost far more than intended. Start with sampling limits and widen only where profiles show risk.
- Token drift. Changing the key or the surrogate name changes every token and breaks joins with history. Version de-identify templates and rotate keys as a planned migration.
- Context in the clear. Removing identifiers does not remove re-identification risk from quasi-identifiers such as postcode, birth date and job title together; generalise those too.
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Where to transform | Inline content methods: raw data never lands | Scan after landing: simpler, but raw data exists at rest |
| Token type | Deterministic AES-SIV: fast, any input | FFX FPE: keeps format, slow and restricted |
| Join key without restore | cryptoHashConfig: one-way | Tokenisation: restorable, so the key is a liability |
| Coverage | Discovery profiles: broad, coarse, scheduled | Inspection jobs: precise, per-dataset, more config |
| Detection threshold | Lower likelihood: fewer misses | Higher likelihood: fewer false positives |
What to do next
- Turn on discovery for one project and read the resulting profiles before designing anything else.
- Write an inspect template with your custom infoTypes, hotword and exclusion rules, and test it against seeded records.
- Pick per field: redact, generalise, hash or tokenise. Default to deterministic encryption for restorable tokens.
- Create a KMS-wrapped data key, restrict decrypt to a single re-identification service account, and alert on key state changes.
- Audit every findings table for
include_quoteand remove stored quotes. - Pair this with Security Command Center, VPC Service Controls and IAM; for redacting PII in agent prompts, see PII redaction in ADK agents.