Sensitive Data Protection is the current name for what Google Cloud launched as Cloud DLP. The rename widened the product from an inspection API into a family that also profiles data estates continuously, but the API endpoint is still dlp.googleapis.com and the client libraries are still dlp_v2, so code and documentation use both names. This article explains how detection actually works, the three ways to run it and when each fits, how de-identification differs from redaction, and how to build a reversible tokenisation flow without leaking the data you were trying to protect.

Pricing, quotas and request size limits change often and differ between the content methods and storage scans, so this article gives none; check the current pricing and quotas pages before sizing a deployment.

How detection works

Everything starts with an infoType, a named detector such as EMAIL_ADDRESS, CREDIT_CARD_NUMBER or US_SOCIAL_SECURITY_NUMBER. Built-in infoTypes combine patterns, checksums (a card number must pass the Luhn check), and context. Each finding comes back with a likelihood from VERY_UNLIKELY to VERY_LIKELY; the min_likelihood setting drops anything below it. Lowering it finds more and produces more false positives; there is no setting that does both.

You extend detection in three ways. Custom infoTypes can be a regular expression (an internal employee ID), a word list (project code names), or a stored dictionary built from a large term list in Cloud Storage or BigQuery. Hotword rules adjust likelihood when a nearby term appears, so a nine-digit number next to "SSN" is raised and one next to "order" is not. Exclusion rules remove findings, for example an email detector that should ignore your own support address. Rules belong in an inspect template, a stored, versioned configuration that every job and request references by name, so a tuning change lands everywhere at once.

Three ways to run it

Three ways to run Sensitive Data Protection, one detection engineApps / pipelinestext, tables, imagesCloud Storage, BigQueryDatastore, hybridOrg / folder / projectscope for discoveryContent methodsinspect / deidentify (sync)Inspection jobsone-off or job triggersDiscoveryscheduled data profilesDetection engineinfoTypes + rulesTransformed datatokens, masksFindings tableBigQueryPub/Sub, SCCnotificationsProfilesrisk + sensitivityTemplates (inspect and de-identify) are shared by all three paths, so detection rules live in one place.
Figure 1. Content methods, inspection jobs and discovery all use the same infoTypes and templates; they differ in what they read and what they produce.

Content methods (inspect_content, deidentify_content, reidentify_content) are synchronous calls on data you send in the request: a string, a table of rows, or an image. Use them inline in pipelines, for example a Dataflow job or a Cloud Run service that tokenises records before they reach the warehouse.

Inspection jobs read data where it sits in Cloud Storage, BigQuery or Datastore and write results through actions: save findings to a BigQuery table, publish to Pub/Sub, or publish a summary to Security Command Center. A job trigger reruns the job on a schedule. Jobs support sampling (rows_limit and sample_method for BigQuery, bytes_limit_per_file and files_limit_percent for Cloud Storage), which is how you keep a scan of a petabyte bucket affordable. Hybrid jobs let you push content from outside Google Cloud into a job and get the same findings output.

Discovery answers "where is sensitive data at all?" across an organisation, folder or project. A scan configuration names the scope, resource type, inspect templates and reprofiling schedule (daily, weekly or monthly, or on schema or template change). Supported sources include BigQuery, Cloud SQL, Cloud Storage and Vertex AI, plus Amazon S3 and Azure Blob Storage. The output is a data profile per project, table, column or file store, carrying predicted infoTypes and calculated data-risk and sensitivity levels, which can be published to BigQuery, Security Command Center, Knowledge Catalog and Pub/Sub.

Inspecting content from code

A minimal inspection with the Python client. Note the include_quote flag; it is discussed under failure modes.

from google.cloud import dlp_v2

dlp = dlp_v2.DlpServiceClient()
parent = "projects/my-project/locations/global"

inspect_config = {
    "info_types": [{"name": "EMAIL_ADDRESS"}, {"name": "PHONE_NUMBER"}],
    "custom_info_types": [{
        "info_type": {"name": "EMPLOYEE_ID"},
        "regex": {"pattern": r"EMP-\d{6}"},
        "likelihood": "LIKELY",
    }],
    "min_likelihood": "POSSIBLE",
    "include_quote": False,           # never store the matched text by default
    "limits": {"max_findings_per_request": 100},
}

resp = dlp.inspect_content(request={
    "parent": parent,
    "inspect_config": inspect_config,
    "item": {"value": "Ticket from jo@example.com, EMP-004211, call 555-0100"},
})
for f in resp.result.findings:
    loc = f.location.byte_range
    print(f.info_type.name, f.likelihood.name, loc.start, loc.end)

De-identification: remove, generalise, pseudonymise or tokenise

Redaction removes a value; de-identification replaces it with something still useful. The transformations fall into three groups, and the choice decides whether the data can ever be joined or restored.

GroupTransformationsReversible?Typical use
Remove or replaceredactConfig, replaceConfig, replaceWithInfoTypeConfig, replaceDictionaryConfigNoFree text shown to humans or models
GeneralisecharacterMaskConfig, fixedSizeBucketingConfig, bucketingConfig, dateShiftConfig, timePartConfigNoAnalytics that need shape, not identity
PseudonymisecryptoHashConfigNo (one-way)Stable join key, never restored
TokenisecryptoDeterministicConfig (AES-SIV), cryptoReplaceFfxFpeConfig (FFX)Yes, with the keyJoinable tokens that an authorised service can restore

Deterministic encryption and format-preserving encryption both produce the same token for the same input under the same key, so tokenised columns still join and group. Google's guidance is to use deterministic encryption unless you must preserve the input's alphabet and length (for example a legacy schema that demands sixteen digits), because FFX is much slower and limits alphabet size and input length. Date shifting keeps intervals within one record consistent when a context field such as a patient ID is supplied, which preserves time-series analysis without real dates.

Reversible tokens with a KMS-wrapped key

Reversible tokenisation needs a key, and the key must not sit in code. The pattern is a data key wrapped by Cloud KMS: generate a 256-bit key once, encrypt it with a KMS key, store only the wrapped bytes, and let the service unwrap it inside each request. The caller's service account needs decrypt permission on the KMS key; anyone without it holds only ciphertext.

crypto_key = {"kms_wrapped": {
    "wrapped_key": wrapped_key_bytes,     # AES-256 key encrypted by KMS, loaded from config
    "crypto_key_name": "projects/p/locations/global/keyRings/dlp/cryptoKeys/tokeniser",
}}

def tokenise(text: str) -> str:
    resp = dlp.deidentify_content(request={
        "parent": parent,
        "inspect_config": {"info_types": [{"name": "EMAIL_ADDRESS"}]},
        "deidentify_config": {"info_type_transformations": {"transformations": [{
            "info_types": [{"name": "EMAIL_ADDRESS"}],
            "primitive_transformation": {"crypto_deterministic_config": {
                "crypto_key": crypto_key,
                "surrogate_info_type": {"name": "EMAIL_TOKEN"},
            }},
        }]}},
        "item": {"value": text},
    })
    return resp.item.value   # "... EMAIL_TOKEN(52):AfQx... ..."

def restore(text: str) -> str:
    resp = dlp.reidentify_content(request={
        "parent": parent,
        "inspect_config": {"custom_info_types": [
            {"info_type": {"name": "EMAIL_TOKEN"}, "surrogate_type": {}}]},
        "reidentify_config": {"info_type_transformations": {"transformations": [{
            "info_types": [{"name": "EMAIL_TOKEN"}],
            "primitive_transformation": {"crypto_deterministic_config": {
                "crypto_key": crypto_key,
                "surrogate_info_type": {"name": "EMAIL_TOKEN"},
            }},
        }]}},
        "item": {"value": text},
    })
    return resp.item.value

The surrogate infoType is what makes restoration possible in free text: the output marks each token with its name and length, and the re-identify call finds tokens by that marker. In structured data you can skip the surrogate and apply the transformation to named columns with record transformations instead.

Worked example: support tickets for analytics and an LLM index

A support organisation wants ticket text in BigQuery for analytics and for a retrieval index used by an LLM assistant, without exposing customer emails, phone numbers or card numbers. The design:

  1. Tickets arrive on Pub/Sub. A Cloud Run service calls deidentify_content with a stored de-identify template: emails and phone numbers are tokenised with deterministic encryption, card numbers are replaced with their infoType name, and free-text names are masked.
  2. The service writes only transformed text to BigQuery and the retrieval index. Analysts can still count tickets per customer, because the same email always yields the same token.
  3. A separate re-identification service, with its own service account holding KMS decrypt, restores a single ticket for an agent handling an escalation. Every call is logged with the requester and ticket ID.
  4. A weekly job trigger inspects the warehouse table itself with the same inspect template and saves findings (without quotes) to a findings table. Any finding means something bypassed the pipeline.
  5. Discovery profiles the whole project monthly and publishes to Security Command Center, so a new table with high sensitivity shows up as a finding rather than a surprise.

The data flow keeps raw values in exactly two places: the Pub/Sub message in flight and the re-identification response. Retention on the topic and the absence of request logging in the Cloud Run service are therefore part of the control, not details.

Operating it: identity, templates and monitoring

Running the service well is mostly about identity and change control. Inspection jobs and discovery read data as the project's Sensitive Data Protection service agent, a Google-managed service account of the form service-PROJECT_NUMBER@dlp-api.iam.gserviceaccount.com. It needs read access to every bucket, dataset or instance in scope, and write access to the findings dataset. When a job reports success with nothing scanned, a missing grant to this agent is the usual cause.

Separate who configures from who calls. Give the small team that owns templates and scan configurations the administrator role, give pipeline service accounts only the user role they need to call content methods, and keep re-identification on its own account. If your data sits inside a VPC Service Controls perimeter, add the DLP API to the same perimeter so tokenisation calls do not become an exfiltration path.

Treat templates as code. Keep their JSON in version control, apply them through CI, and reference them by full resource name from jobs and services so a rollback is one deploy. Monitor three signals: job state changes (failed jobs often mean revoked access), finding counts per job compared with the previous run (a sudden drop is as suspicious as a spike), and new high-sensitivity profiles from discovery.

# Version-controlled de-identify template, applied from CI
dlp.create_deidentify_template(request={
    "parent": "projects/my-project/locations/global",
    "template_id": "tickets-v3",
    "deidentify_template": {"display_name": "Ticket tokeniser v3",
                            "deidentify_config": TICKET_DEID_CONFIG},
})

Failure modes

  • Findings that leak. With include_quote enabled, the findings table stores the matched text, creating a new, well-indexed copy of the sensitive data. Leave it off except for short tuning runs on restricted datasets.
  • Silent misses. A high min_likelihood, a missing locale-specific infoType, or a format your regex does not cover all produce zero findings, which looks like success. Seed test records you know are sensitive and assert they are caught.
  • Lost keys. Destroying or disabling the KMS key version makes every token permanent. Treat it as you would a database encryption key: restricted access, alerts on state changes, documented recovery.
  • Unbounded scans. An inspection job over a whole bucket without sampling can cost far more than intended. Start with sampling limits and widen only where profiles show risk.
  • Token drift. Changing the key or the surrogate name changes every token and breaks joins with history. Version de-identify templates and rotate keys as a planned migration.
  • Context in the clear. Removing identifiers does not remove re-identification risk from quasi-identifiers such as postcode, birth date and job title together; generalise those too.

Trade-offs

DecisionOption AOption B
Where to transformInline content methods: raw data never landsScan after landing: simpler, but raw data exists at rest
Token typeDeterministic AES-SIV: fast, any inputFFX FPE: keeps format, slow and restricted
Join key without restorecryptoHashConfig: one-wayTokenisation: restorable, so the key is a liability
CoverageDiscovery profiles: broad, coarse, scheduledInspection jobs: precise, per-dataset, more config
Detection thresholdLower likelihood: fewer missesHigher likelihood: fewer false positives

What to do next

  1. Turn on discovery for one project and read the resulting profiles before designing anything else.
  2. Write an inspect template with your custom infoTypes, hotword and exclusion rules, and test it against seeded records.
  3. Pick per field: redact, generalise, hash or tokenise. Default to deterministic encryption for restorable tokens.
  4. Create a KMS-wrapped data key, restrict decrypt to a single re-identification service account, and alert on key state changes.
  5. Audit every findings table for include_quote and remove stored quotes.
  6. Pair this with Security Command Center, VPC Service Controls and IAM; for redacting PII in agent prompts, see PII redaction in ADK agents.
Key takeaway: Sensitive Data Protection is one detection engine, infoTypes plus likelihood plus rules kept in templates, behind three ways of running it: synchronous content methods for pipelines, inspection jobs for data at rest, and discovery for organisation-wide profiles. Transform before data lands when you can, use deterministic AES-SIV encryption with a KMS-wrapped key when tokens must be restorable, keep quotes out of findings, and test detection with seeded records so a miss cannot pass for a clean result.