Dataplex is Google Cloud's governance layer for analytical data. It does four jobs that teams otherwise stitch together by hand: it keeps a catalog of what data exists and what it means, it runs data quality and profiling scans against tables, it records lineage between datasets, and it can group storage into lakes and zones for management. When an upstream column silently changes meaning, Dataplex is where that becomes visible and, wired in, where it stops a bad load.

Naming first, because it trips up every search. The product was called Dataplex, then Dataplex Universal Catalog. Since 10 April 2026 Google calls it Knowledge Catalog. The API is still dataplex.googleapis.com, the CLI is still gcloud dataplex and the IAM roles still start with roles/dataplex. The older Data Catalog service was deprecated in February 2025 and shut down on 1 June 2026, so anything you read that tells you to tag tables in Data Catalog is out of date. This article uses "Dataplex" for the service and its API.

Three planes over your storage

One governance plane over many stores: catalog, scans and lineageBigQuery tables@bigquery entry groupCloud Storagebuckets, object tablesOther sourcesconnectors, import jobsCatalog (entries)entry = one assetentry type: required aspectsaspect = typed metadataaspect type = its schemaentry group = IAM boundarysearch + business glossarylakes / zones / assetsmetadatadiscoveryimportData quality scanrules, thresholds, scoreData profile scanstats per columnData lineageprocess, run, eventPipeline gate (CI, Composer)run scan, read passed + score, block or publishjob resultBigQuery exportresults table, dashboardsProduct name since 2026-04-10: Knowledge Catalog. API, gcloud and IAM names stay dataplex.
Sources feed the catalog; scans and lineage attach results to catalog entries; a pipeline gate reads scan results.

Think of Dataplex as three planes over the same storage. The metadata plane is the catalog: one entry per table, file set or model, with typed metadata attached. The verification plane is the scan engine, which runs SQL against BigQuery-backed data and writes results back to the catalog and to BigQuery. The provenance plane is lineage, which records which job read which table and wrote which other table.

None of these planes moves your data, so Dataplex can be adopted one table at a time.

The catalog model

The catalog has five nouns, and once they click the rest of the product reads naturally.

NounWhat it isAnalogy
EntryOne data asset: a BigQuery table, an object table, a model, or a custom asset you registerA row in an inventory
Entry groupA regional container of entries; the administrative and IAM boundaryA folder with its own permissions
Aspect typeA schema for a block of metadata: field names, types, validationA class definition
AspectAn instance of an aspect type attached to an entry, a column path or an entry linkAn object of that class
Entry typeA template that says which aspect types an entry of this kind must carryAn interface the entry must implement

Google maintains system entry groups such as @bigquery, so every BigQuery table you can see already has an entry without you doing anything. Your work is to define the aspect types your organisation cares about and attach aspects. A typical first aspect type is ownership: owning team, on-call alias, data classification and a retention class. Attach it to every production table and you can answer "who do I page about this table?" from search instead of from tribal memory.

Aspects can also attach to a column path, which is where classification belongs: sensitivity is a property of customers.email, not of the whole table.

Entry types are the enforcement hook. For custom assets, define an entry type that requires your ownership aspect, and an entry of that type cannot be created without an owner.

Lakes, zones and assets

Lakes, zones and assets are the older organising layer, and you will still meet them. A lake is a management domain, often one per business domain. A zone groups assets inside a lake, and comes in two kinds: raw zones for data as it lands and curated zones for cleaned, schema-stable data. An asset attaches a Cloud Storage bucket or a BigQuery dataset to a zone. Discovery, enabled by default on new zones and assets, scans the attached storage and registers tables it finds so they can be queried.

Ingestion of lake and zone entities into Data Catalog was shut down on 30 September 2025. Lakes are optional for everything else in this article: entries, aspects, scans and lineage work on BigQuery tables directly. Use lakes for domain grouping and discovery of Cloud Storage files.

How data quality scans decide pass and fail

A data quality scan is a saved job: a data source, a list of rules, optional sampling and filtering, and actions to run afterwards. Each run produces a job whose result says, per rule, how many rows were evaluated and how many passed, plus an overall boolean and a score from 0 to 100.

Rules come in three shapes, and the distinction decides how thresholds behave.

  • Row-level rules test each row and pass if the fraction of passing rows meets the rule's threshold. These are nonNullExpectation, rangeExpectation, setExpectation, regexExpectation and the custom rowConditionExpectation.
  • Aggregate rules compute one value over the data and pass or fail as a whole: uniquenessExpectation, statisticRangeExpectation (mean, min or max within a range) and the custom tableConditionExpectation.
  • SQL assertions (sqlAssertion) run your statement and fail if it returns any rows. Write ${data()} where the table name would go; Dataplex substitutes the source with your row filter, sampling and incremental window already applied, so the assertion checks exactly what the other rules check.

Every rule also carries a dimension for reporting. The seven standard dimensions are completeness, validity, uniqueness, consistency, accuracy, freshness and volume. A scan may hold up to 1,000 rules. When a rule fails, the result includes failingRowsQuery, a ready-made query that returns the offending rows, which turns "the scan is red" into "here are the 312 rows" in one click.

Before writing rules, run a data profile scan on the same table (gcloud dataplex datascans create data-profile). It reports per-column statistics such as null ratios, distinct counts and value ranges. Profile first, then set thresholds from what the data actually looks like on a good day; rules written from memory either fire constantly or never.

Worked example: an orders table

A retailer loads about 2.4 million orders a day into sales.orders, partitioned by order_ts. Downstream, finance reports read the table at 06:00. The team wants five checks, scanned incrementally on the ingested_at load timestamp so each run only reads new rows. Here is the spec, using the field names from the DataQualityRule API:

# orders_dq.yaml
rowFilter: "channel != 'TEST'"
rules:
  - column: order_id
    dimension: UNIQUENESS
    uniquenessExpectation: {}
  - column: customer_id
    dimension: COMPLETENESS
    nonNullExpectation: {}
    threshold: 1.0
  - column: amount_minor
    dimension: VALIDITY
    ignoreNull: true
    rangeExpectation:
      minValue: "0"
      maxValue: "5000000"
    threshold: 0.999
  - column: status
    dimension: VALIDITY
    setExpectation:
      values: ["NEW", "PAID", "SHIPPED", "CANCELLED", "REFUNDED"]
    threshold: 1.0
  - dimension: CONSISTENCY
    name: shipped-after-created
    sqlAssertion:
      sqlStatement: >
        SELECT order_id FROM ${data()}
        WHERE shipped_at IS NOT NULL AND shipped_at < created_at
postScanActions:
  bigqueryExport:
    resultsTable: //bigquery.googleapis.com/projects/acme-dq/datasets/dq_results/tables/scan_results
gcloud dataplex datascans create data-quality orders-dq \
  --project=acme-sales --location=europe-west1 \
  --data-source-resource="//bigquery.googleapis.com/projects/acme-sales/datasets/sales/tables/orders" \
  --data-quality-spec-file=orders_dq.yaml \
  --incremental-field=ingested_at \
  --on-demand \
  --service-account=dq-runner@acme-dq.iam.gserviceaccount.com

Now follow one morning's run by hand. Of 2,400,000 new rows, all have a customer id, so the null rule passes at 1.0. 1,800 rows have amounts above the cap because a bulk B2B order type started this week; the pass ratio is 1 − 1,800 / 2,400,000 = 0.99925, which clears 0.999, so the range rule passes and the outliers show up in the exported results for someone to look at. Then 312 rows carry a new status, PARTIALLY_REFUNDED, shipped by the payments team without telling anyone. The set rule's ratio is 0.99987, below its threshold of 1.0, so it fails and the job's overall passed is false.

That is the point of threshold 1.0 on enumerations. A new enum value is a contract change upstream, and the finance query grouping by status would silently drop those orders. The scan turns that drift into a page before 06:00.

Turning a scan into a pipeline gate

A scan that only feeds a dashboard is a smoke detector with no siren. Make it a gate: run the scan after the load, wait for the job, and block publication if it failed. The script below separates the two kinds of failure, because they need different responses. Exit 1 means the data is bad; exit 2 means the check itself did not complete, and nobody knows whether the data is bad.

#!/usr/bin/env bash
# dq_gate.sh: run an on-demand scan and fail the pipeline on bad data.
set -Eeuo pipefail
trap 'echo "dq gate: gcloud call failed" >&2; exit 2' ERR   # any CLI/API error = check did not run
SCAN=orders-dq LOC=europe-west1 PROJECT=acme-sales

JOB=$(gcloud dataplex datascans run "$SCAN" --location="$LOC" \
        --project="$PROJECT" --format="value(job.name)")
JOB_ID=${JOB##*/}

STATE=PENDING
for _ in $(seq 1 60); do                      # up to 30 minutes
  STATE=$(gcloud dataplex datascans jobs describe "$JOB_ID" \
            --datascan="$SCAN" --location="$LOC" --project="$PROJECT" \
            --format="value(state)")
  case "$STATE" in SUCCEEDED|FAILED|CANCELLED) break ;; esac
  sleep 30
done

if [ "$STATE" != "SUCCEEDED" ]; then
  echo "dq gate: scan job $JOB_ID ended in state $STATE" >&2
  exit 2                                      # check did not run: do not publish, do not blame data
fi

PASSED=$(gcloud dataplex datascans jobs describe "$JOB_ID" \
           --datascan="$SCAN" --location="$LOC" --project="$PROJECT" \
           --view=FULL --format="value(dataQualityResult.passed)")
if [ "$PASSED" != "True" ]; then
  echo "dq gate: data quality failed; see failingRowsQuery in job $JOB_ID" >&2
  exit 1
fi
echo "dq gate: passed"

Three details matter. The run method only works on on-demand scans, which is why the scan was created with --on-demand. The ERR trap maps any failing gcloud call, such as a permissions error, to exit 2 rather than letting it masquerade as bad data. The --view=FULL flag is required to get the result at all; the default basic view omits it. And the gate treats a timeout as failure, never as success: a gate that passes when the scan hangs is worse than no gate, because it advertises safety it does not provide.

Stronger still: load into a staging table, scan that, and swap it into place only when the scan passes, so readers never see bad rows at all.

Lineage

Lineage answers two questions: what feeds this table, and what breaks if I change it. Dataplex records lineage automatically for supported services such as BigQuery jobs, as a graph of processes, runs and lineage events between fully qualified names. For anything it cannot see, such as a custom Spark job writing to Cloud Storage, you report lineage yourself through the Data Lineage API: create a process for the job, a run for each execution, and an event naming its sources and targets.

When orders-dq fails, walk lineage downstream to find affected reports and upstream to find the job that wrote the bad rows.

Operating scans

Scans run as a service account, so give that account read access on the source tables and write access only on the results dataset. Keep scan definitions in version control, as YAML specs applied by CI, rather than edited in the console. Export results to BigQuery and chart score per table per day: a slow slide from 99.9 to 98 is a warning no single run gives you.

Scans are billed for the processing they do, so sampling (samplingPercent) and incremental scans are your cost levers. Sample for statistical rules on huge tables; never sample for uniqueness or for enum checks with threshold 1.0, where a sample can miss the one bad row that matters.

Failure modes

SymptomCauseFix
Scan always passes, incidents still happenThresholds set below normal noise, or rules on the wrong columnsProfile first; set thresholds from good-day data; add rules after each incident
Gate passes while the scan is still runningScript treats a timeout or unknown state as successExit non-zero on anything other than SUCCEEDED
run fails on a scheduled scanRun is only supported for on-demand scansCreate gate scans with --on-demand
Uniqueness passes on sampled dataA sample rarely contains both duplicatesNo sampling on uniqueness or exact enum rules
Incremental scan misses late rowsRows arrive with an old order_tsUse an ingestion timestamp as the incremental field
Permission denied in scan jobRunner account lacks read on source or write on resultsGrant per table and dataset; avoid project-wide editor
Old automation tags nothingIt targets Data Catalog, shut down 1 June 2026Port tags to aspect types and aspects

Trade-offs

Dataplex versus code-first tests such as dbt: Dataplex rules need no runtime of your own and publish results to the catalog; code-first tools give richer logic and local testing. Many teams use dbt tests in CI and Dataplex scans on published tables.

Strict versus lenient thresholds is a paging trade-off. Threshold 1.0 catches contract changes the hour they happen but pages on single bad rows; 0.999 tolerates noise but lets a small, real regression through. Use 1.0 for keys, enumerations and referential checks, and ratios for measurements.

For more on the tables these scans protect, see BigQuery in depth and BigQuery partitioning and clustering; for orchestration, Cloud Composer; for the service account model, GCP IAM; and for finding PII to record as aspects, Cloud DLP.

What to do next

  1. Pick the three tables whose silent corruption would hurt most, and run a data profile scan on each.
  2. Define one ownership aspect type (team, on-call, classification) and attach it to those tables.
  3. Write a quality spec per table: keys unique, required columns non-null, enumerations at threshold 1.0, ranges at measured ratios.
  4. Create the scans with --on-demand and a dedicated runner service account with least privilege.
  5. Put a gate like dq_gate.sh between load and publish, with exit 1 for bad data and exit 2 for a failed check.
  6. Export results to BigQuery and chart score per table per day.
  7. Search your automation for Data Catalog calls and port them to aspects.
  8. After the next data incident, add the rule that would have caught it before closing the postmortem.
Key takeaway: Dataplex, now named Knowledge Catalog with unchanged API and CLI names, gives analytical data a catalog of typed metadata, rule-based quality scans and lineage. Model ownership and sensitivity as aspects, profile tables before writing rules, use threshold 1.0 for keys and enumerations, and make an on-demand scan a gate between load and publish that fails on bad data and on a check that did not finish. Move anything still pointed at Data Catalog, which shut down on 1 June 2026.