Dataplex is Google Cloud's governance layer for analytical data. It does four jobs that teams otherwise stitch together by hand: it keeps a catalog of what data exists and what it means, it runs data quality and profiling scans against tables, it records lineage between datasets, and it can group storage into lakes and zones for management. When an upstream column silently changes meaning, Dataplex is where that becomes visible and, wired in, where it stops a bad load.
Naming first, because it trips up every search. The product was called Dataplex, then Dataplex Universal Catalog. Since 10 April 2026 Google calls it Knowledge Catalog. The API is still dataplex.googleapis.com, the CLI is still gcloud dataplex and the IAM roles still start with roles/dataplex. The older Data Catalog service was deprecated in February 2025 and shut down on 1 June 2026, so anything you read that tells you to tag tables in Data Catalog is out of date. This article uses "Dataplex" for the service and its API.
Three planes over your storage
Think of Dataplex as three planes over the same storage. The metadata plane is the catalog: one entry per table, file set or model, with typed metadata attached. The verification plane is the scan engine, which runs SQL against BigQuery-backed data and writes results back to the catalog and to BigQuery. The provenance plane is lineage, which records which job read which table and wrote which other table.
None of these planes moves your data, so Dataplex can be adopted one table at a time.
The catalog model
The catalog has five nouns, and once they click the rest of the product reads naturally.
| Noun | What it is | Analogy |
|---|---|---|
| Entry | One data asset: a BigQuery table, an object table, a model, or a custom asset you register | A row in an inventory |
| Entry group | A regional container of entries; the administrative and IAM boundary | A folder with its own permissions |
| Aspect type | A schema for a block of metadata: field names, types, validation | A class definition |
| Aspect | An instance of an aspect type attached to an entry, a column path or an entry link | An object of that class |
| Entry type | A template that says which aspect types an entry of this kind must carry | An interface the entry must implement |
Google maintains system entry groups such as @bigquery, so every BigQuery table you can see already has an entry without you doing anything. Your work is to define the aspect types your organisation cares about and attach aspects. A typical first aspect type is ownership: owning team, on-call alias, data classification and a retention class. Attach it to every production table and you can answer "who do I page about this table?" from search instead of from tribal memory.
Aspects can also attach to a column path, which is where classification belongs: sensitivity is a property of customers.email, not of the whole table.
Entry types are the enforcement hook. For custom assets, define an entry type that requires your ownership aspect, and an entry of that type cannot be created without an owner.
Lakes, zones and assets
Lakes, zones and assets are the older organising layer, and you will still meet them. A lake is a management domain, often one per business domain. A zone groups assets inside a lake, and comes in two kinds: raw zones for data as it lands and curated zones for cleaned, schema-stable data. An asset attaches a Cloud Storage bucket or a BigQuery dataset to a zone. Discovery, enabled by default on new zones and assets, scans the attached storage and registers tables it finds so they can be queried.
Ingestion of lake and zone entities into Data Catalog was shut down on 30 September 2025. Lakes are optional for everything else in this article: entries, aspects, scans and lineage work on BigQuery tables directly. Use lakes for domain grouping and discovery of Cloud Storage files.
How data quality scans decide pass and fail
A data quality scan is a saved job: a data source, a list of rules, optional sampling and filtering, and actions to run afterwards. Each run produces a job whose result says, per rule, how many rows were evaluated and how many passed, plus an overall boolean and a score from 0 to 100.
Rules come in three shapes, and the distinction decides how thresholds behave.
- Row-level rules test each row and pass if the fraction of passing rows meets the rule's
threshold. These arenonNullExpectation,rangeExpectation,setExpectation,regexExpectationand the customrowConditionExpectation. - Aggregate rules compute one value over the data and pass or fail as a whole:
uniquenessExpectation,statisticRangeExpectation(mean, min or max within a range) and the customtableConditionExpectation. - SQL assertions (
sqlAssertion) run your statement and fail if it returns any rows. Write${data()}where the table name would go; Dataplex substitutes the source with your row filter, sampling and incremental window already applied, so the assertion checks exactly what the other rules check.
Every rule also carries a dimension for reporting. The seven standard dimensions are completeness, validity, uniqueness, consistency, accuracy, freshness and volume. A scan may hold up to 1,000 rules. When a rule fails, the result includes failingRowsQuery, a ready-made query that returns the offending rows, which turns "the scan is red" into "here are the 312 rows" in one click.
Before writing rules, run a data profile scan on the same table (gcloud dataplex datascans create data-profile). It reports per-column statistics such as null ratios, distinct counts and value ranges. Profile first, then set thresholds from what the data actually looks like on a good day; rules written from memory either fire constantly or never.
Worked example: an orders table
A retailer loads about 2.4 million orders a day into sales.orders, partitioned by order_ts. Downstream, finance reports read the table at 06:00. The team wants five checks, scanned incrementally on the ingested_at load timestamp so each run only reads new rows. Here is the spec, using the field names from the DataQualityRule API:
# orders_dq.yaml
rowFilter: "channel != 'TEST'"
rules:
- column: order_id
dimension: UNIQUENESS
uniquenessExpectation: {}
- column: customer_id
dimension: COMPLETENESS
nonNullExpectation: {}
threshold: 1.0
- column: amount_minor
dimension: VALIDITY
ignoreNull: true
rangeExpectation:
minValue: "0"
maxValue: "5000000"
threshold: 0.999
- column: status
dimension: VALIDITY
setExpectation:
values: ["NEW", "PAID", "SHIPPED", "CANCELLED", "REFUNDED"]
threshold: 1.0
- dimension: CONSISTENCY
name: shipped-after-created
sqlAssertion:
sqlStatement: >
SELECT order_id FROM ${data()}
WHERE shipped_at IS NOT NULL AND shipped_at < created_at
postScanActions:
bigqueryExport:
resultsTable: //bigquery.googleapis.com/projects/acme-dq/datasets/dq_results/tables/scan_resultsgcloud dataplex datascans create data-quality orders-dq \
--project=acme-sales --location=europe-west1 \
--data-source-resource="//bigquery.googleapis.com/projects/acme-sales/datasets/sales/tables/orders" \
--data-quality-spec-file=orders_dq.yaml \
--incremental-field=ingested_at \
--on-demand \
--service-account=dq-runner@acme-dq.iam.gserviceaccount.comNow follow one morning's run by hand. Of 2,400,000 new rows, all have a customer id, so the null rule passes at 1.0. 1,800 rows have amounts above the cap because a bulk B2B order type started this week; the pass ratio is 1 − 1,800 / 2,400,000 = 0.99925, which clears 0.999, so the range rule passes and the outliers show up in the exported results for someone to look at. Then 312 rows carry a new status, PARTIALLY_REFUNDED, shipped by the payments team without telling anyone. The set rule's ratio is 0.99987, below its threshold of 1.0, so it fails and the job's overall passed is false.
That is the point of threshold 1.0 on enumerations. A new enum value is a contract change upstream, and the finance query grouping by status would silently drop those orders. The scan turns that drift into a page before 06:00.
Turning a scan into a pipeline gate
A scan that only feeds a dashboard is a smoke detector with no siren. Make it a gate: run the scan after the load, wait for the job, and block publication if it failed. The script below separates the two kinds of failure, because they need different responses. Exit 1 means the data is bad; exit 2 means the check itself did not complete, and nobody knows whether the data is bad.
#!/usr/bin/env bash
# dq_gate.sh: run an on-demand scan and fail the pipeline on bad data.
set -Eeuo pipefail
trap 'echo "dq gate: gcloud call failed" >&2; exit 2' ERR # any CLI/API error = check did not run
SCAN=orders-dq LOC=europe-west1 PROJECT=acme-sales
JOB=$(gcloud dataplex datascans run "$SCAN" --location="$LOC" \
--project="$PROJECT" --format="value(job.name)")
JOB_ID=${JOB##*/}
STATE=PENDING
for _ in $(seq 1 60); do # up to 30 minutes
STATE=$(gcloud dataplex datascans jobs describe "$JOB_ID" \
--datascan="$SCAN" --location="$LOC" --project="$PROJECT" \
--format="value(state)")
case "$STATE" in SUCCEEDED|FAILED|CANCELLED) break ;; esac
sleep 30
done
if [ "$STATE" != "SUCCEEDED" ]; then
echo "dq gate: scan job $JOB_ID ended in state $STATE" >&2
exit 2 # check did not run: do not publish, do not blame data
fi
PASSED=$(gcloud dataplex datascans jobs describe "$JOB_ID" \
--datascan="$SCAN" --location="$LOC" --project="$PROJECT" \
--view=FULL --format="value(dataQualityResult.passed)")
if [ "$PASSED" != "True" ]; then
echo "dq gate: data quality failed; see failingRowsQuery in job $JOB_ID" >&2
exit 1
fi
echo "dq gate: passed"Three details matter. The run method only works on on-demand scans, which is why the scan was created with --on-demand. The ERR trap maps any failing gcloud call, such as a permissions error, to exit 2 rather than letting it masquerade as bad data. The --view=FULL flag is required to get the result at all; the default basic view omits it. And the gate treats a timeout as failure, never as success: a gate that passes when the scan hangs is worse than no gate, because it advertises safety it does not provide.
Stronger still: load into a staging table, scan that, and swap it into place only when the scan passes, so readers never see bad rows at all.
Lineage
Lineage answers two questions: what feeds this table, and what breaks if I change it. Dataplex records lineage automatically for supported services such as BigQuery jobs, as a graph of processes, runs and lineage events between fully qualified names. For anything it cannot see, such as a custom Spark job writing to Cloud Storage, you report lineage yourself through the Data Lineage API: create a process for the job, a run for each execution, and an event naming its sources and targets.
When orders-dq fails, walk lineage downstream to find affected reports and upstream to find the job that wrote the bad rows.
Operating scans
Scans run as a service account, so give that account read access on the source tables and write access only on the results dataset. Keep scan definitions in version control, as YAML specs applied by CI, rather than edited in the console. Export results to BigQuery and chart score per table per day: a slow slide from 99.9 to 98 is a warning no single run gives you.
Scans are billed for the processing they do, so sampling (samplingPercent) and incremental scans are your cost levers. Sample for statistical rules on huge tables; never sample for uniqueness or for enum checks with threshold 1.0, where a sample can miss the one bad row that matters.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Scan always passes, incidents still happen | Thresholds set below normal noise, or rules on the wrong columns | Profile first; set thresholds from good-day data; add rules after each incident |
| Gate passes while the scan is still running | Script treats a timeout or unknown state as success | Exit non-zero on anything other than SUCCEEDED |
run fails on a scheduled scan | Run is only supported for on-demand scans | Create gate scans with --on-demand |
| Uniqueness passes on sampled data | A sample rarely contains both duplicates | No sampling on uniqueness or exact enum rules |
| Incremental scan misses late rows | Rows arrive with an old order_ts | Use an ingestion timestamp as the incremental field |
| Permission denied in scan job | Runner account lacks read on source or write on results | Grant per table and dataset; avoid project-wide editor |
| Old automation tags nothing | It targets Data Catalog, shut down 1 June 2026 | Port tags to aspect types and aspects |
Trade-offs
Dataplex versus code-first tests such as dbt: Dataplex rules need no runtime of your own and publish results to the catalog; code-first tools give richer logic and local testing. Many teams use dbt tests in CI and Dataplex scans on published tables.
Strict versus lenient thresholds is a paging trade-off. Threshold 1.0 catches contract changes the hour they happen but pages on single bad rows; 0.999 tolerates noise but lets a small, real regression through. Use 1.0 for keys, enumerations and referential checks, and ratios for measurements.
For more on the tables these scans protect, see BigQuery in depth and BigQuery partitioning and clustering; for orchestration, Cloud Composer; for the service account model, GCP IAM; and for finding PII to record as aspects, Cloud DLP.
What to do next
- Pick the three tables whose silent corruption would hurt most, and run a data profile scan on each.
- Define one ownership aspect type (team, on-call, classification) and attach it to those tables.
- Write a quality spec per table: keys unique, required columns non-null, enumerations at threshold 1.0, ranges at measured ratios.
- Create the scans with
--on-demandand a dedicated runner service account with least privilege. - Put a gate like
dq_gate.shbetween load and publish, with exit 1 for bad data and exit 2 for a failed check. - Export results to BigQuery and chart score per table per day.
- Search your automation for Data Catalog calls and port them to aspects.
- After the next data incident, add the rule that would have caught it before closing the postmortem.