Most organisations on AWS reach the same point: the data lake has hundreds of tables across a dozen accounts, nobody can find the right one, and every access request is a ticket that a platform engineer resolves by hand in Lake Formation or Redshift. Amazon DataZone addresses that. It is a catalog with business context, plus a publish-and-subscribe workflow that turns an approved request into the actual grants.

This article explains DataZone's model from first principles (domains, projects, environments and assets), shows how metadata gets in and how access goes out, automates both with boto3, and walks through a two-team example. Since 2025 DataZone also underpins Amazon SageMaker Catalog in SageMaker Unified Studio, so the same concepts apply there. Service details were checked against the AWS documentation on 3 October 2026.

How data flows through DataZone

Producer projectSales analyticsData source runGlue database metadataProject inventorycurate: names, glossary, formsDomain catalogpublished listingsConsumer projectMarketingSubscription requestowner approves or rejectsFulfilmentLake Formation / Redshift grantsUnmanaged assetEventBridge event, you grantimportpublishsearchsubscribeapproveother typesManaged assets (Glue tables, Redshift tables and views) are granted automatically into the consumer's environments; others are handed to your automation.
A producer imports and curates metadata, publishes it to the domain catalog, and a consumer's approved subscription turns into grants in its environments.

The model: domains, projects and environments

Everything sits inside a domain, the top-level container for assets, users and projects. You can run one domain for the enterprise or one per business unit. AWS accounts holding data are associated with the domain; an account owner must accept the association, and one account can belong to several domains.

Inside a domain, domain units mirror your organisation, such as Finance and Marketing, and carry authorization policies that say who may create projects, domain units, glossaries, metadata forms and custom asset types, and who may create environments from which blueprints. This is how a central team delegates without handing out domain admin.

A project is a team working on a use case. It owns assets and is the unit that subscribes to data. Members have one of five roles:

RolePublish and ingestRequest subscriptionsApprove requestsAdd members
OwnerYesYesYesYes
ContributorYesYesYesNo
StewardYesNoYesNo
ConsumerNoYesNoNo
ViewerNoNoNoNo

An environment is the set of real AWS resources a project works in, such as a Glue database, an S3 location and an Athena workgroup, plus the IAM principals allowed to use them. Environments are created from environment profiles, which are blueprints with preset parameters such as account and Region. The default blueprints are Data Lake (Lake Formation-managed tables queried with Athena), Data Warehouse (your own Redshift clusters) and SageMaker. Subscribed data lands in a project's environments, which is why environments matter for access.

Getting metadata in: data sources, inventory and publishing

An asset is one data object, physical like a table or file or virtual like a view, validated against an asset type. Assets enter a project's inventory, visible only to that project, through a data source: a connection to the Glue Data Catalog or Redshift that imports technical metadata (table names, columns, types) on each data source run. Assets can also be created through the API, including custom types for models, dashboards or on-premises tables.

Producers then curate the inventory: business names, descriptions, a readme, glossary terms on the asset and its columns, and metadata forms, which are typed templates a domain admin defines for things like data classification or retention. Every edit creates a new inventory version. Publishing puts the latest version in the domain catalog. If you edit after publishing, you must publish again; the catalog keeps the older version until you do.

A data source can publish on import. That is convenient and dangerous: tables appear in the catalog with no business metadata, discoverable by every domain user. Use it only for sources whose curation is automated.

Getting access out: subscriptions and fulfilment

A consumer finds a listing and asks for access on behalf of a project, giving a reason. The owning project is notified and approves or rejects. On approval DataZone runs fulfilment: for managed assets (Glue tables and Redshift tables and views) it creates the Lake Formation or Redshift grants that make the data queryable in each of the consumer project's environments. Where it puts the grant is set by the environment's subscription target, a location plus the role DataZone uses.

For unmanaged assets DataZone cannot grant anything itself. It publishes an EventBridge event with the subscription details; your automation grants access, then calls the API to mark the grant complete so members are told they can use the data. Owners can also narrow what a subscriber sees with asset filters on rows and columns, so a region-limited team gets only its rows.

Automating it with boto3

The console works for a pilot, but production domains are run as code. This boto3 sketch registers a Glue source, runs it, finds a listing, and handles a request. The identifiers are placeholders, and the API reference does not enumerate the type string, so confirm GLUE against a data source your domain already lists:

import boto3, time

dz = boto3.client("datazone")
DOMAIN, SALES, MKT = "dzd_example123", "prj_sales", "prj_marketing"

# 1. Producer: import technical metadata from one Glue database, without auto-publishing.
src = dz.create_data_source(
    domainIdentifier=DOMAIN, projectIdentifier=SALES, environmentIdentifier="env_sales_lake",
    name="sales-curated", type="GLUE", publishOnImport=False,
    configuration={"glueRunConfiguration": {"relationalFilterConfigurations": [{
        "databaseName": "sales_curated",
        "filterExpressions": [{"type": "INCLUDE", "expression": "orders*"},
                              {"type": "EXCLUDE", "expression": "*_tmp"}]}]}},
)
run = dz.start_data_source_run(domainIdentifier=DOMAIN, dataSourceIdentifier=src["id"])
while dz.get_data_source_run(domainIdentifier=DOMAIN, identifier=run["id"])["status"] in ("REQUESTED", "RUNNING"):
    time.sleep(15)

# 2. Consumer: search the catalog and request access for the marketing project.
hits = dz.search_listings(domainIdentifier=DOMAIN, searchText="orders", maxResults=10)["items"]
listing = hits[0]["assetListing"]
req = dz.create_subscription_request(
    domainIdentifier=DOMAIN,
    subscribedPrincipals=[{"project": {"identifier": MKT}}],
    subscribedListings=[{"identifier": listing["listingId"]}],
    requestReason="Campaign attribution for Q4; read-only, EU rows only",
)

# 3. Owner automation: approve requests whose reason and requester meet policy.
dz.accept_subscription_request(domainIdentifier=DOMAIN, identifier=req["id"],
                               decisionComment="Approved under data-sharing policy DS-7")

In a real pipeline, step 3 would be a human approval or a policy check, not a blind accept. For unmanaged assets, an EventBridge rule triggers a Lambda that performs the grant, such as an S3 bucket policy or a database role, and then calls update_subscription_grant_status. Lineage can be added by sending OpenLineage run events from your jobs with post_lineage_event, so the catalog shows which job produced which table.

Worked example: sharing orders with marketing

A retailer creates one domain with two domain units, Sales and Marketing. Sales owns the sales_curated Glue database with an orders table partitioned by date and region, already registered in its S3 data lake and governed by Lake Formation.

The Sales data engineer creates the data source above. The run imports orders and orders_daily and skips orders_tmp because of the exclude filter. A Sales steward adds the business name Customer orders, the glossary term Net revenue on the amount column, and fills the domain's Classification form with Confidential. Then they publish.

A Marketing analyst searches for orders, finds the listing, reads the glossary definition, and requests access for the Marketing project with a reason. The Sales owner creates an asset filter limited to region = 'EU' and the non-PII columns, and approves with that filter. DataZone grants the Marketing project's data lake environment read access through Lake Formation. The analyst queries it in Athena in their own environment, sees only EU rows, and never needed a ticket.

Three weeks later Sales adds a column and edits the description. Until someone republishes, Marketing sees the old version in the catalog. That is the most common confusion in a new DataZone rollout.

Lineage, quality and search

A catalog earns trust when a consumer can answer two questions without asking anyone: where did this table come from, and is it any good today? DataZone has a hook for each.

For provenance, DataZone accepts lineage as OpenLineage run events. Only run events are supported, which fits how jobs already report: each event names the job, the run, its input datasets and its output datasets. Emit one event when a job starts and one when it completes, and DataZone links the job to the tables it read and wrote. Spark, Airflow and dbt all have OpenLineage integrations; a small adapter can forward their events to post_lineage_event, passing a clientToken so retries do not create duplicates. Dataset names in the events must resolve to the same tables your data sources imported, or the graph will show orphan nodes. Agree the naming convention before the first job reports.

For quality, a Glue data source can import AWS Glue Data Quality results alongside the schema; the configuration flag is autoImportDataQualityResult. Consumers then see recent rule outcomes on the asset page instead of discovering a broken feed in a dashboard.

Search is the third piece. search_listings accepts free text plus structured filters with EQ, LT, GT and TEXT_SEARCH operators combined with and/or, and glossary terms on assets and columns are searchable. That is why curated glossary terms matter more than long descriptions: they are what people actually filter by.

Operating a domain

  • Treat the domain as infrastructure. Keep domain units, metadata form types, glossaries, blueprint configurations and environment profiles in code, reviewed like any other change. Hand-built domains drift, and nobody can say why a policy exists.
  • Separate admin from approval. Domain admins configure; asset-owning projects approve. If the platform team approves every request, you have rebuilt the ticket queue with a nicer interface.
  • Make reasons useful. Require a request reason that names the use case and an end date, and review standing subscriptions on a schedule; revoke what is no longer used.
  • Watch the event stream. Route DataZone events through EventBridge to a queue you monitor, so stuck fulfilments, failed runs and new requests are visible rather than buried in the portal.
  • Keep the source of truth clear. Lake Formation and Redshift still hold the actual permissions. Reconcile them against DataZone's subscriptions periodically and investigate any grant that has no matching subscription.

Failure modes

  • Approved but cannot query. The consumer project has no environment matching the asset (a Redshift asset and only a data lake environment), or the subscription target's role lacks permission. Check the grant status on the subscription.
  • Grants that do not restrict. Lake Formation grants only control access when the tables are actually governed by Lake Formation. If broad IAM-based access to the S3 location remains, people can bypass the catalog. Audit the S3 bucket policies and IAM roles, not just the DataZone view.
  • Catalog full of raw tables. Publish on import was switched on for a broad filter. Turn it off, unpublish, and make curation a gate.
  • Stale listings. Inventory edits were never republished, or the data source run is failing; watch lastRunStatus for FAILED and PARTIALLY_SUCCEEDED.
  • Unmanaged requests stuck in progress. No automation consumes the EventBridge events, so nobody grants access or updates the status.
  • Ownership limbo. A project's only owner left. Use the domain unit's ownership-assumption policy to name a successor.

Trade-offs

OptionGives youGaps
Glue Data Catalog + Lake Formation aloneTechnical schemas and fine-grained grantsNo business context, no request workflow
DataZoneBusiness catalog, glossary, forms, request and approval, automatic grants for managed assetsGrants only Glue and Redshift natively; projects and environments to administer
SageMaker Unified StudioDataZone's catalog plus notebooks, SQL and model tools in one placeA larger platform to adopt
Third-party catalogsCross-cloud and SaaS coverageSeparate permission model to integrate with AWS grants

Technical metadata still comes from Glue, so Glue ETL jobs and crawlers remain part of the pipeline. Warehouse publishers should read Amazon Redshift for the data sharing that Redshift fulfilment relies on.

What to do next

  1. Pick the domain layout: one domain with domain units per business area, unless legal separation requires more.
  2. Associate the data-holding accounts and confirm their tables are governed by Lake Formation, not open IAM access.
  3. Define two or three metadata forms (classification, owner, retention) and a starter glossary before publishing anything.
  4. Create environment profiles from the Data Lake or Data Warehouse blueprint for each consumer pattern.
  5. Register one Glue data source with tight include and exclude filters and publish on import off; curate and publish by hand.
  6. Run one real subscription end to end, with an asset filter, and confirm the consumer sees exactly the permitted rows.
  7. Automate: alert on failed data source runs, and build the EventBridge-to-Lambda path for unmanaged assets.
Key takeaway: DataZone turns data sharing into a workflow: producers import, curate and publish assets, consumers find them and request access for a project, and approval creates real Lake Formation or Redshift grants. It works when tables are genuinely governed by Lake Formation, curation happens before publishing, and the unmanaged-asset path is automated. Start with one domain, a few metadata forms and one end-to-end subscription.