Spark, Hive and Trino find tables through a Hive metastore: a service that maps a name like sales.orders to a schema, a file format, a storage location and a list of partitions. On a self-managed Dataproc cluster that metastore lives on the cluster's master node, so deleting the cluster deletes every table definition with it. That one fact is why ephemeral clusters need an external metastore, and Dataproc Metastore is Google Cloud's managed version of it.

The second half of this topic, Data Catalog, has changed under it. Dataproc Metastore can sync table metadata to Data Catalog for search and discovery, but Data Catalog was discontinued on 1 June 2026 (the date was originally 30 January 2026), and its role has passed to Knowledge Catalog, the name Google has used since 10 April 2026 for what was Dataplex Universal Catalog. This article covers the metastore in depth, how engines share it, federation, and what to do with catalog integration now.

What a Hive metastore does

A metastore holds metadata, never data. For each table it records the database, columns and types, the input and output formats, the serialiser, a storage location such as gs://lake/sales/orders/, table properties, and for partitioned tables one row per partition with its own location. When a query runs, the engine asks the metastore for the table, asks again for the partitions that match the filter, and then reads the files itself from Cloud Storage.

Two consequences follow. First, the metastore is on the query path: every query start and every partition lookup is a metastore call, so a slow or overloaded metastore makes every engine slow at planning time, even though no data flows through it. Second, the metastore and the files can disagree. Writing files to a partition folder does not register the partition; dropping an external table does not delete the files. Most operational problems are a mismatch between the two.

The managed service and its fixed choices

Tables outlive clusters: one managed metastore, many engines, one catalogNightly ETL clusterSpark, ephemeralAd-hoc clusterSpark SQL, TrinoServerless Sparkbatch jobsDataproc MetastoreHive metastore APIThrift :9083 or gRPC :443backing DB: Cloud SQL or Spannerdatabases, tables, partitionsCloud StorageParquet / ORC filesreads and writes datametastore stores locations, not dataFederationgRPC; ranked backendsData Catalog synclegacy; service shut downKnowledge Catalogcatalogs metastore metadataData Catalog was discontinued on 1 June 2026 (Policy Tag Manager remains).Search and discovery now live in Knowledge Catalog, formerly Dataplex Universal Catalog.
Clusters come and go; table definitions live in the managed metastore, data lives in Cloud Storage, and discovery metadata flows to the catalog.

Dataproc Metastore runs the Hive metastore server for you, backed by a database you never manage directly. The choices you make at creation time are the ones that are hard to change later:

SettingOptionsWhat to know
GenerationDataproc Metastore 1 or 2Version 1 sizes by tier (developer or enterprise); version 2 scales horizontally by a scaling factor, with autoscaling
Endpoint protocolthrift (default) or grpcThrift uses port 9083; gRPC uses 443. A gRPC service cannot be switched back to Thrift; you would create a new one
Database typemysql (Cloud SQL, default) or spannerChosen at creation; there is no flag to change it afterwards
Hive metastore versionFor example 3.1.2Must be compatible with your engines; federation requires 3.1.2 or 2.3.6
Release channelstable or canaryCanary gets features earlier and is not for production
EncryptionGoogle-managed or a customer-managed keyCatalog sync cannot be combined with a customer-managed key

Pick gRPC if you will use federation or associate the metastore with a Knowledge Catalog lake, because both require it. Pick Thrift if you have clients outside Dataproc that only speak the classic Hive protocol and no need for federation. The backing database follows the same logic as any choice between Spanner and Cloud SQL: horizontal scale and regional availability against simplicity.

Creating the service and attaching engines

Create the service, attach clusters to it, and point any other engine at its endpoint. The commands below use flags documented for gcloud metastore services create; check the reference for your gcloud version before scripting them.

# A gRPC metastore, protected against accidental deletion.
gcloud metastore services create lake-hms \
  --location=us-central1 \
  --tier=enterprise \
  --hive-metastore-version=3.1.2 \
  --endpoint-protocol=grpc \
  --network=projects/my-proj/global/networks/data-vpc \
  --deletion-protection

# Every ephemeral cluster attaches to the same service.
gcloud dataproc clusters create etl-nightly \
  --region=us-central1 \
  --dataproc-metastore=projects/my-proj/locations/us-central1/services/lake-hms

# Back up metadata and service configuration before risky changes.
gcloud metastore services backups create pre-migration-0412 \
  --location=us-central1 --service=lake-hms

Clusters created with --dataproc-metastore are configured to use the service automatically. Self-managed Spark or Trino outside Dataproc set hive.metastore.uris to the Thrift endpoint URI of a Thrift service. A gRPC endpoint needs a client that speaks gRPC; for Dataproc clusters Google documents a proxy initialisation action for this. The service sits on a VPC network, so clients need a network path to it, which is a VPC design question to settle before the first cluster.

Tables and partitions done safely

With the metastore shared, a table created by one cluster is visible to every other. The pattern that keeps files and metadata consistent is to register partitions from the job that writes them, not by scanning storage later.

from pyspark.sql import SparkSession

spark = (SparkSession.builder
         .appName("orders-load")
         .enableHiveSupport()          # use the attached metastore
         .getOrCreate())

spark.sql("""
CREATE EXTERNAL TABLE IF NOT EXISTS sales.orders (
  order_id STRING, customer_id STRING, amount DECIMAL(12,2), status STRING)
PARTITIONED BY (dt DATE)
STORED AS PARQUET
LOCATION 'gs://lake/sales/orders/'
""")

def load_day(df, day):
    path = f"gs://lake/sales/orders/dt={day}/"
    df.write.mode("overwrite").parquet(path)                  # 1. files first
    spark.sql(f"ALTER TABLE sales.orders ADD IF NOT EXISTS "
              f"PARTITION (dt='{day}') LOCATION '{path}'")    # 2. then metadata

Writing files first and metadata second means a reader can never see a partition whose files are missing. The alternative, MSCK REPAIR TABLE, lists the whole table location and adds whatever it finds. That is fine for a one-off recovery but expensive as a daily habit on a large table, and it registers half-written folders if a job failed mid-write.

Worked example: three clusters, one metastore

A retail analytics team runs three kinds of compute: a nightly ETL cluster that lives for two hours, an ad-hoc cluster for analysts during working hours, and serverless Spark batches for feature engineering. Before the change each cluster had its own metastore, so the ETL job re-created table definitions on every run and analysts kept a script of CREATE TABLE statements that drifted from reality.

They create one gRPC metastore and attach all three. Table definitions now survive cluster deletion, and a schema change made by ETL is visible to analysts on their next query. The largest table is sales.orders with three years of daily partitions, about 1,100 partitions. That is small for a metastore. A clickstream table partitioned by day and hour across 50 sites, however, reaches 3 x 365 x 24 x 50 = 1.3 million partitions, and an analyst query with no partition filter asks the metastore for all of them. Planning takes minutes and the metastore's load spikes for every other user.

The team fixes it in the data model, not the metastore size: coarser partitions (day and site, with hour as a clustered column inside files) bring the count down to about 55,000, and the Hive setting hive.metastore.limit.partition.request, passed through the service's metastore configuration, turns accidental full fetches into fast errors. They then turn on table and column comments so analysts can find tables in Knowledge Catalog instead of asking in chat.

Federation

Federation gives clients one gRPC endpoint over several metadata sources. Backends are listed with a rank; when two backends contain a database with the same name, the lower rank wins. A Dataproc Metastore backend is written dpms: followed by the service's resource name, and BigQuery and Knowledge Catalog lakes (in Preview) are also supported backends.

gcloud metastore federations create lake-fed \
  --location=us-central1 \
  --hive-metastore-version=3.1.2 \
  --backends=1=dpms:projects/my-proj/locations/us-central1/services/lake-hms,2=dpms:projects/my-proj/locations/us-central1/services/finance-hms

Federation requires gRPC backend services, a federation version that the backends' Hive versions are at least as new as, and backends in the same region. Use it when separate teams own separate metastores and a few consumers need a combined view. Do not use rank order as a substitute for naming discipline: a database that silently shadows another is a debugging session waiting to happen.

From Data Catalog sync to Knowledge Catalog

A metastore answers "where is table X?" for engines. A catalog answers "which table holds customer orders, who owns it, and is it sensitive?" for people. For years the link between the two on Google Cloud was Dataproc Metastore's Data Catalog sync, the --data-catalog-sync flag. It copied databases (name and description) and tables (name, description and schema, including column descriptions) into Data Catalog, with a first full ingestion of up to six hours and no extra charge, and it could not be combined with a customer-managed encryption key.

That path is now legacy. Data Catalog was deprecated in early 2025 with a discontinuation date first set for 30 January 2026; Google's deprecation page now gives 1 June 2026, and that date has passed. Its search and metadata APIs are gone. One part survives: Policy Tag Manager, the taxonomies and policy tags used for column-level access control in BigQuery, is not deprecated and is still served under the datacatalog.googleapis.com endpoint, so do not rip out policy-tag code by mistake.

Knowledge Catalog, formerly Dataplex Universal Catalog, is where discovery lives now, and Google lists Dataproc Metastore among the systems whose metadata it catalogs, next to BigQuery and Cloud Storage. Check the Knowledge Catalog documentation for how your service's metadata reaches it, rather than assuming the old sync flag does the job. For a team working on this today, that means three things. Any automation that still searches or tags entries through Data Catalog APIs is already broken and needs porting, not planning. Ownership, sensitivity and glossary terms belong in Knowledge Catalog's model, not in Hive table properties that no one searches. And table and column comments in your CREATE TABLE statements matter more than any integration, because they are the descriptions every catalog ingests, whatever it is called next.

Operations and failure modes

SymptomCauseFix
Tables vanish when a cluster is deletedCluster used its local metastoreAttach every cluster with --dataproc-metastore; block cluster templates without it
Query planning takes minutesToo many partitions, or unfiltered queriesCoarser partitions; limit partitions fetched per query
Partition exists but query returns nothingMetadata registered before files were written, or files movedWrite files first; never move files under a registered partition
New data invisible to readersFiles written, partition never addedAdd partitions in the writing job; alert on folders without partitions
Cannot enable federationService uses ThriftCreate a gRPC service and migrate via export and import
Catalog jobs fail since June 2026They call the discontinued Data Catalog search or entry APIsPort them to Knowledge Catalog; keep Policy Tag Manager calls as they are
Clients cannot connectNo network path to the service's VPCFix routing and firewall rules; test from a client subnet

Access to the metastore is controlled with IAM roles on the service, so grant job service accounts the narrowest role that lets them read or write metadata, and remember that metastore access does not grant access to the files: the same accounts also need storage permissions (GCP IAM). Take a backup before upgrades or bulk schema changes, and export metadata to Cloud Storage on a schedule so a bad DROP DATABASE ... CASCADE is recoverable.

Trade-offs

A managed metastore versus a self-run Hive metastore on Cloud SQL: managed removes patching, scaling and availability work, and costs more per hour than a small self-run instance. Self-run gives full control of Hive versions and configuration. For more than one or two clusters, the managed service is usually cheaper in engineer time.

Hive tables in a metastore versus native BigQuery tables: Hive tables keep data in open formats that Spark, Trino and others read directly; BigQuery gives a serverless engine with its own storage optimisations. Many lakes keep raw and intermediate data in Hive tables and publish curated data to BigQuery. For lakes built on Apache Iceberg, also evaluate Google's BigLake metastore, which is designed around open table formats. For the clusters themselves, see Dataproc architecture.

What to do next

  1. List every cluster, serverless job and external engine that reads your lake tables, and which metastore each uses today.
  2. Decide on the endpoint protocol first: gRPC if federation or lake association is in your future, because it cannot be switched later.
  3. Create the service with deletion protection, and attach every cluster through --dataproc-metastore.
  4. Move partition registration into writing jobs: files first, then ADD PARTITION. Retire scheduled MSCK REPAIR.
  5. Count partitions per table and redesign any table heading toward millions.
  6. Schedule metadata backups and exports, and test one restore.
  7. Add table and column comments to every CREATE TABLE, and confirm the tables appear in Knowledge Catalog search.
  8. Search your code for Data Catalog search and entry API calls, which no longer work, and port them; leave policy-tag calls alone.
Key takeaway: Dataproc Metastore is a managed Hive metastore that lets table definitions outlive clusters and be shared by every engine. Choose the endpoint protocol and database type deliberately, because they are fixed at creation; register partitions from the writing job; keep partition counts bounded; back up metadata. For discovery, build on Knowledge Catalog: Data Catalog was discontinued on 1 June 2026, and only its Policy Tag Manager survives.