Spark, Hive and Trino find tables through a Hive metastore: a service that maps a name like sales.orders to a schema, a file format, a storage location and a list of partitions. On a self-managed Dataproc cluster that metastore lives on the cluster's master node, so deleting the cluster deletes every table definition with it. That one fact is why ephemeral clusters need an external metastore, and Dataproc Metastore is Google Cloud's managed version of it.
The second half of this topic, Data Catalog, has changed under it. Dataproc Metastore can sync table metadata to Data Catalog for search and discovery, but Data Catalog was discontinued on 1 June 2026 (the date was originally 30 January 2026), and its role has passed to Knowledge Catalog, the name Google has used since 10 April 2026 for what was Dataplex Universal Catalog. This article covers the metastore in depth, how engines share it, federation, and what to do with catalog integration now.
What a Hive metastore does
A metastore holds metadata, never data. For each table it records the database, columns and types, the input and output formats, the serialiser, a storage location such as gs://lake/sales/orders/, table properties, and for partitioned tables one row per partition with its own location. When a query runs, the engine asks the metastore for the table, asks again for the partitions that match the filter, and then reads the files itself from Cloud Storage.
Two consequences follow. First, the metastore is on the query path: every query start and every partition lookup is a metastore call, so a slow or overloaded metastore makes every engine slow at planning time, even though no data flows through it. Second, the metastore and the files can disagree. Writing files to a partition folder does not register the partition; dropping an external table does not delete the files. Most operational problems are a mismatch between the two.
The managed service and its fixed choices
Dataproc Metastore runs the Hive metastore server for you, backed by a database you never manage directly. The choices you make at creation time are the ones that are hard to change later:
| Setting | Options | What to know |
|---|---|---|
| Generation | Dataproc Metastore 1 or 2 | Version 1 sizes by tier (developer or enterprise); version 2 scales horizontally by a scaling factor, with autoscaling |
| Endpoint protocol | thrift (default) or grpc | Thrift uses port 9083; gRPC uses 443. A gRPC service cannot be switched back to Thrift; you would create a new one |
| Database type | mysql (Cloud SQL, default) or spanner | Chosen at creation; there is no flag to change it afterwards |
| Hive metastore version | For example 3.1.2 | Must be compatible with your engines; federation requires 3.1.2 or 2.3.6 |
| Release channel | stable or canary | Canary gets features earlier and is not for production |
| Encryption | Google-managed or a customer-managed key | Catalog sync cannot be combined with a customer-managed key |
Pick gRPC if you will use federation or associate the metastore with a Knowledge Catalog lake, because both require it. Pick Thrift if you have clients outside Dataproc that only speak the classic Hive protocol and no need for federation. The backing database follows the same logic as any choice between Spanner and Cloud SQL: horizontal scale and regional availability against simplicity.
Creating the service and attaching engines
Create the service, attach clusters to it, and point any other engine at its endpoint. The commands below use flags documented for gcloud metastore services create; check the reference for your gcloud version before scripting them.
# A gRPC metastore, protected against accidental deletion.
gcloud metastore services create lake-hms \
--location=us-central1 \
--tier=enterprise \
--hive-metastore-version=3.1.2 \
--endpoint-protocol=grpc \
--network=projects/my-proj/global/networks/data-vpc \
--deletion-protection
# Every ephemeral cluster attaches to the same service.
gcloud dataproc clusters create etl-nightly \
--region=us-central1 \
--dataproc-metastore=projects/my-proj/locations/us-central1/services/lake-hms
# Back up metadata and service configuration before risky changes.
gcloud metastore services backups create pre-migration-0412 \
--location=us-central1 --service=lake-hmsClusters created with --dataproc-metastore are configured to use the service automatically. Self-managed Spark or Trino outside Dataproc set hive.metastore.uris to the Thrift endpoint URI of a Thrift service. A gRPC endpoint needs a client that speaks gRPC; for Dataproc clusters Google documents a proxy initialisation action for this. The service sits on a VPC network, so clients need a network path to it, which is a VPC design question to settle before the first cluster.
Tables and partitions done safely
With the metastore shared, a table created by one cluster is visible to every other. The pattern that keeps files and metadata consistent is to register partitions from the job that writes them, not by scanning storage later.
from pyspark.sql import SparkSession
spark = (SparkSession.builder
.appName("orders-load")
.enableHiveSupport() # use the attached metastore
.getOrCreate())
spark.sql("""
CREATE EXTERNAL TABLE IF NOT EXISTS sales.orders (
order_id STRING, customer_id STRING, amount DECIMAL(12,2), status STRING)
PARTITIONED BY (dt DATE)
STORED AS PARQUET
LOCATION 'gs://lake/sales/orders/'
""")
def load_day(df, day):
path = f"gs://lake/sales/orders/dt={day}/"
df.write.mode("overwrite").parquet(path) # 1. files first
spark.sql(f"ALTER TABLE sales.orders ADD IF NOT EXISTS "
f"PARTITION (dt='{day}') LOCATION '{path}'") # 2. then metadataWriting files first and metadata second means a reader can never see a partition whose files are missing. The alternative, MSCK REPAIR TABLE, lists the whole table location and adds whatever it finds. That is fine for a one-off recovery but expensive as a daily habit on a large table, and it registers half-written folders if a job failed mid-write.
Worked example: three clusters, one metastore
A retail analytics team runs three kinds of compute: a nightly ETL cluster that lives for two hours, an ad-hoc cluster for analysts during working hours, and serverless Spark batches for feature engineering. Before the change each cluster had its own metastore, so the ETL job re-created table definitions on every run and analysts kept a script of CREATE TABLE statements that drifted from reality.
They create one gRPC metastore and attach all three. Table definitions now survive cluster deletion, and a schema change made by ETL is visible to analysts on their next query. The largest table is sales.orders with three years of daily partitions, about 1,100 partitions. That is small for a metastore. A clickstream table partitioned by day and hour across 50 sites, however, reaches 3 x 365 x 24 x 50 = 1.3 million partitions, and an analyst query with no partition filter asks the metastore for all of them. Planning takes minutes and the metastore's load spikes for every other user.
The team fixes it in the data model, not the metastore size: coarser partitions (day and site, with hour as a clustered column inside files) bring the count down to about 55,000, and the Hive setting hive.metastore.limit.partition.request, passed through the service's metastore configuration, turns accidental full fetches into fast errors. They then turn on table and column comments so analysts can find tables in Knowledge Catalog instead of asking in chat.
Federation
Federation gives clients one gRPC endpoint over several metadata sources. Backends are listed with a rank; when two backends contain a database with the same name, the lower rank wins. A Dataproc Metastore backend is written dpms: followed by the service's resource name, and BigQuery and Knowledge Catalog lakes (in Preview) are also supported backends.
gcloud metastore federations create lake-fed \
--location=us-central1 \
--hive-metastore-version=3.1.2 \
--backends=1=dpms:projects/my-proj/locations/us-central1/services/lake-hms,2=dpms:projects/my-proj/locations/us-central1/services/finance-hmsFederation requires gRPC backend services, a federation version that the backends' Hive versions are at least as new as, and backends in the same region. Use it when separate teams own separate metastores and a few consumers need a combined view. Do not use rank order as a substitute for naming discipline: a database that silently shadows another is a debugging session waiting to happen.
From Data Catalog sync to Knowledge Catalog
A metastore answers "where is table X?" for engines. A catalog answers "which table holds customer orders, who owns it, and is it sensitive?" for people. For years the link between the two on Google Cloud was Dataproc Metastore's Data Catalog sync, the --data-catalog-sync flag. It copied databases (name and description) and tables (name, description and schema, including column descriptions) into Data Catalog, with a first full ingestion of up to six hours and no extra charge, and it could not be combined with a customer-managed encryption key.
That path is now legacy. Data Catalog was deprecated in early 2025 with a discontinuation date first set for 30 January 2026; Google's deprecation page now gives 1 June 2026, and that date has passed. Its search and metadata APIs are gone. One part survives: Policy Tag Manager, the taxonomies and policy tags used for column-level access control in BigQuery, is not deprecated and is still served under the datacatalog.googleapis.com endpoint, so do not rip out policy-tag code by mistake.
Knowledge Catalog, formerly Dataplex Universal Catalog, is where discovery lives now, and Google lists Dataproc Metastore among the systems whose metadata it catalogs, next to BigQuery and Cloud Storage. Check the Knowledge Catalog documentation for how your service's metadata reaches it, rather than assuming the old sync flag does the job. For a team working on this today, that means three things. Any automation that still searches or tags entries through Data Catalog APIs is already broken and needs porting, not planning. Ownership, sensitivity and glossary terms belong in Knowledge Catalog's model, not in Hive table properties that no one searches. And table and column comments in your CREATE TABLE statements matter more than any integration, because they are the descriptions every catalog ingests, whatever it is called next.
Operations and failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Tables vanish when a cluster is deleted | Cluster used its local metastore | Attach every cluster with --dataproc-metastore; block cluster templates without it |
| Query planning takes minutes | Too many partitions, or unfiltered queries | Coarser partitions; limit partitions fetched per query |
| Partition exists but query returns nothing | Metadata registered before files were written, or files moved | Write files first; never move files under a registered partition |
| New data invisible to readers | Files written, partition never added | Add partitions in the writing job; alert on folders without partitions |
| Cannot enable federation | Service uses Thrift | Create a gRPC service and migrate via export and import |
| Catalog jobs fail since June 2026 | They call the discontinued Data Catalog search or entry APIs | Port them to Knowledge Catalog; keep Policy Tag Manager calls as they are |
| Clients cannot connect | No network path to the service's VPC | Fix routing and firewall rules; test from a client subnet |
Access to the metastore is controlled with IAM roles on the service, so grant job service accounts the narrowest role that lets them read or write metadata, and remember that metastore access does not grant access to the files: the same accounts also need storage permissions (GCP IAM). Take a backup before upgrades or bulk schema changes, and export metadata to Cloud Storage on a schedule so a bad DROP DATABASE ... CASCADE is recoverable.
Trade-offs
A managed metastore versus a self-run Hive metastore on Cloud SQL: managed removes patching, scaling and availability work, and costs more per hour than a small self-run instance. Self-run gives full control of Hive versions and configuration. For more than one or two clusters, the managed service is usually cheaper in engineer time.
Hive tables in a metastore versus native BigQuery tables: Hive tables keep data in open formats that Spark, Trino and others read directly; BigQuery gives a serverless engine with its own storage optimisations. Many lakes keep raw and intermediate data in Hive tables and publish curated data to BigQuery. For lakes built on Apache Iceberg, also evaluate Google's BigLake metastore, which is designed around open table formats. For the clusters themselves, see Dataproc architecture.
What to do next
- List every cluster, serverless job and external engine that reads your lake tables, and which metastore each uses today.
- Decide on the endpoint protocol first: gRPC if federation or lake association is in your future, because it cannot be switched later.
- Create the service with deletion protection, and attach every cluster through
--dataproc-metastore. - Move partition registration into writing jobs: files first, then
ADD PARTITION. Retire scheduledMSCK REPAIR. - Count partitions per table and redesign any table heading toward millions.
- Schedule metadata backups and exports, and test one restore.
- Add table and column comments to every
CREATE TABLE, and confirm the tables appear in Knowledge Catalog search. - Search your code for Data Catalog search and entry API calls, which no longer work, and port them; leave policy-tag calls alone.