Cloudera Data Warehouse (CDW) is Cloudera's way of running Hive, Impala and Trino as a containerised service instead of as daemons on a long-lived Hadoop cluster. If you already know the engines, the interesting questions are what changes when they run this way: where the metadata lives, how compute is sized and scaled, what auto-suspend does to caches and latency, and how you split workloads so one team's ETL job does not stall another team's dashboards.

This article explains CDW from first principles, describes its three layers, walks through sizing and auto-scaling with the numbers Cloudera documents, works an example for a BI team and an ETL team sharing data, and ends with failure modes and a checklist. Product details were checked against Cloudera's documentation in October 2026; CDW changes often, so confirm specifics against the docs for your release.

Architecture at a glance

ClientsJDBC/ODBC, BI tools, shellsCDW control planecreate, size, suspend, upgrade warehousesKubernetes in your environment (EKS or AKS on public cloud)Impala Virtual Warehousecoordinators + executor groupsHive Virtual WarehouseHiveServer + executorsTrino Virtual Warehousecoordinator + workersDatabase CatalogHive metastore: tables, views, permissionsObject storage in the data lakeS3 or ADLS: ORC, Parquet, Iceberg filesSQLmanagescan
The control plane manages Virtual Warehouses that run as containers in Kubernetes; each warehouse runs one engine against a shared Database Catalog, and all data lives in object storage.

From first principles: separating storage, metadata and compute

On a classic cluster, storage, metadata and compute share the same machines and the same lifecycle. Adding capacity for one team means adding nodes for everyone, an upgrade affects every workload at once, and idle capacity is still paid for at 3 a.m. Separating these concerns is the idea behind every cloud warehouse, and CDW applies it to the Cloudera engines in three steps.

  1. Data lives in object storage as open files, so it outlives any compute and any engine can read it.
  2. Metadata, meaning table definitions, partitions, views and permissions, lives in a shared metastore, so every engine sees the same tables.
  3. Compute is a disposable set of containers sized for one workload, started when needed and stopped when idle.

The trade is that compute no longer sits next to its data. Every scan reads from object storage over the network, so local caches on the executors become essential to performance, and anything that empties them, such as suspending or resizing, has a cost you must plan for.

The three layers and choosing an engine

Cloudera's documentation describes CDW as three decoupled layers. The control plane is the management application: you activate CDW on an environment you have registered, then create, size, suspend, upgrade and delete warehouses through it. The Database Catalog is a logical collection of table and view metadata and security permissions, backed by a Hive metastore that points at data in the environment's object storage. The Virtual Warehouse is an instance of compute, running in Kubernetes (EKS or AKS on public cloud), that executes queries against one Database Catalog and exposes JDBC/ODBC and HiveServer-compatible endpoints.

When you create a Virtual Warehouse you choose its engine, and that choice should follow the workload.

EngineStrengthsTypical use
ImpalaLow-latency MPP execution, no per-query start-up, high concurrency of short queriesInteractive analysis and BI dashboards
HiveFull SQL and ACID semantics, long complex queries, query isolation for big scansETL, complex reports, enterprise dashboards
TrinoConnectors and federation across heterogeneous sourcesQueries joining lake tables with external databases

Several warehouses can attach to the same Database Catalog, which is how CDW gives teams isolated compute over shared data: the dashboards and the nightly ETL see the same tables but never compete for the same executors.

Sizing: executors and executor groups

Sizing is by executors, the unit of compute. Cloudera's sizing page for CDW on cloud lists these T-shirt sizes, with a custom option for anything else:

SizeExecutorsWhen to use
XSMALL2Evaluation and learning; the documented recommendation for trying CDW
SMALL10Modest concurrency or data volumes
MEDIUM20Typical departmental workloads
LARGE40Large scans or high concurrency
Custom1 to 100Matching an existing on-premises cluster; required for Trino auto-scaling

The size is also the size of one executor group, the block in which auto-scaling adds and removes capacity. A MEDIUM warehouse scales in groups of 20 executors. The practical way to choose is to start from the executor count of the cluster that runs the workload today, then adjust for query complexity and data scanned per query, measured from your own query history rather than guessed.

Auto-scaling and auto-suspend

CDW scales differently per engine, and the difference matters for planning.

  • Impala: concurrency auto-scaling. When queries queue because the current executor groups are busy, the auto-scaler adds another executor group, up to the maximum you set, and removes groups when they fall idle. Coordinators plan queries and hand fragments to executors. Cloudera's private-cloud documentation describes two coordinators by default for availability and about three large queries per executor group; treat that as a sizing hint, not a contract. Cloudera also documents Workload Aware Auto-Scaling for Impala, which scales across executor group sets of different sizes according to the workload.
  • Hive: concurrency plus query isolation. Hive scales for concurrency in the same way for BI-style queries. For ETL-style queries it can also run a big query in isolation on a dedicated executor group. The planner estimates the bytes a query will read; if isolation is enabled and the estimate exceeds hive.query.isolation.scan.size.threshold, it starts a separate group sized for that query, capped by hive.query.isolation.max.nodes.per.query, which the docs give as twice the T-shirt size by default.
  • Trino: auto-scaling is available only for warehouses created with a custom size.
  • Auto-suspend. The AutoSuspend Timeout is how long a warehouse may idle before it suspends and stops consuming compute. When a query arrives at a suspended warehouse, the auto-scaler sees the queued query and starts an executor group; the query waits for it. Disabling auto-suspend keeps the last group running when idle.

Worked example: dashboards and ETL on shared tables

A retail company has two workloads on the same lake tables. The BI team runs dashboards for about 60 analysts from 08:00 to 18:00, mostly short aggregations with bursts at 09:00 and after lunch. The data engineering team runs a nightly ETL from 01:00 to 04:00 that includes one weekly job scanning about 6 TB. Today both share a 30-node cluster that runs all day.

Design: one Database Catalog for the shared tables, an Impala warehouse for BI and a Hive warehouse for ETL. The BI warehouse is SMALL (10 executors) with a maximum of three executor groups, so bursts can reach 30 executors, and an auto-suspend timeout long enough to cover lunch-hour lulls so caches stay warm during the working day. The ETL warehouse is MEDIUM (20 executors) with a short auto-suspend timeout, since nobody waits on its first query, and with query isolation on. For the isolation threshold, Cloudera's own example multiplies the core group's executors by the per-executor data cache size: 20 x 200 GB = 4 TB. Under that example the weekly 6 TB job runs isolated on up to 40 executors without evicting the nightly jobs' cached data.

Compute, in executor-hours per weekday, is simple arithmetic: BI uses 10 executors for 10 hours plus 20 burst executors for about 2 hours, 140 in total; ETL uses 20 for 3 hours, 60 in total, plus about 40 on the weekly job's day. That is about 200 to 240 executor-hours against 720 for 30 always-on nodes. Executors and nodes are not priced the same, so convert with your own rates before claiming savings, but the shape is the point: the capacity follows the workload.

-- Impala Virtual Warehouse: an Iceberg table partitioned by day, then statistics
CREATE TABLE sales.orders (
  order_id BIGINT, customer_id BIGINT, amount DECIMAL(12,2), ts TIMESTAMP)
PARTITIONED BY SPEC (day(ts))
STORED AS ICEBERG;

COMPUTE STATS sales.orders;          -- the planner and admission control need row counts

-- Hive Virtual Warehouse: the nightly rollup reading the same table
INSERT OVERWRITE TABLE sales.daily_revenue
SELECT to_date(ts) AS d, SUM(amount) FROM sales.orders
WHERE ts >= date_sub(current_date, 1) GROUP BY to_date(ts);

Connecting is the same as for any HiveServer2 endpoint. Copy the exact JDBC URL from the warehouse rather than assembling one by hand.

# Hive: copy the JDBC URL from the Virtual Warehouse's options menu in the CDW UI
beeline -u '<jdbc-url-copied-from-the-ui>' -n "$WORKLOAD_USER" -p

# Impala: the shell speaks HiveServer2 over HTTP on port 443
pip install impala-shell
impala-shell --protocol='hs2-http' --ssl -i '<impala-endpoint-host>:443' -u "$WORKLOAD_USER" -l

Operating CDW

  • Watch queueing, not just CPU. Queued queries are the signal that drives Impala and Hive scaling. Persistent queueing at the maximum group count means the maximum is too low or the queries are too heavy; tune queries first.
  • Keep statistics current. Admission control and join planning depend on them; run COMPUTE STATS or ANALYZE TABLE after large loads.
  • Mind the small files. Object-store listing and per-file overhead dominate when ETL writes many tiny files. Compact, or size output files deliberately.
  • Metadata freshness across engines. When Hive writes a table Impala reads, Impala must learn about it, either from metastore event processing or an explicit REFRESH. Check which applies in your release.
  • Plan upgrades per warehouse. Each warehouse runs its own engine version, so upgrade one, run its queries, then do the rest.
  • Tag every warehouse. Creation accepts optional key-value tags; use them for team and cost centre so executor-hours can be charged back to the workload that spent them.

A useful habit is to read scaling as a latency budget. A query that arrives when every group is busy pays the time to start a new group on top of its own run time, and the new group starts with an empty cache, so its first scans run at object-storage speed. For a dashboard with a two-second target that penalty is visible to users; for a nightly job it is noise. That is why the same scaling settings are rarely right for both kinds of warehouse, and why the first thing to check when interactive users report slow mornings is whether the warehouse suspended overnight and how long the first query waited in the queue. Query profiles from the engine show queueing time separately from execution time; look there before resizing anything.

Failure modes

  • Cold start after suspend. The first morning query waits for executors to start and then reads from object storage with empty caches. Pre-warm before business hours or lengthen the timeout for interactive warehouses.
  • Runaway scale-out. A badly written dashboard query multiplied by many users triggers new executor groups. Set the maximum deliberately and use admission control limits.
  • Isolation threshold mis-set. Too low and many ordinary queries start their own groups and pay start-up latency; too high and big scans evict everyone's cache.
  • One catalog, one blast radius. All warehouses share a metastore; a bad DDL change or a burst of partition operations affects every engine using it.
  • Auto-suspend disabled by default habit. Warehouses left running overnight are the most common cost surprise.

Trade-offs

Against a classic CDP Base cluster, CDW gives isolation per workload, elastic capacity and independent upgrades, and costs you data locality, cold-start latency after suspend, and a Kubernetes platform to understand. Against fully managed cloud warehouses, it keeps data in open formats in your own storage, readable by several engines with one permission model, at the price of more knobs to tune. Smaller warehouses with auto-scaling waste less money; larger always-on warehouses give steadier latency.

Related reading

For the engines under CDW: Hive LLAP, Impala admission control and the Impala data cache. For the shared metadata layer read the Hive metastore, for the table format Impala with Iceberg, and for other vendors cloud-native analytics compared.

What to do next

  1. Inventory your workloads by latency need, concurrency and data scanned per query, and assign each to Impala, Hive or Trino.
  2. Start each warehouse at the size closest to the cluster that runs it today, with a maximum group count you can afford.
  3. Set auto-suspend per warehouse: short for batch, longer for interactive, and plan a pre-warm for morning peaks.
  4. Enable Hive query isolation for ETL and derive the threshold from your executor count and cache size.
  5. Migrate one dashboard and one ETL job first, compare latency and executor-hours against the old cluster, then move the rest.
  6. Alert on queued queries at maximum scale, warehouses that never suspend, and stale table statistics.
Key takeaway: CDW separates data in object storage, metadata in a Hive-metastore-backed Database Catalog, and compute in Virtual Warehouses that run Hive, Impala or Trino as containers. Size warehouses in executors, from 2 for XSMALL to 40 for LARGE or a custom count, knowing that the size is also the auto-scaling step. Impala adds executor groups when queries queue; Hive does the same and can also isolate very large scans on a dedicated group. Auto-suspend saves money but empties caches, so tune it per workload. Give each team its own warehouse over a shared catalog, and measure queueing, statistics and executor-hours.