Compression in Hive is not one setting. A single query can touch three different codecs: the one inside the columnar files it reads and writes, the one applied to shuffle data between stages, and, for row-oriented formats, the one wrapped around whole output files. Each has different defaults, different configuration keys and a different trade-off between CPU, disk and network. Most surprises, such as a 40 GB table read by a single mapper, a cluster that suddenly cannot read last night's partition, or a codec change that did nothing, come from mixing these up.

A useful mental model: a codec is a dial between CPU and bytes. Every byte you save on disk is a byte you do not read from storage, push across the network or hold in a cache, but every byte must be compressed once by a writer and decompressed by every reader, every time a query touches it. Data written once and scanned thousands of times rewards spending more CPU at write time for faster, smaller reads; data shuffled once between two stages rewards the cheapest codec that still shrinks it.

This article separates the three, explains how each codec behaves and why splittability matters, and walks a realistic conversion of a gzip text landing zone into compressed ORC. File layout details belong to the ORC format and Parquet format articles; here we assume you know a file has stripes or row groups and focus on the bytes the codec sees. Property names and values were checked against the ORC, Hive and Impala documentation; where a default depends on your version, the article says so rather than guessing.

Advertisement

The three places compression happens

Three places a Hive query compresses dataSource tableORC / Parquet / textMap / Tez vertexscan, filter, projectShufflesort, partition, spillReduce vertexaggregate, joinTarget tablefiles written to storagedecompress1. Storage codecorc.compress, parquet.compression2. Intermediate codechive.exec.compress.intermediatetez.runtime.compress(.codec)3. Output codeccolumnar: table property of targettext / SequenceFile: hive.exec.compress.outputEach knob is independent: a fast codec for shuffle and a dense codec for storage is the usual pairing.
Storage, intermediate and output compression are configured separately. The storage codec of a columnar target comes from that table's properties; the shuffle codec comes from session or engine settings.

Storage compression lives inside ORC and Parquet files. The writer encodes each column (dictionary, run-length, delta) and then compresses the encoded bytes in independent chunks. Because every chunk is independently decodable and the footer records offsets, a reader can still split the file and skip stripes or row groups. This is the codec that decides your storage bill and most of your scan cost.

Intermediate compression applies to data that one stage writes for the next: map output being sorted and shuffled, or spills to local disk. It lives for minutes and never reaches the warehouse, so speed matters far more than ratio.

Output compression in the classic sense applies to row-oriented outputs, text and SequenceFile, where Hadoop wraps the stream in a codec. For columnar targets the table's own properties win; the generic output switch matters mostly for text exports and legacy tables.

Codecs compared

CodecTypical characterSplittable as a text fileGood for
SnappyVery fast, moderate ratioNoHot data, shuffle, interactive queries
LZ4Very fast decompression, moderate ratioNoShuffle, hot data
ZLIB / DEFLATE / gzipGood ratio, slower, CPU heavy to writegzip: noCold or archive data inside ORC
ZSTDRatio close to or better than zlib at much higher speed; tunable levelNoDefault choice for new columnar tables if every reader supports it
bzip2High ratio, very slowYesRare: raw text archives that must stay splittable
LZOFast; GPL licensed, shipped separatelyOnly with an index fileLegacy text pipelines

Treat these as directions, not numbers: the ratio depends heavily on your data after columnar encoding, which has already removed much of the redundancy. Low-cardinality columns that ORC dictionary-encodes compress well with anything; high-entropy columns such as UUIDs or hashes barely compress at all. Always measure on a real partition before standardising.

Advertisement

Splittability, and the 40 GB single mapper

A file is splittable if a reader can start decoding from the middle. Plain gzip is a single stream; the only way to read byte 30,000,000 is to decompress everything before it. So when Hive plans a scan over a 40 GB .gz text file, it must give the whole file to one task. Your 200-node cluster processes it on one core. Raw Snappy-framed text has the same problem.

Splittability is a property of the container, not just the codec. ORC, Parquet and block-compressed SequenceFiles compress in independent chunks with sync points or footers, so they split regardless of codec. That is why the practical rule is: keep data in raw text only at the edge, convert to a columnar format immediately, and let the columnar format carry the codec. If you must keep compressed text, keep individual files small, around one HDFS block or object-store part, so lack of splitting does not matter.

ORC: setting and tuning the codec

ORC reads its codec from table properties. Valid values for orc.compress are NONE, ZLIB, SNAPPY, LZO, LZ4 and ZSTD. The default depends on the ORC library bundled with your engine: it was ZLIB for many years, and current ORC documentation lists ZSTD. Do not rely on either; set it explicitly so every writer, whether Hive, Spark or a streaming job, produces the same files.

CREATE TABLE sales.orders_orc (
  order_id     BIGINT,
  customer_id  BIGINT,
  status       STRING,
  amount       DECIMAL(12,2),
  created_at   TIMESTAMP
)
PARTITIONED BY (dt STRING)
STORED AS ORC
TBLPROPERTIES (
  'orc.compress'      = 'ZSTD',
  'orc.compress.size' = '262144'   -- compression chunk size in bytes (the documented default)
);

orc.compress.size is the size of each independently compressed chunk. Larger chunks give the codec more context and slightly better ratios; smaller ones reduce the bytes a reader must decompress to reach one value and the memory each open column stream needs. The default is a sensible balance; change it only after measuring, and remember that wide tables multiply per-column buffers by the number of columns being written.

If any reader in your estate is old, ZSTD is the risky choice: an engine whose ORC library predates ZSTD support cannot read those files. Inventory readers (Hive, Impala, Spark, Presto or Trino, external tools) before switching. In that situation SNAPPY or ZLIB are the safe fallbacks.

Parquet: Hive and Impala settings

Hive reads the Parquet codec from the parquet.compression table property, supported since Hive 1.1.0, or from the same key set in the session or hive-site.xml. Typical values are UNCOMPRESSED, SNAPPY, GZIP and, with newer Parquet libraries, ZSTD. If neither is set, what you get depends on your distribution, so set it.

Impala does not use the table property for its own writes. It uses the COMPRESSION_CODEC query option, whose allowed values are NONE, SNAPPY (the default), GZIP, ZSTD and LZ4. ZSTD takes an optional level from 1 to 22, defaulting to 3:

-- Hive
CREATE TABLE sales.events_pq (...) STORED AS PARQUET
TBLPROPERTIES ('parquet.compression' = 'SNAPPY');

-- Impala: codec chosen per session, applied to every Parquet file this INSERT writes
SET COMPRESSION_CODEC=zstd:6;
INSERT INTO sales.events_pq PARTITION (dt='2026-10-01')
SELECT ... FROM staging.events_raw WHERE dt='2026-10-01';

The split means one table can legitimately contain Snappy files written by Impala and ZSTD files written by Hive. That is fine for readers, because Parquet records the codec per column chunk, but it confuses anyone estimating storage. Pick one codec per table and configure every writer to use it.

Text, SequenceFile and intermediate settings

For row-oriented outputs, such as an export to a text table for a downstream system, enable output compression and choose a codec class:

SET hive.exec.compress.output=true;
SET mapreduce.output.fileoutputformat.compress.codec=org.apache.hadoop.io.compress.GzipCodec;
-- For SequenceFile targets, compress blocks of records, not each record:
SET mapreduce.output.fileoutputformat.compress.type=BLOCK;

-- Shuffle and spill data between stages
SET hive.exec.compress.intermediate=true;
-- On Tez, the shuffle codec is governed by the Tez runtime keys:
SET tez.runtime.compress=true;
SET tez.runtime.compress.codec=org.apache.hadoop.io.compress.SnappyCodec;

Intermediate compression usually pays for itself on joins and aggregations that shuffle tens of gigabytes, because network and local disk are slower than Snappy or LZ4. It can cost a little on small queries that are CPU bound. Early Tez releases named these tez.runtime.intermediate-output.*; check what your distribution's tez-site.xml already sets before overriding it. How shuffle fits into the Tez DAG is covered in Hive on Tez execution.

Worked example: converting a gzip landing zone to ORC

A team lands 2 TB per day of gzip-compressed CSV, each file 30 to 60 GB. Queries over the raw table run with one task per file and spend most of their time decompressing. The fix is a daily conversion job that rewrites each partition into the ORC table above:

SET hive.exec.dynamic.partition.mode=nonstrict;
SET hive.exec.compress.intermediate=true;

INSERT OVERWRITE TABLE sales.orders_orc PARTITION (dt)
SELECT CAST(order_id AS BIGINT),
       CAST(customer_id AS BIGINT),
       status,
       CAST(amount AS DECIMAL(12,2)),
       CAST(created_at AS TIMESTAMP),
       dt
FROM   landing.orders_csv_gz
WHERE  dt = '2026-10-01';

The conversion itself is slow, because each big gzip file is still read by one task. Fixing that at the source, by asking the producer for many smaller files or for ORC or Parquet directly, removes the bottleneck. Once in ORC, downstream scans split freely, read only the referenced columns and skip stripes using statistics. Measure before and after with the same query: bytes read, task count, CPU seconds and wall time. Also check the file sizes produced, since an over-partitioned conversion can trade one problem for the small files problem.

Changing a table's codec safely

Running ALTER TABLE t SET TBLPROPERTIES ('orc.compress'='ZSTD') changes only what future writes produce. Existing files keep the codec they were written with, because the codec is recorded in each file, and readers handle a mix without trouble. To convert existing data, rewrite it: INSERT OVERWRITE partition by partition for non-transactional tables, or rely on a major compaction to rewrite base files for ACID tables once the property is set.

Then verify instead of assuming. For ORC, hive --orcfiledump <path> prints the compression kind and chunk size from the file footer; for Parquet, the parquet-cli meta command shows the codec per column chunk. Sample a few files from old and new partitions after every change.

Failure modes

FailureSymptomFix
Huge gzip text filesOne task per file, long tailsConvert to ORC or Parquet; keep raw files small
Reader lacks the codecOlder engine errors on newly written partitionsInventory readers before switching to ZSTD or LZ4; fall back to SNAPPY or ZLIB
Native library missingCodec class errors on some nodes onlyInstall native Snappy, ZSTD or LZO libraries cluster-wide; LZO needs separate GPL packages
Codec change seems ignoredStorage unchanged after ALTEROnly new files change; rewrite or compact
Writers disagreeSame table, different codecs per partitionSet the table property and every engine's session option
High level ZSTD on hot pathSlow inserts, CPU-bound writersUse a low level for ingest and recompress cold partitions later

Trade-offs and a sensible default policy

Compression trades CPU for bytes, and in Hive bytes usually cost more: they are read from disk or object storage, sent over the network and held in caches. A reasonable starting policy is ZSTD at a low level for columnar tables if all readers support it, otherwise Snappy for hot tables and ZLIB for cold ones; Snappy or LZ4 for shuffle; and gzip only for text exports leaving the cluster. For cost-sensitive Impala deployments, pair this with the advice in Impala cost optimisation.

Tiering gets you both. Write recent partitions with a fast codec so ingest keeps up and interactive queries decompress cheaply, then have a scheduled job rewrite partitions older than, say, thirty days with a denser codec or a higher ZSTD level. Because the codec is recorded per file, the table needs no schema change and readers never notice; only the storage graph does. Keep the rewrite partition-scoped and idempotent so a failed run can simply be repeated.

What to do next

  1. List every table's format and its explicit codec property; flag tables relying on defaults.
  2. Find compressed text files larger than one block and plan their conversion to a columnar format.
  3. Inventory every engine that reads your tables and confirm which codecs each supports before adopting ZSTD.
  4. Enable intermediate compression with Snappy or LZ4 and compare shuffle-heavy queries before and after.
  5. Set one codec per table and configure Hive's table property and Impala's COMPRESSION_CODEC to match.
  6. After any change, inspect sample files with orcfiledump or parquet-cli and rewrite old partitions deliberately.
Key takeaway: Hive compresses in three independent places: inside ORC and Parquet files, on shuffle data between stages, and around row-oriented output files. Columnar formats stay splittable whatever the codec, so keep data columnar, choose ZSTD or Snappy for storage based on what every reader supports, use a fast codec for shuffle, set codecs explicitly rather than trusting version-dependent defaults, and remember that a codec change applies only to files written afterwards.