Time-series data looks like a perfect fit for HBase: append-heavy writes, reads that are range scans over sorted keys, and a storage engine that sorts everything for you. Naive designs fail predictably. One cell per point inflates storage several-fold because every cell repeats its full key. Time-first row keys send all writes to one region. Queries over weeks of data at full resolution scan millions of cells to draw a chart a few hundred pixels wide.

This article builds a metrics store on HBase from first principles. It covers the choice between series-first and time-first row keys, why closed time buckets should be packed into a single compressed cell, a working delta-of-delta and XOR encoder in the style of Facebook's Gorilla, the read path that merges packed and late-arriving points, routing between raw data and rollups, and a sizing exercise. Device-fleet modelling, latest-state tables and rollup pipelines are covered in HBase for IoT, and OpenTSDB's specific format in OpenTSDB in depth; this page is about the storage layout and read path you would build yourself.

A metrics store on HBase: raw points, packed hourly blocks, rollups and a routerCollectors100k points/sBufferedMutatorbatched Putstable metricsfamily draw points, TTL 2 daysfamily pone packed block per rowFIFO compactionlong retentionrow key: salt | metric | series | hourwritePackerevery 10 minutesread closed hourswrite blockRollup job1-minute, 1-hourmetrics_1m, metrics_1hcount, sum, min, maxSeries indextags to series idsQuery routerpicks a resolutionseries idscoarse rangesraw + packed
Raw points land in a short-lived family. A packer rewrites each closed hour as one compressed block in a second family, a rollup job maintains coarser tables, and a router decides which resolution answers each query.

The shape of time-series data

A time series is a sequence of (timestamp, value) points identified by a metric name and a set of tags, such as cpu.user{host=web-17,dc=fra}. Each distinct tag set is a series. Writes arrive roughly in time order across all series; reads ask for a few series, or an aggregate over many, across a time range; old data is read rarely and at lower resolution.

HBase stores cells sorted by row key, then family, qualifier and timestamp, and each cell carries its full coordinates. A cell holding an 8-byte double with a 17-byte row key and a 2-byte qualifier needs roughly 48 bytes in the KeyValue format before data block encoding and compression.

The shape of time-series data

A time series is a sequence of (timestamp, value) points identified by a metric name and a set of tags, such as cpu.user{host=web-17,dc=fra}. Each distinct tag set is a series. Writes arrive roughly in time order across all series; reads ask for a few series, or an aggregate over many, across a time range; old data is read rarely and at lower resolution.

HBase stores cells sorted by row key, then family, qualifier and timestamp, and each cell carries its full coordinates. A cell holding an 8-byte double with a 17-byte row key and a 2-byte qualifier needs roughly 48 bytes in the KeyValue format before data block encoding and compression.

Row keys: series-first or time-first

Use bucketed rows: one row per series per time bucket, with points as columns within the row. An hour is a common bucket; at a 10-second interval it holds 360 points. The open question is the order of the key parts.

LayoutRow keyCheap queryExpensive query
Series-firstsalt | metric | series | hourOne series over a long range: one contiguous scanAll series of a metric in one hour: one seek per series
Time-firstsalt | metric | hour | seriesAll series of a metric in a window: a contiguous scan per saltOne series over a month: one seek per hourly bucket

OpenTSDB uses the time-first order with tags after the base time, which suits dashboards that aggregate many series. A store whose dominant query is "these 50 series for the last 30 days" should go series-first. Either way, the metric must not be the leading component on its own, or every write for a busy metric lands on one region; a salt byte derived from the series spreads the load. Derive the salt from the series alone, not from the time, so a series' rows stay in one salt range and one series scan stays contiguous. The rules for choosing and pre-splitting salts are in HBase salting.

This article uses series-first: a 1-byte salt (hash of series id modulo 16), a 4-byte metric id, an 8-byte series id taken from a hash of the sorted tag set, and a 4-byte bucket start in epoch seconds. A series index table maps tag values to series ids. With 8-byte hashes, a collision among one million series has a probability of about 3 in 100 million; a registry that assigns ids sequentially removes it.

Writing raw points

Writes go to family d, one cell per point, qualifier = 2-byte offset of the point within the hour, value = 8-byte double. Use a BufferedMutator so Puts are batched per region server:

static final byte[] D = Bytes.toBytes("d");

byte[] rowKey(int metricId, long seriesId, long bucketStart) {
    byte salt = (byte) Math.floorMod(Long.hashCode(seriesId), 16);
    return ByteBuffer.allocate(17).put(salt).putInt(metricId)
            .putLong(seriesId).putInt((int) bucketStart).array();  // epoch s, valid to 2038
}

void write(BufferedMutator mutator, int metricId, long seriesId,
           long tsSeconds, double value) throws IOException {
    long bucket = tsSeconds - Math.floorMod(tsSeconds, 3600);
    Put put = new Put(rowKey(metricId, seriesId, bucket));
    put.addColumn(D, Bytes.toBytes((short) (tsSeconds - bucket)), Bytes.toBytes(value));
    mutator.mutate(put);   // flushed by buffer size or a periodic flush()
}

The cell timestamp is left to the server, so it records ingestion time, not event time. Event time lives in the row key and qualifier. That separation matters for TTL, which is measured against cell timestamps; the trap of setting cell timestamps to event time is covered in HBase TTL and versions.

Packing closed buckets

Once an hour is closed, plus an allowance for late points, its 360 cells are rewritten as a single cell in family p. The encoding follows the Gorilla paper (Pelkonen et al., VLDB 2015). Timestamps are stored as delta-of-deltas, which is zero for a perfectly regular interval and costs one bit. Each value is XORed with the previous one; slowly changing values share their high bits, so the XOR is mostly zeros and only its meaningful bits are written. The paper reports an average of 1.37 bytes per point on Facebook's monitoring data. Your data will differ, so measure it.

import struct

class BitWriter:
    def __init__(self):
        self.buf, self.acc, self.n = bytearray(), 0, 0
    def write(self, value, bits):
        for i in range(bits - 1, -1, -1):
            self.acc = (self.acc << 1) | ((value >> i) & 1)
            self.n += 1
            if self.n == 8:
                self.buf.append(self.acc)
                self.acc, self.n = 0, 0
    def getvalue(self):
        tail = bytes([self.acc << (8 - self.n)]) if self.n else b""
        return bytes(self.buf) + tail

def bits_of(x):
    return struct.unpack(">Q", struct.pack(">d", x))[0]

def encode_block(bucket_start, points):
    """points: sorted list of (ts_seconds, float) inside one bucket."""
    w = BitWriter()
    w.write(len(points), 16)
    t0, v0 = points[0]
    w.write(t0 - bucket_start, 14)        # offsets up to 16383 s
    w.write(bits_of(v0), 64)
    prev_t, prev_delta, prev_v = t0, 0, bits_of(v0)
    prev_lead, prev_trail = 65, 0         # no reusable window yet
    for t, v in points[1:]:
        delta = t - prev_t
        dod = delta - prev_delta
        if dod == 0:
            w.write(0, 1)
        elif -63 <= dod <= 64:
            w.write(0b10, 2); w.write(dod + 63, 7)
        elif -255 <= dod <= 256:
            w.write(0b110, 3); w.write(dod + 255, 9)
        elif -2047 <= dod <= 2048:
            w.write(0b1110, 4); w.write(dod + 2047, 12)
        else:
            w.write(0b1111, 4); w.write(dod & 0xFFFFFFFF, 32)
        prev_t, prev_delta = t, delta
        cur = bits_of(v)
        x = cur ^ prev_v
        if x == 0:
            w.write(0, 1)                 # value unchanged
        else:
            lead = min(64 - x.bit_length(), 31)
            trail = (x & -x).bit_length() - 1
            if lead >= prev_lead and trail >= prev_trail:
                w.write(0b10, 2)          # fits the previous window
                w.write(x >> prev_trail, 64 - prev_lead - prev_trail)
            else:
                sig = 64 - lead - trail
                w.write(0b11, 2); w.write(lead, 5); w.write(sig - 1, 6)
                w.write(x >> trail, sig)
                prev_lead, prev_trail = lead, trail
        prev_v = cur
    return w.getvalue()

The decoder mirrors each branch. Round-trip test it with NaN, negative zero and irregular gaps. The packer is idempotent: it reads every raw cell and any existing block for the row, merges them by timestamp with raw cells winning, encodes, and writes the block. Running it twice produces the same bytes, so a crashed or repeated run is harmless.

Table layout and compaction

The two families have different lifecycles, so they get different storage settings:

create 'metrics',
  {NAME => 'd', VERSIONS => 1, TTL => 172800,
   DATA_BLOCK_ENCODING => 'FAST_DIFF', COMPRESSION => 'SNAPPY',
   CONFIGURATION => {
     'hbase.hstore.defaultengine.compactionpolicy.class' =>
       'org.apache.hadoop.hbase.regionserver.compactions.FIFOCompactionPolicy',
     'hbase.hstore.blockingStoreFiles' => '1000'}},
  {NAME => 'p', VERSIONS => 1, COMPRESSION => 'NONE', BLOOMFILTER => 'ROW'},
  SPLITS => (1..15).map { |i| [i].pack('C') }

Family d holds two days of raw points. FIFO compaction, documented in the HBase reference guide, never rewrites data; it only deletes store files whose cells have all expired, which saves the compaction I/O that raw points would otherwise cost. It requires a non-default TTL and MIN_VERSIONS of 0, and the guide's example raises hbase.hstore.blockingStoreFiles because files accumulate between expiries. Family p skips compression because packed blocks are already dense. Its long retention is a good candidate for date-tiered compaction, whose settings are covered in HBase compaction tuning. SPLITS pre-splits the table on the 16 salt values.

The read path

A query names a metric, a tag filter, a time range and a step, for example the average of cpu.user for dc=fra over 7 days at 5-minute steps. The read path runs in four stages.

  1. Resolve the tag filter against the series index to a list of series ids.
  2. Choose the resolution. Seven days at 5-minute steps is 2,016 output points per series; raw data would decode 60,480 points per series to produce them. The router picks the coarsest table whose resolution is no coarser than the step, here the 1-minute rollup.
  3. Build one row range per series (the same shape for raw and rollup tables), from the bucket containing the start to the bucket after the end, and scan them together with a MultiRowRangeFilter, which seeks between ranges instead of reading the gaps.
  4. Decode, merge late raw cells over packed points, aggregate into steps, then combine across series.
List<MultiRowRangeFilter.RowRange> ranges = new ArrayList<>();
for (long sid : seriesIds) {
    ranges.add(new MultiRowRangeFilter.RowRange(
            rowKey(metricId, sid, startBucket), true,
            rowKey(metricId, sid, endBucket + 3600), false));
}
Scan scan = new Scan()
        .addFamily(Bytes.toBytes("p"))
        .addFamily(Bytes.toBytes("d"))
        .setFilter(new MultiRowRangeFilter(ranges))
        .setCaching(200);
try (ResultScanner rs = table.getScanner(scan)) {
    for (Result r : rs) {
        mergeIntoSteps(r, stepSeconds, accumulator);  // packed block + raw cells
    }
}

Rollup rows store count, sum, min and max rather than an average, because averages of averages are wrong when buckets hold different numbers of points, while sums and counts combine exactly. Percentiles do not combine; store a mergeable sketch per bucket if you need them. For wide fan-out, run one scan per salt in parallel; move aggregation into a coprocessor only once client-side aggregation is proven to be the bottleneck. Scanner settings are covered in HBase scans.

Late points and repacking

Points can arrive after their hour has been packed. They still land in family d, and readers merge them over the block, so queries are correct immediately. To fold them in, the packer uses cell timestamps, which record ingestion time: a scan of family d with setTimeRange(lastRun, now) finds exactly the rows touched since the last run, and HBase skips store files whose recorded time range does not overlap. Each touched row whose bucket is closed is repacked. Because the raw family's TTL counts from ingestion time, every late point gets the full two days to be packed, however old its event time. Large backfills of old data are better sent through a separate path that packs whole buckets directly, rather than repacking thousands of old rows one point at a time.

Worked example: sizing one million series

One million series at a 10-second interval: 100,000 points per second, 8.64 billion per day.

  • Raw family. At about 48 bytes per cell before encoding, that is about 415 GB per day of logical data; FAST_DIFF and Snappy reduce it several-fold, but measure it. Two days of TTL bound it.
  • Packed family. At an assumed 2 bytes per point, 17 GB per day, plus 24 million rows per day at about 40 bytes of cell overhead each, roughly 1 GB. A year is about 6.7 TB of logical data, about 20 TB on HDFS with three replicas.
  • Without packing. A year of raw cells is about 150 TB of logical data before encoding, more than twenty times the packed size.
  • Write load. 100,000 Puts per second across ten region servers is about 10,000 per server, a comfortable rate for batched writes. The packer rereads each closed hour once: 360 million points, about 17 GB of logical raw data per hour before encoding.

Failure modes

  • Hot region. A time-first key without a salt sends every write to the newest region. Check per-region request counts after go-live.
  • Packer falling behind. If packing lags past the raw TTL, points are lost. Alert on the age of the oldest unpacked closed bucket.
  • Encoder bugs. A decoder that misreads one branch corrupts every later point in the block. Keep round-trip tests in CI and a version byte at the start of each block.
  • Unbounded series. A tag holding request ids or user ids creates millions of series with a handful of points each, defeating bucketing. Enforce cardinality limits at ingest.

Trade-offs

ChoiceGainCost
Packed blocksOver twenty times less storage than unencoded raw cells; fewer cells to scanA packer job, an encoder to maintain, late-data handling
Series-first keysCheap long-range reads of few seriesWide aggregates seek once per series
FIFO on raw dataNo compaction I/O for short-lived pointsRequires TTL; files accumulate until expiry
Rollup tablesFast long-range queriesExtra writes and storage; fixed aggregate set
Building on HBase over a dedicated TSDBReuses an existing cluster and its operationsYou own the encoder, router and query layer

What to do next

  1. Write down your three most common queries and pick series-first or time-first keys from them.
  2. Define the row key with a series-derived salt and pre-split the table on salt values.
  3. Create separate raw and packed families, with a TTL and FIFO compaction on the raw family.
  4. Implement the encoder and decoder with round-trip tests, then measure bytes per point on a day of your own data.
  5. Build an idempotent packer driven by time-ranged scans, and alert on unpacked closed buckets.
  6. Add 1-minute and 1-hour rollups storing count, sum, min and max, and a router that picks resolution from the step.
  7. Load-test with production cardinality and verify write distribution across regions.
Key takeaway: Store time series in HBase as one row per series per time bucket, with a series-derived salt and a key order chosen from your dominant queries. Land raw points in a short-lived FIFO-compacted family, pack closed buckets into one delta-of-delta and XOR encoded cell, merge late points on read and repack them, keep rollups as sums and counts, and route each query to the coarsest resolution that still serves its step.