Redpanda stores each partition as a Raft log on local disk, which makes it fast and also makes retention expensive: keeping thirty days of a busy topic means thirty days on every replica's NVMe. Tiered Storage breaks that link. Closed log segments are uploaded to an object store such as Amazon S3, local disk keeps only the recent part of the log, and consumers that ask for older offsets are served from the bucket through a local cache. Producers and consumers use the ordinary Kafka API throughout.
This article explains how Redpanda's implementation behaves, so you can size it, configure it and know what will hurt. The general Kafka design, KIP-405 with its remote log manager and plugins, is covered in the Kafka tiered storage article; here the details are Redpanda's, checked against its documentation in October 2026. Tiered Storage requires a Redpanda enterprise license.
From local log to two tiers
A Redpanda partition is a sequence of segment files. The newest, active segment receives appends; when it reaches the topic's segment.bytes it is closed and a new one starts. Every replica holds the same segments, replicated by Raft. Without Tiered Storage, the only way to drop data is to delete the oldest local segments when retention says so, so retention is bounded by the smallest disk.
With Tiered Storage, the partition leader uploads closed segments to the bucket along with a partition manifest that lists which offset ranges live in which objects. The log then has two tiers: a local tail on every replica, and a complete history in object storage held once. Local retention can be short because nothing is lost when old local segments are removed; the bucket still has them.
The upload path, precisely
Three rules decide when data reaches the bucket. Redpanda only uploads segments that contain offsets smaller than the last stable offset, so uncommitted transactional data is never archived. A segment is uploaded when it closes at segment.bytes, or when it has been idle for cloud_storage_segment_max_upload_interval_sec, which forces low-traffic partitions to upload too; if that property is null, metadata is still uploaded but a segment waits until it is full. And only the leader uploads, so the bucket holds one copy, not one per replica.
The consequence is that object storage lags the local log, by up to one segment or one upload interval per partition. That lag is your recovery point if the whole cluster is lost and you restore from the bucket, so choose the interval with that in mind: a 10-minute interval on a quiet partition means up to 10 minutes of data that exists only on the replicas' local disks.
As partitions grow, the manifest grows too. Redpanda archives older manifest metadata into spillover manifests in the bucket, controlled by cloud_storage_spillover_manifest_size, so very long retention does not mean an ever-growing manifest in memory.
Two retention layers
Tiered topics have two independent retention settings, and confusing them is the most common misconfiguration.
| Setting | Applies to | Default for new topics | Meaning |
|---|---|---|---|
retention.local.target.ms / .bytes | Local disk on each replica | One day | How much recent log to keep locally; a target, not a hard cap |
retention.ms / retention.bytes | The whole log, including the bucket | Seven days | When data is deleted for good |
retention_local_target_capacity_percent / _bytes | Cluster-wide local disk | Version dependent | Overall target for log data on a node |
retention_local_strict | Cluster | Version dependent | Whether housekeeping trims local data back to the configured retention |
The local values are targets because Redpanda manages local disk as a whole: under space pressure it can trim older local data that is already in the bucket, and it keeps actively used data and the next sequential segments for readers locally. A setting of retention.ms=-1 with tiered storage means infinite retention in the bucket, which is legitimate for an event store but should be a decision, not an accident.
The read path: chunks and the cache
When a fetch asks for an offset that is no longer on local disk, the broker consults the manifest, finds the object, and downloads the needed part into the Tiered Storage cache directory, cloud_storage_cache_directory. Reads are chunked: rather than whole segments, Redpanda downloads chunks of cloud_storage_cache_chunk_size, 16 MiB by default, so a consumer reading a few records from a large segment does not pull the whole object. The cache is capped by cloud_storage_cache_size or cloud_storage_cache_size_percent and evicts old chunks when full.
This shapes performance. The first read of a cold range pays object-store latency, typically tens of milliseconds per request rather than the sub-millisecond of local NVMe, but sequential consumers amortize it because each chunk serves many fetches. Large fetch sizes help. Many concurrent backfills over different ranges are the bad case: their working set exceeds the cache, chunks are evicted before reuse, and the same objects are downloaded repeatedly.
import time
from confluent_kafka import Consumer, TopicPartition
c = Consumer({
"bootstrap.servers": "redpanda-0:9092",
"group.id": "orders-reprocess-2026-10",
"enable.auto.commit": False,
"fetch.max.bytes": 16 * 1024 * 1024, # large fetches suit remote reads
"max.partition.fetch.bytes": 8 * 1024 * 1024,
})
# Start each partition at a timestamp that is only in object storage
ts_ms = int((time.time() - 20 * 86400) * 1000) # 20 days ago: past local, inside total
parts = [TopicPartition("orders", p, ts_ms) for p in range(12)]
c.assign(c.offsets_for_times(parts, timeout=10))
while True:
msgs = c.consume(num_messages=500, timeout=5.0)
if not msgs:
break
handle(msgs) # idempotent sink
c.commit(asynchronous=False)This backfill reads a 20-day-old range from every partition. It is ordinary Kafka client code: Tiered Storage is invisible to clients except in latency. Run one large backfill at a time, size the cache for its working set, and make the sink idempotent so a restart from the last committed offset is safe.
Configuring it
# Cluster: point Redpanda at a bucket (an enterprise license is required)
rpk cluster config set cloud_storage_enabled true
rpk cluster config set cloud_storage_bucket my-redpanda-tiered
rpk cluster config set cloud_storage_region eu-west-1
rpk cluster config set cloud_storage_credentials_source aws_instance_metadata
# Topic, v26.1 and later: one switch
rpk topic create orders -p 12 -r 3 \
-c redpanda.storage.mode=tiered \
-c retention.local.target.ms=86400000 \
-c retention.ms=2592000000
# Topic on older clusters, or when the storage mode is unset: the legacy pair
rpk topic alter-config orders \
--set redpanda.remote.write=true \
--set redpanda.remote.read=true
# Check what the topic actually ended up with
rpk topic describe orders -cFrom v26.1 the recommended per-topic switch is redpanda.storage.mode=tiered; the older redpanda.remote.write and redpanda.remote.read properties still control topics whose storage mode is unset, and the cluster properties cloud_storage_enable_remote_write and cloud_storage_enable_remote_read only set creation-time defaults for such topics. A topic that has write without read uploads data but cannot serve it back, which is occasionally useful for archiving and usually a mistake. redpanda.remote.delete decides whether deleting the topic deletes its objects; leave it on unless something else owns the bucket's lifecycle. Supported stores are Amazon S3, Google Cloud Storage, Azure Blob Storage and Azure Data Lake Storage.
Worked example: sizing a 30-day topic
An orders stream ingests 50 MB/s, replication factor 3, and must be replayable for 30 days. Here is the arithmetic for keeping one day locally.
def size(ingress_mb_s, rf, local_days, total_days, segment_mib):
day_tb = ingress_mb_s * 86_400 / 1e6 # logical TB per day
without_ts = day_tb * total_days * rf # every byte on every replica
with_ts_local = day_tb * local_days * rf
with_ts_bucket = day_tb * total_days # one copy, uploaded by the leader
puts_per_day = day_tb * 1e12 / (segment_mib * 2**20)
return dict(
without_ts_disk_tb=round(without_ts, 1),
with_ts_disk_tb=round(with_ts_local, 1),
bucket_tb=round(with_ts_bucket, 1),
segment_puts_per_day=round(puts_per_day),
)
print(size(ingress_mb_s=50, rf=3, local_days=1, total_days=30, segment_mib=256))
# {'without_ts_disk_tb': 388.8, 'with_ts_disk_tb': 13.0,
# 'bucket_tb': 129.6, 'segment_puts_per_day': 16093}Without Tiered Storage, 30 days at RF3 needs about 389 TB of NVMe across the cluster before headroom. With it, local disk needs about 13 TB plus cache and headroom, while the bucket holds about 130 TB once, at object-storage prices. With 256 MiB segments, the upload rate is around 16,000 segment PUTs a day plus manifest uploads, which is small; it becomes large only if you shrink segments or set a very short upload interval across thousands of partitions. Reads are the variable cost: a full 30-day replay downloads the 130 TB again, so budget request and egress charges for each planned backfill, and keep the brokers in the same region as the bucket.
Recovery, read replicas and compaction
Because the bucket holds a complete, manifested copy of each partition, it doubles as a recovery source. Creating a topic with redpanda.remote.recovery=true rebuilds it from the objects in the bucket, which is how you restore a deleted topic or move one to a new cluster. The recovered topic contains what was uploaded, so anything inside the upload lag is gone.
Remote Read Replicas use the same data for read scaling: a separate cluster with access to the same bucket creates a topic with redpanda.remote.readreplica=<bucket>, and serves a read-only mirror. It trails the origin by at least the upload lag, so it suits analytics and reprocessing, not low-latency consumers, and it keeps that load off the production brokers entirely.
Compacted topics are supported. With the default implementation, tiered_v1, segments are compacted locally and re-uploaded. Setting default_redpanda_storage_mode_tiered_impl=tiered_v2 enables a beta implementation where compaction runs on the data in object storage; treat it as beta. The semantics of compaction itself are in the log compaction article.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Local disk keeps filling | Bucket unreachable or credentials expired, so segments cannot be uploaded and stay local | Alert on upload lag; test credential rotation |
| Old offsets return nothing | Remote read disabled, or total retention shorter than expected | Describe the topic; check retention.ms, not only local |
| Backfill very slow, broker CPU busy | Cache smaller than concurrent backfill working set | Serialize backfills, grow cache, larger fetches |
| Surprising object-store bill | Tiny segments or short upload interval on many partitions; cross-region egress | Larger segments, same-region bucket |
| Data loss after restore | Restored only what had been uploaded | Shorten upload interval for critical topics |
| Topic deleted, bucket still full or emptied unexpectedly | redpanda.remote.delete not what you assumed | Set it explicitly per topic |
Monitor upload backlog, cache hit rate and remote request errors from the broker metrics, using the names in your version's metrics reference, and alert when the bucket falls more than a few upload intervals behind. Tail consumers do not notice an outage of the object store; disks filling up an hour later do.
Trade-offs
Tiered Storage trades disk cost for read latency on old data, an operational dependency on the object store, and a licensed feature. It makes brokers lighter, because decommissioning or rebalancing moves only the local tail, and it turns the bucket into a backup and a read-scaling source. It does not make tail latency better, does not replace replication, since recent data is still protected only by Raft, and does not suit workloads that constantly read random old offsets. If you are comparing it with Kafka's or Pulsar's tiering, the platform comparison covers the storage models side by side.
What to do next
- Write down each topic's replay requirement and set total retention from it, then set local retention to the window your tail and lagging consumers actually need, often hours to a day.
- Enable Tiered Storage at cluster level with instance credentials, then per topic with
redpanda.storage.mode=tieredon v26.1 or the legacy write and read properties on older versions, and confirm withrpk topic describe -c. - Pick the upload interval per topic from its acceptable recovery point and check the PUT cost it implies across all partitions.
- Size the cache for your largest planned backfill and run backfills one at a time with large fetches and idempotent sinks.
- Alert on upload backlog and remote errors, and rehearse a topic restore with remote recovery into a scratch cluster.
- Review how Raft replication protects the local tail by reading the replication article, then decide whether a read replica cluster should take analytics load.