Your RDS dashboard shows DiskQueueDepth climbing from 2 to 60, queries are slow, and somebody suggests buying more IOPS. Sometimes that is right. Often it is not: a deep queue is a symptom, and the same number can mean the volume is out of IOPS, the instance has hit its EBS bandwidth limit, a gp2 burst bucket is empty, the throughput cap is binding on large reads, or the database simply stopped fitting in memory. Each cause has a different fix with a different price.
This article explains what the metric measures, how it relates to IOPS and latency through Little's law, which limits produce a queue, how to tell them apart from CloudWatch and engine statistics, and which fixes to try in which order. It applies to RDS for MySQL, MariaDB, PostgreSQL, Oracle, SQL Server and Db2, which all store data on EBS volumes. Aurora uses a different distributed storage layer; see the Aurora guide for that model.
What the metric measures
AWS defines DiskQueueDepth as the number of I/O requests in the queue waiting to be serviced: requests the database has submitted that have not yet been sent to the device because it is busy with other requests. RDS reports it as an average over each one-minute interval. Typical values range from zero to several hundred, and the unit is simply a count of outstanding I/Os.
Three consequences follow from the definition. First, it is an average, so a ten-second stall inside a quiet minute shows up as a modest number; look at maximum statistics and at latency alongside it. Second, the queue is not bad in itself, because storage needs several I/Os in flight to reach its rated IOPS. The RDS documentation even lists a queue depth consistently below 1 as a sign that the application is not driving enough I/O to use what it pays for. Third, time spent in the queue is part of the latency your queries feel, so a growing queue at constant IOPS always means growing latency.
Queue depth, IOPS and latency
Little's law links the three numbers you see in CloudWatch. For a stable system, the average number of requests in flight equals the arrival rate multiplied by the average time each request spends in the system. For storage that reads as:
queue_depth ~= IOPS x average_latency_seconds
3,000 IOPS x 0.002 s = 6 healthy: busy volume, fast completions
3,000 IOPS x 0.020 s = 60 same work, ten times the latency
900 IOPS x 0.045 s = 40 a gp2 volume pinned at baselineUse the law as a consistency check. Add ReadIOPS and WriteIOPS, take the IOPS-weighted average of ReadLatency and WriteLatency, which CloudWatch reports in seconds, and multiply. If the product is close to the reported queue depth, the metrics agree and the question becomes why latency rose. If IOPS is flat at a round number while the queue grows, something is capping IOPS. If IOPS is low, latency is high and the queue is moderate, look at throughput: a few large I/Os can saturate bandwidth while the IOPS count looks harmless.
The four limits that create a queue
An RDS I/O can be throttled at four places. Knowing their numbers turns a vague alarm into a specific diagnosis.
| Limit | Numbers that matter | CloudWatch signal |
|---|---|---|
| gp3 volume IOPS and throughput | 3,000 IOPS and 125 MiB/s below the striping threshold; 12,000 IOPS and 500 MiB/s at or above it, provisionable up to 64,000 IOPS and 4,000 MiB/s | TotalIOPS or throughput flat at the configured value |
| gp2 baseline and burst | 3 IOPS per GiB, minimum 100; volumes under 1,000 GiB burst to 3,000 IOPS while credits last | BurstBalance falling toward zero |
| io1 or io2 provisioned IOPS | What you provisioned, up to 256,000 IOPS on io2 Block Express | IOPS flat at the provisioned figure |
| Instance EBS limits | Per-class IOPS and bandwidth caps; smaller classes have burstable EBS performance | EBSIOBalance% or EBSByteBalance% falling; IOPS capped below the volume figure |
The threshold for gp3 striping is 400 GiB for Db2, MariaDB, MySQL and PostgreSQL and 200 GiB for Oracle; SQL Server does not stripe and lets you provision gp3 performance at any size. The instance limit wins whenever it is lower than the storage: the RDS documentation gives the example of a class capped at 40,000 IOPS attached to four 64,000-IOPS volumes, which still delivers 40,000, not 256,000. Check the EBS-optimised limits for your class in the EC2 documentation before provisioning storage performance it cannot use. The volume types themselves are compared in the EBS volume types guide.
Diagnosing a high queue
The script below pulls the relevant metrics for the last few hours, applies Little's law, and names the most likely limit. It uses get_metric_data so one call fetches everything.
import boto3, datetime as dt
cw = boto3.client("cloudwatch")
DB = "orders-prod"
METRICS = ["DiskQueueDepth", "ReadIOPS", "WriteIOPS", "ReadLatency",
"WriteLatency", "ReadThroughput", "WriteThroughput",
"BurstBalance", "EBSIOBalance%", "EBSByteBalance%", "FreeableMemory"]
end = dt.datetime.now(dt.timezone.utc)
queries = [{
"Id": f"m{i}",
"Label": name,
"MetricStat": {"Metric": {"Namespace": "AWS/RDS", "MetricName": name,
"Dimensions": [{"Name": "DBInstanceIdentifier", "Value": DB}]},
"Period": 60, "Stat": "Average"},
} for i, name in enumerate(METRICS)]
resp = cw.get_metric_data(MetricDataQueries=queries,
StartTime=end - dt.timedelta(hours=3), EndTime=end)
series = {r["Label"]: dict(zip(r["Timestamps"], r["Values"]))
for r in resp["MetricDataResults"]}
if not series.get("DiskQueueDepth"):
raise SystemExit(f"no DiskQueueDepth datapoints for {DB}: check the instance ID and Region")
worst = max(series["DiskQueueDepth"], key=series["DiskQueueDepth"].get)
v = {name: series.get(name, {}).get(worst) for name in METRICS}
iops = (v["ReadIOPS"] or 0) + (v["WriteIOPS"] or 0)
lat = ((v["ReadIOPS"] or 0) * (v["ReadLatency"] or 0) +
(v["WriteIOPS"] or 0) * (v["WriteLatency"] or 0)) / max(iops, 1)
mib_s = ((v["ReadThroughput"] or 0) + (v["WriteThroughput"] or 0)) / 2**20
print(f"{worst:%H:%M} queue={v['DiskQueueDepth']:.1f} iops={iops:.0f} "
f"lat={lat*1000:.1f}ms little={iops*lat:.1f} MiB/s={mib_s:.0f}")
if v["BurstBalance"] is not None and v["BurstBalance"] < 10:
print("gp2 burst credits exhausted: volume is at baseline IOPS")
elif any(v[k] is not None and v[k] < 10 for k in ("EBSIOBalance%", "EBSByteBalance%")):
print("instance EBS burst exhausted: the DB class is the limit")
elif mib_s > 0 and iops and mib_s * 1024 / iops > 64:
print("large I/Os: check throughput against the volume and instance caps")
else:
print("compare IOPS with provisioned and instance limits, then check cache hit ratio")Once storage is ruled in or out, look inside the engine. On PostgreSQL, pg_stat_statements shows shared_blks_read per statement, which identifies the queries causing physical reads, and wait events such as IO:DataFileRead show sessions waiting on them. On MySQL, compare Innodb_buffer_pool_reads with Innodb_buffer_pool_read_requests for the miss rate. RDS Performance Insights, whose capabilities AWS is folding into CloudWatch Database Insights, groups load by these wait events. Enhanced Monitoring adds operating-system disk statistics at up to one-second granularity, which catches bursts that one-minute averages flatten. CloudWatch in depth explains the metric statistics and alarm options used here.
Worked example: a gp2 burst cliff
A PostgreSQL instance holds 300 GiB on gp2. Its baseline is 900 IOPS, 3 per GiB, and it can burst to 3,000 while credits last. Every night at 01:00 a reporting job scans several large tables. For the first forty minutes everything is fine: IOPS sits near 3,000, latency near 2 ms, queue around 6. Then BurstBalance reaches zero, IOPS drops to a flat 900, latency rises to 45 ms and DiskQueueDepth climbs to about 40, exactly what Little's law predicts for 900 IOPS at 45 ms. The daytime API, which shares the volume, times out until the job finishes.
The obvious fix, more provisioned IOPS, is not available at this size: below 400 GiB, gp3 gives a fixed 3,000 IOPS and 125 MiB/s with nothing extra to provision. Converting to gp3 already helps a lot, because 3,000 becomes the sustained baseline rather than a burst. The change runs online, but changing storage type itself consumes I/O capacity and performance sits between the old and new specification while the volume is in the optimizing state, so never start it during the nightly job. Growing storage to 400 GiB would unlock 12,000 IOPS, but it also moves the instance from one volume to four, which RDS implements by copying data to new volumes. The documentation warns that this can consume a lot of IOPS, raise latency significantly and take several hours while the instance stays in Modifying, so schedule it as a planned change, not as an incident response.
The cheapest durable fix was elsewhere. Two of the reporting queries lacked an index on their date filter, and pg_stat_statements showed them responsible for most block reads. With the indexes added, the nightly read volume fell by roughly an order of magnitude in this case, and the job then also moved to a read replica so it no longer shared a volume with the API.
Fixes, cheapest first
Work through fixes from cheapest to most expensive, and stop as soon as latency is back inside your objective.
- Do less I/O. Add missing indexes, remove full scans, batch small writes into fewer commits, and fix chatty ORM patterns. Nothing else is as cheap.
- Cache more. If the working set has outgrown memory, a class with more RAM turns reads into buffer hits. Falling
FreeableMemoryand risingReadIOPStogether point here. - Move reads away. Send reporting and analytics to a read replica; its storage type is independent of the primary.
- Smooth writes. On PostgreSQL, larger
max_wal_sizevalues spread checkpoints out; on MySQL, review InnoDB flushing settings. Write bursts at checkpoints are a common cause of periodic queue spikes. - Raise the storage limit. Convert gp2 to gp3; above the striping threshold, provision IOPS and throughput; for latency-sensitive production, use io2 Block Express. A dedicated log volume, available only with Provisioned IOPS storage, separates transaction log writes from table I/O.
- Raise the instance limit. If
EBSIOBalance%orEBSByteBalance%drains, storage changes cannot help; move to a larger or newer class.
Alarms that mean something
Alarm on what users feel. A queue-depth alarm alone fires during healthy batch jobs and stays quiet when a slow volume serves a small queue. A better pair is a latency alarm with a threshold taken from your own baseline, plus early-warning alarms on the burst balances, which give you time to act before the cliff.
aws cloudwatch put-metric-alarm \
--alarm-name orders-prod-read-latency \
--namespace AWS/RDS --metric-name ReadLatency \
--dimensions Name=DBInstanceIdentifier,Value=orders-prod \
--statistic Average --period 60 \
--evaluation-periods 10 --datapoints-to-alarm 5 \
--threshold 0.010 --comparison-operator GreaterThanThreshold \
--alarm-actions arn:aws:sns:us-east-1:123456789012:db-oncall
aws cloudwatch put-metric-alarm \
--alarm-name orders-prod-burst-balance \
--namespace AWS/RDS --metric-name BurstBalance \
--dimensions Name=DBInstanceIdentifier,Value=orders-prod \
--statistic Minimum --period 300 --evaluation-periods 3 \
--threshold 20 --comparison-operator LessThanThreshold \
--alarm-actions arn:aws:sns:us-east-1:123456789012:db-oncallLatency metrics are in seconds, so 0.010 means 10 ms. Watch for two situations that produce misleading readings. After a restore from snapshot, blocks are fetched from S3 the first time they are read, so a freshly restored instance can show high read latency and queue depth until it warms up. And during a storage modification, performance sits between the old and new configuration, so do not benchmark until the operation completes.
Trade-offs
More storage performance is the fastest fix to apply and the most expensive to keep: provisioned IOPS and throughput are billed whether you use them or not. Query and index work is cheap to run but needs engineering time and a way to find the guilty statements. A larger instance class buys memory and EBS bandwidth together, which is efficient when both are short and wasteful when only one is. Read replicas add capacity but also replication lag, which some reads cannot tolerate. General background on instance classes, storage autoscaling and Multi-AZ is in the RDS operator's guide.
What to do next
- Run the diagnostic script against your busiest instance and record queue, IOPS, latency and the Little's law product at peak.
- Write down the four limits for that instance: volume IOPS, volume throughput, burst balance if gp2, and the class's EBS limits.
- Convert remaining gp2 volumes to gp3 during a quiet period.
- Find the top statements by physical reads with
pg_stat_statementsor Performance Insights and fix the worst one. - Replace queue-depth-only alarms with latency alarms plus burst-balance warnings.
- Schedule any change that crosses the striping threshold as planned maintenance.
- Move reporting and batch jobs off the primary to a read replica.