Every low-latency audio problem is a buffering problem. Sound is a continuous stream, and computers process it in chunks. Each chunk has to be ready before the hardware needs it, so software keeps some audio queued ahead. Queue too little and the device runs dry: an underrun, heard as a click or dropout. Queue too much and every sound arrives late: a guitarist hears their note after they play it, two people on a call talk over each other. Buffer management is the craft of choosing how much to queue, where, and how to react when the choice turns out wrong.
This article covers buffering on the device side and inside your application: the latency arithmetic, the callback model, the main platform settings, how to hand audio between threads without locks, and how to adapt buffer sizes at run time. Network buffering for packets that arrive with jitter is a separate problem with its own design, covered in the adaptive jitter buffer article; the two stack, and their latencies add.
Latency is frames divided by sample rate
Audio hardware consumes samples at a fixed rate. At 48 kHz, one frame (one sample per channel) lasts about 20.8 microseconds, so a block of N frames lasts N divided by 48,000 seconds. Every buffer is a delay of that size whenever it is full. That one formula lets you reason about any audio stack without trusting marketing figures.
| Frames per period | At 44.1 kHz | At 48 kHz | Typical use |
|---|---|---|---|
| 32 | 0.73 ms | 0.67 ms | Specialist hardware, rarely stable on general systems |
| 64 | 1.45 ms | 1.33 ms | Live monitoring, instruments, tuned systems |
| 128 | 2.90 ms | 2.67 ms | Low-latency default for music apps |
| 256 | 5.80 ms | 5.33 ms | Voice chat, games |
| 480 | 10.9 ms | 10.0 ms | Common shared-mode engine period |
| 1024 | 23.2 ms | 21.3 ms | Playback where latency barely matters |
Round-trip latency, from microphone to speaker, is the sum of every stage that holds frames: converter filters, the input buffer, queues between threads, block-based processing, the output buffer and driver safety offsets. Measure it end to end with a loopback cable and a click track; an API's reported figure is usually only part of it.
The callback model and its deadline
Most modern audio APIs are pull-based. The driver wakes a high-priority thread once per period and calls your function to fill or drain one period of frames. That function, the audio callback, has a hard deadline: it must return before the hardware finishes playing what is already queued. With a 128-frame period at 48 kHz the deadline is 2.67 ms, every time, forever. Being late once is an audible glitch.
The output buffer is usually a small number of periods. With two periods, the hardware plays one while your callback fills the other, so output latency from the buffer alone is up to two periods. Adding periods buys tolerance for scheduling hiccups at the cost of latency. The figure shows where frames wait in a typical full-duplex app.
Platform knobs
Each platform exposes the same two numbers, period size and number of periods, under different names. Request a size, then read back what you actually got, because drivers round requests to what the hardware supports.
| Platform | Period setting | Queue depth setting | Notes |
|---|---|---|---|
| ALSA (Linux) | snd_pcm_hw_params_set_period_size_near | snd_pcm_hw_params_set_buffer_size_near | Buffer size is periods times period size |
| JACK | Frames per period | Periods per buffer | One graph period for every client |
| PipeWire | Quantum | Set by the graph | Clients can request a latency such as PIPEWIRE_LATENCY=256/48000 |
| Core Audio (macOS) | kAudioDevicePropertyBufferFrameSize | Managed by the HAL | Device and stream safety offsets add to it |
| WASAPI (Windows) | Engine period; smaller via IAudioClient3 | Buffer duration at initialize | Shared mode defaults to a 10 ms engine period |
| AAudio (Android) | Burst, from AAudioStream_getFramesPerBurst | AAudioStream_setBufferSizeInFrames | Request the low-latency performance mode |
On Windows, shared mode mixes all applications in the audio engine at a fixed period, traditionally 10 ms. IAudioClient3 lets an application query the periods the driver supports with GetSharedModeEnginePeriod and open a stream at a smaller one with InitializeSharedAudioStream, when the driver advertises smaller periods. Exclusive mode bypasses the engine entirely and gives the lowest latency, at the cost of locking other applications out of the device. On Android, a buffer of two bursts is the usual starting point, adjusted at run time as shown later.
Rules for the real-time thread
The callback runs on a thread the operating system schedules with real-time or elevated priority. That priority only helps if the callback never waits for anything with lower priority. Anything that can block for an unbounded time is forbidden inside it:
- No locks shared with non-real-time threads. If a background thread holds the mutex and gets descheduled, the callback waits for it. This is priority inversion, and it causes the rare, unreproducible glitches that take weeks to find.
- No memory allocation or freeing. General-purpose allocators take locks and can touch the operating system. Allocate everything before the stream starts.
- No file, network or console I/O, including logging. Count events with atomics and let another thread report them.
- No unbounded loops. A DSP chain that costs 60 percent of the period on average and 140 percent on a cache-cold path will glitch.
- No garbage-collector pauses; keep the callback in native code or a runtime that never pauses that thread.
Budget against the worst case, not the average. Processing that takes 1 ms at the 99.99th percentile inside a 1.33 ms period leaves too little headroom on a general-purpose laptop.
Lock-free rings between threads
Real applications have work that cannot run in the callback: decoding a compressed stream, running a speech model, receiving network packets. The standard pattern puts that work on an ordinary thread and connects it to the callback through a single-producer single-consumer ring buffer. Each side owns one index; each reads the other's index with acquire ordering and publishes its own with release ordering. No locks, no allocation, and both operations take bounded time.
#include <atomic>
#include <algorithm>
#include <cstddef>
// Single producer, single consumer. N must be a power of two.
template <std::size_t N>
class SpscRing {
static_assert((N & (N - 1)) == 0, "N must be a power of two");
float buf_[N];
alignas(64) std::atomic<std::size_t> head_{0}; // written only by the producer
alignas(64) std::atomic<std::size_t> tail_{0}; // written only by the consumer
public:
std::size_t write(const float* src, std::size_t n) { // producer thread
std::size_t h = head_.load(std::memory_order_relaxed);
std::size_t t = tail_.load(std::memory_order_acquire);
n = std::min(n, N - (h - t)); // free space
for (std::size_t i = 0; i < n; ++i) buf_[(h + i) & (N - 1)] = src[i];
head_.store(h + n, std::memory_order_release); // publish the frames
return n;
}
std::size_t read(float* dst, std::size_t n) { // consumer thread
std::size_t t = tail_.load(std::memory_order_relaxed);
std::size_t h = head_.load(std::memory_order_acquire);
n = std::min(n, h - t); // frames available
for (std::size_t i = 0; i < n; ++i) dst[i] = buf_[(t + i) & (N - 1)];
tail_.store(t + n, std::memory_order_release); // hand the space back
return n;
}
std::size_t fill() const {
return head_.load(std::memory_order_acquire) - tail_.load(std::memory_order_acquire);
}
};The power-of-two capacity makes wraparound a mask, the ever-increasing indexes keep full and empty unambiguous, and the alignment keeps the indexes on separate cache lines. Exactly one writer and one reader; mix multiple sources on the worker side first.
Watermarks: the ring fill is your latency
A ring between worker and callback adds latency equal to its fill level, not its capacity. Capacity is just the ceiling. The worker decides the fill by producing up to a target and then waiting. Choose the target from the worker's worst-case wake-up delay plus its worst-case block cost; two periods is a reasonable start for a local decoder, more if the worker itself waits on a network.
SpscRing<8192> playout; // about 170 ms of mono audio at 48 kHz
std::atomic<uint32_t> underruns{0};
// Called by the audio API on its real-time thread, once per period.
void on_output(float* out, std::size_t frames) {
std::size_t got = playout.read(out, frames);
if (got < frames) { // the worker fell behind: we still owe the device audio
std::fill(out + got, out + frames, 0.0f);
underruns.fetch_add(1, std::memory_order_relaxed); // logged by another thread, never here
}
}
// Worker thread: keep the ring between a low and a high watermark.
void worker_loop() {
const std::size_t target = 2 * PERIOD; // the latency you choose to pay
while (running) {
while (playout.fill() < target) {
std::size_t n = render_block(block, PERIOD); // decode, synthesize, mix
playout.write(block, n);
}
wait_for_wakeup(); // condition variable signalled by a timer, not the callback
}
}Two details matter. On underrun the callback writes silence and counts the event; it never waits for data. And the worker is woken by its own timer or by a condition variable that only non-real-time threads touch. Signalling a condition variable from the callback would mean taking its lock on the real-time thread.
Adapting at run time
The right buffer size depends on the device, the driver, the thermal state and whatever else is running, so good applications measure and adapt. The usual policy is asymmetric: grow quickly when underruns happen, shrink slowly and only after a long clean stretch. Growing is audible once; shrinking too eagerly causes repeated glitches. On Android, AAudio exposes the counters needed to do this directly:
// Android AAudio: start at two bursts, grow by one burst when the device reports xruns.
int32_t burst = AAudioStream_getFramesPerBurst(stream);
AAudioStream_setBufferSizeInFrames(stream, 2 * burst);
int32_t last_xruns = 0;
// On a normal thread, every few hundred milliseconds:
int32_t xruns = AAudioStream_getXRunCount(stream);
if (xruns > last_xruns) {
int32_t size = AAudioStream_getBufferSizeInFrames(stream);
AAudioStream_setBufferSizeInFrames(stream, size + burst); // clamped to capacity by the API
last_xruns = xruns;
}Use the same policy for your own rings: raise the worker's target by a period after an underrun, lower it by a period after a clean minute, cap growth, and expose the current size in diagnostics.
Two clocks, one stream
When input and output are different devices, such as a USB microphone and Bluetooth headphones, each runs from its own crystal. Two nominal 48 kHz clocks typically differ by tens of parts per million, so a ring between them fills or drains by a few frames per second. With a ring holding 256 frames, the stream glitches every few minutes no matter how carefully you sized it. The fix is to measure the fill trend and resample by the tiny ratio needed to hold it steady, which is the subject of clock drift compensation and depends on a good resampler. Echo cancellers suffer especially from drift, because they assume a fixed delay between the playback reference and the microphone.
Worked example: a 10 ms round trip for live voice effects
Target: a voice-effects app at 48 kHz with no more than 10 ms from microphone to headphones. Start with the budget. Converters cost about 1 ms total. That leaves 9 ms, or 432 frames, for everything software holds.
Choose 128-frame periods with two-period input and output buffers. Input waits up to one period before the callback sees it: 2.67 ms. Processing in the same callback adds no queue, as long as it finishes in time. Output holds up to two periods: 5.33 ms. Total software holding is 8 ms, so measured round trip lands near 9 ms. The processing budget is now fixed: the effect chain must finish well inside 2.67 ms per call at its worst, which rules out a neural model in the callback.
A neural denoiser on a worker thread needs its own ring, adding at least one model block plus wake-up delay. With a 10 ms block the round trip reaches about 20 ms, over target. The options are a smaller block, a smaller hop or a looser target, and the arithmetic tells you that before you write code.
Failure modes
- Rare clicks under load. A lock or allocation in the callback, exposed when another thread holds the lock. Audit the callback for any blocking call.
- Latency that grows over a session. Clock drift filling a ring, or a growth policy that never shrinks.
- Glitches only on battery or after a few minutes. CPU frequency scaling and thermal throttling stretch processing time; budget for the slow state.
- Requested size ignored. The driver rounded your request. Always read back the period and buffer actually granted.
- Bluetooth output adds codec and radio buffering far larger than any period you set; do not promise live monitoring over it.
Trade-offs
Smaller periods mean lower latency, more CPU overhead and less tolerance for scheduling jitter. More periods mean fewer glitches and more latency. Exclusive mode lowers latency and blocks other apps. Work in the callback avoids a ring but must be cheap and bounded; work on a worker can be anything and costs at least one block. The right setting is the smallest buffer that survives your worst realistic load, plus a run-time policy for when reality is worse. For the whole call pipeline, see real-time audio architecture.
What to do next
- Write your latency budget as frames: converters, input, queues, output. Check it against the target before choosing any algorithm.
- Measure round-trip latency with a loopback cable and compare it with the budget.
- Read back the period and buffer size the driver actually granted and log them at stream start.
- Audit the callback for locks, allocation, I/O and unbounded loops; replace shared state with atomics and SPSC rings.
- Time the callback at the 99.99th percentile under load, on battery, and after the device warms up.
- Add an underrun counter and an asymmetric grow-fast, shrink-slow buffer policy, with a cap.
- If input and output are different devices, add drift compensation before shipping.