A real-time audio program has one job that dominates every other design decision: the audio device asks for the next block of samples at a fixed rhythm, and the program must hand them over before the deadline every single time. Miss once and the listener hears a click, a dropout or a stutter. Average speed does not matter; the worst case does.
Rust is a good fit because it has no garbage collector to pause the audio thread and its type system catches data races at compile time. It is not a complete fit: nothing in the type system stops you allocating, locking or doing I/O inside the callback, and those are exactly the things that cause glitches. This article shows how to structure a Rust audio pipeline so the callback stays deterministic, using the cpal crate for device I/O and the rtrb crate for lock-free ring buffers, and how to measure whether it works. For the signal-processing stages themselves, see real-time audio architecture.
The deadline arithmetic
The device consumes audio in blocks of a fixed number of frames (one frame is one sample per channel). The time you have to produce a block is the block size divided by the sample rate. At 48 kHz:
| Block size (frames) | Deadline per callback | Typical use |
|---|---|---|
| 64 | 1.33 ms | Live instruments, monitoring |
| 128 | 2.67 ms | Low-latency effects, games |
| 256 | 5.33 ms | Voice calls, most interactive apps |
| 512 | 10.67 ms | Playback, streaming, heavy DSP |
| 1024 | 21.33 ms | Background playback |
The deadline is not your budget. The operating system, the driver and other processes share the same core, so a common rule is to keep the callback's worst-case time well under half the deadline. Total latency is larger still: it adds the device's own buffering, any ring buffer between your producer and the callback, and the hardware converters. A block size is therefore a trade between latency and safety margin, and the right value is the largest one your use case tolerates.
Three thread roles
A robust pipeline separates work by how predictable it is. The audio callback runs on a thread created by the audio backend and must only do bounded, predictable work. A worker thread does everything that can stall: decoding files, receiving network packets, running a neural model, resampling large chunks. A control thread owns the user interface and configuration. A fourth, small role is often added: a reclaim thread that frees memory the audio thread has finished with.
The rule that makes this work is that every channel into the callback is wait-free for the callback: it can always complete its read in bounded time, whatever the other thread is doing. A mutex fails that test. If the control thread holds a lock and is descheduled, the audio thread waits for a thread with lower priority, which is priority inversion, and the deadline passes.
What the callback must never do
| Operation | Why it is dangerous | Rust trap |
|---|---|---|
| Allocate or free memory | The allocator may take a lock or ask the OS for pages | Vec::push, String formatting, Box::new, cloning a Vec, dropping the last Arc to a big object |
| Lock a mutex | Priority inversion; unbounded wait | Mutex::lock, RwLock, and channels that lock internally |
| System calls and I/O | Unbounded latency | println!, logging to a file, reading files, sockets |
| Unbounded loops | Worst case is unknown | Iterating a collection whose length the control thread controls |
| Panic | Unwinding allocates and tears down the stream | Indexing out of bounds, unwrap on a full ring buffer |
The 'Arc drop' row deserves emphasis because it looks innocent. If the callback holds the last reference to a large buffer and replaces it, Rust drops the old value right there, in the audio thread, and frees memory. The fix is shown below: never let the last reference die on the audio thread.
To catch allocations during development, the assert_no_alloc crate wraps the global allocator and reports any allocation made inside a closure you mark; enable it in debug builds around the body of your callback.
Opening a stream with cpal
cpal is the common cross-platform crate for audio device I/O in Rust. You choose a host (the platform audio API), a device, and a stream configuration, then hand it a data callback that fills each output buffer. The code below targets cpal 0.18, where build_output_stream takes the configuration by value plus an optional initialisation timeout; older releases took a reference to the configuration, so check the signature for the version you pin.
use cpal::traits::{DeviceTrait, HostTrait, StreamTrait};
use rtrb::RingBuffer;
fn main() -> anyhow::Result<()> {
let host = cpal::default_host();
let device = host.default_output_device().expect("no output device");
let mut config: cpal::StreamConfig = device.default_output_config()?.into();
config.buffer_size = cpal::BufferSize::Fixed(256); // request; the backend may refuse
let channels = config.channels as usize;
// Preallocate everything the callback touches, before the stream starts.
let (mut producer, mut consumer) = RingBuffer::<f32>::new(48_000); // ~1 s mono headroom
let shared = std::sync::Arc::new(Shared::default());
let cb_shared = shared.clone();
let stream = device.build_output_stream(
config,
move |out: &mut [f32], _info: &cpal::OutputCallbackInfo| {
let gain = cb_shared.gain(); // one atomic load per block
for frame in out.chunks_mut(channels) {
let s = match consumer.pop() {
Ok(v) => v * gain,
Err(_) => { cb_shared.count_underrun(); 0.0 } // silence, never block
};
for ch in frame.iter_mut() { *ch = s; }
}
},
|err| eprintln!("stream error: {err}"), // runs off the audio path
None,
)?;
stream.play()?;
spawn_decoder(producer); // worker thread fills the ring
run_ui(shared);
Ok(())
}Notice what the callback does: one atomic load, a loop bounded by the buffer length, a wait-free pop and a counter increment. An underrun produces silence and a counted event instead of blocking. A fixed buffer size is a request; some backends ignore it, so read the size the callback actually receives and size your processing for the largest you observe. Popping one sample at a time is clear but slow for large blocks; rtrb also offers chunk APIs such as read_chunk that let you copy a contiguous region at once.
Moving samples: sizing the ring buffer
The ring buffer between the worker and the callback absorbs the worker's irregular timing. It is single-producer single-consumer, so one side only advances the write index and the other only advances the read index, and neither ever waits. rtrb's RingBuffer::new allocates once and returns a Producer and a Consumer that you move to their threads; push and pop return a Result instead of blocking, so a full or empty buffer is an ordinary branch.
Worked example. A worker decodes compressed audio in packets of 960 frames (20 ms at 48 kHz) and occasionally stalls for up to 15 ms when it reads from disk. The callback consumes 256 frames every 5.33 ms. To survive the stall without an underrun, the ring must hold at least 15 ms, 720 frames, of audio at the moment the stall begins, plus one packet of slack because the worker writes in 960-frame bursts. A target fill of about 2,000 frames (around 42 ms) is a sensible starting point; it adds that much latency, which is fine for playback and too much for live monitoring. In a live path you cannot hide a 15 ms stall with buffering at a 5 ms deadline, so the stall itself has to go: move the disk read to another thread or prefetch.
Keep the worker writing ahead by watching the fill level (rtrb's producer reports how many slots are free) and decoding more when it falls below target. If producer and consumer run on different clocks, for example a network stream and a local sound card, the fill level drifts over minutes; that needs adaptive resampling, covered in clock drift compensation and resampling.
Parameters and commands without locks
Small continuous parameters, such as gain or a filter cutoff, are best shared through atomics. Rust has no atomic f32, but you can store the bits of an f32 in an AtomicU32.
use std::sync::atomic::{AtomicU32, AtomicU64, Ordering};
#[derive(Default)]
struct Shared { gain_bits: AtomicU32, underruns: AtomicU64 }
impl Shared {
fn set_gain(&self, g: f32) { self.gain_bits.store(g.to_bits(), Ordering::Relaxed) }
fn gain(&self) -> f32 { f32::from_bits(self.gain_bits.load(Ordering::Relaxed)) }
fn count_underrun(&self) { self.underruns.fetch_add(1, Ordering::Relaxed); }
}
// Inside the DSP: smooth towards the target so a jump does not click.
struct Smoothed { current: f32, coeff: f32 }
impl Smoothed {
fn next(&mut self, target: f32) -> f32 {
self.current += self.coeff * (target - self.current); // one-pole low-pass
self.current
}
}Relaxed ordering is enough for an independent value like gain, because nothing else depends on seeing it in a particular order. Smoothing matters: applying a new gain abruptly at a block boundary produces an audible step, known as zipper noise, so ramp towards the target over a few milliseconds.
Discrete events, such as 'start this note' or 'load preset 3', go through a second SPSC ring of a small Copy enum. The callback drains it at the start of each block, with a fixed maximum per block so a flood of commands cannot blow the deadline.
Swapping large objects safely
Some changes are too big for an atomic: a new impulse response for a convolution reverb, a rebuilt processing graph, a model's weights. Build the new object on the worker thread, where allocation is fine, and send a Box of it to the callback through a ring buffer. The callback swaps it in with std::mem::replace and pushes the old Box into a garbage ring instead of letting it drop. A reclaim thread pops the garbage ring periodically and drops the old objects there.
// In the callback, at the start of a block:
if let Ok(new_graph) = graph_in.pop() {
let old = std::mem::replace(&mut graph, new_graph);
if garbage_out.push(old).is_err() {
// Garbage ring full: keep the old one alive rather than free it here.
// Size the ring so this never happens, and count it if it does.
}
}
graph.process(out);The same rule applies to Arc: if both the callback and another thread hold an Arc, make sure the other thread is always the one that drops last. Crates exist that package this deferred-reclamation pattern, but the hand-written version above is short enough to own.
Inside the DSP loop
- Allocate every scratch buffer at start-up, sized for the largest block the backend can deliver, and slice it per call.
- Process in blocks rather than sample by sample where possible; it lets the compiler vectorise and amortises per-call overhead.
- Guard against denormals: when filters decay towards silence, values can become subnormal floats that are very slow on some CPUs. Enable flush-to-zero for the audio thread where your platform allows it, or add a tiny offset in feedback paths.
- Avoid panics: prefer iterators and chunks_exact over manual indexing, and never unwrap a ring-buffer result in the callback.
- Keep anything with unpredictable cost, such as neural inference, on a worker and feed results through a ring, accepting one block of added latency.
Measuring: xruns and callback time
An xrun is an underrun or overrun at the device. You cannot fix what you do not count, so measure two things from inside the callback without logging there: how often the ring was empty, and how long each callback took relative to its deadline. Read a monotonic clock at entry and exit, and record the duration into a fixed array of histogram buckets with atomic increments. A control-thread timer reads the counters and logs them once a second.
Test under stress: run with a busy CPU, while the window is being resized, after the laptop wakes from sleep, and with Bluetooth output, which often uses larger and less predictable buffers. A pipeline that never glitches on an idle desktop proves little.
Thread priority depends on the backend. Callback threads created by platform audio APIs are typically given elevated scheduling, but this varies by platform; if you run your own time-critical threads, the audio_thread_priority crate requests real-time priority from each operating system.
Failure modes and trade-offs
| Symptom | Likely cause | Fix |
|---|---|---|
| Periodic clicks under load | Worker too slow or ring too small | Raise target fill; profile the worker; prefetch |
| Rare glitches, idle machine | Allocation, lock or Arc drop in the callback | Run with assert_no_alloc; audit drops; move I/O out |
| Glitch when changing a setting | Unsmoothed parameter or a big object freed in the callback | Smooth; use the garbage ring |
| Slow drift to underrun or overflow over minutes | Producer and device clocks differ | Adaptive resampling on fill level |
| CPU spikes as sound fades out | Denormals in feedback filters | Flush-to-zero or offset |
The core trade-off is latency against robustness: every millisecond of buffering absorbs jitter and is heard as delay. The second is simplicity against throughput: per-sample code is easy to read, while block processing is faster. Lock-free structures are well covered in lock-free threading if you want to understand why SPSC rings need no compare-and-swap.
What to do next
- Pick a block size from the deadline table and set a callback budget of under half that deadline.
- Open a cpal stream that outputs a sine wave from a preallocated oscillator, and confirm the block sizes the callback really receives.
- Add an rtrb ring fed by a worker thread, size it with the worked example, and count underruns with an atomic.
- Move every parameter to atomics with smoothing, and every discrete event to a bounded command ring.
- Add the garbage ring before your first hot-swap feature, not after the first glitch report.
- Wrap the callback in assert_no_alloc in debug builds, add a callback-duration histogram, and stress-test under CPU load and device changes.