A real-time audio program has one job that dominates every other design decision: the audio device asks for the next block of samples at a fixed rhythm, and the program must hand them over before the deadline every single time. Miss once and the listener hears a click, a dropout or a stutter. Average speed does not matter; the worst case does.

Rust is a good fit because it has no garbage collector to pause the audio thread and its type system catches data races at compile time. It is not a complete fit: nothing in the type system stops you allocating, locking or doing I/O inside the callback, and those are exactly the things that cause glitches. This article shows how to structure a Rust audio pipeline so the callback stays deterministic, using the cpal crate for device I/O and the rtrb crate for lock-free ring buffers, and how to measure whether it works. For the signal-processing stages themselves, see real-time audio architecture.

Advertisement

The deadline arithmetic

The device consumes audio in blocks of a fixed number of frames (one frame is one sample per channel). The time you have to produce a block is the block size divided by the sample rate. At 48 kHz:

Block size (frames)Deadline per callbackTypical use
641.33 msLive instruments, monitoring
1282.67 msLow-latency effects, games
2565.33 msVoice calls, most interactive apps
51210.67 msPlayback, streaming, heavy DSP
102421.33 msBackground playback

The deadline is not your budget. The operating system, the driver and other processes share the same core, so a common rule is to keep the callback's worst-case time well under half the deadline. Total latency is larger still: it adds the device's own buffering, any ring buffer between your producer and the callback, and the hardware converters. A block size is therefore a trade between latency and safety margin, and the right value is the largest one your use case tolerates.

Three thread roles

A robust pipeline separates work by how predictable it is. The audio callback runs on a thread created by the audio backend and must only do bounded, predictable work. A worker thread does everything that can stall: decoding files, receiving network packets, running a neural model, resampling large chunks. A control thread owns the user interface and configuration. A fourth, small role is often added: a reclaim thread that frees memory the audio thread has finished with.

Three thread roles; only lock-free, allocation-free paths cross into the audio callbackControl threadUI, config, commandsWorker threaddecode, network, MLReclaim threaddrops retired objectsAtomicsgain, mute, cutoffCommand ringSPSC, preallocatedSample ringSPSC f32 framesGarbage ringold boxes go backAudio callbackpop commandsread paramspull frames / DSPwrite device bufferbump xrun counterstorepushpushloadpoppopretiredpop + dropDevice / driverdeadline every block
Control and worker threads talk to the audio callback only through atomics and single-producer single-consumer ring buffers. Objects the callback no longer needs are sent back through a garbage ring and dropped on another thread.

The rule that makes this work is that every channel into the callback is wait-free for the callback: it can always complete its read in bounded time, whatever the other thread is doing. A mutex fails that test. If the control thread holds a lock and is descheduled, the audio thread waits for a thread with lower priority, which is priority inversion, and the deadline passes.

Advertisement

What the callback must never do

OperationWhy it is dangerousRust trap
Allocate or free memoryThe allocator may take a lock or ask the OS for pagesVec::push, String formatting, Box::new, cloning a Vec, dropping the last Arc to a big object
Lock a mutexPriority inversion; unbounded waitMutex::lock, RwLock, and channels that lock internally
System calls and I/OUnbounded latencyprintln!, logging to a file, reading files, sockets
Unbounded loopsWorst case is unknownIterating a collection whose length the control thread controls
PanicUnwinding allocates and tears down the streamIndexing out of bounds, unwrap on a full ring buffer

The 'Arc drop' row deserves emphasis because it looks innocent. If the callback holds the last reference to a large buffer and replaces it, Rust drops the old value right there, in the audio thread, and frees memory. The fix is shown below: never let the last reference die on the audio thread.

To catch allocations during development, the assert_no_alloc crate wraps the global allocator and reports any allocation made inside a closure you mark; enable it in debug builds around the body of your callback.

Opening a stream with cpal

cpal is the common cross-platform crate for audio device I/O in Rust. You choose a host (the platform audio API), a device, and a stream configuration, then hand it a data callback that fills each output buffer. The code below targets cpal 0.18, where build_output_stream takes the configuration by value plus an optional initialisation timeout; older releases took a reference to the configuration, so check the signature for the version you pin.

use cpal::traits::{DeviceTrait, HostTrait, StreamTrait};
use rtrb::RingBuffer;

fn main() -> anyhow::Result<()> {
    let host = cpal::default_host();
    let device = host.default_output_device().expect("no output device");
    let mut config: cpal::StreamConfig = device.default_output_config()?.into();
    config.buffer_size = cpal::BufferSize::Fixed(256);   // request; the backend may refuse
    let channels = config.channels as usize;

    // Preallocate everything the callback touches, before the stream starts.
    let (mut producer, mut consumer) = RingBuffer::<f32>::new(48_000);  // ~1 s mono headroom
    let shared = std::sync::Arc::new(Shared::default());
    let cb_shared = shared.clone();

    let stream = device.build_output_stream(
        config,
        move |out: &mut [f32], _info: &cpal::OutputCallbackInfo| {
            let gain = cb_shared.gain();                  // one atomic load per block
            for frame in out.chunks_mut(channels) {
                let s = match consumer.pop() {
                    Ok(v) => v * gain,
                    Err(_) => { cb_shared.count_underrun(); 0.0 }   // silence, never block
                };
                for ch in frame.iter_mut() { *ch = s; }
            }
        },
        |err| eprintln!("stream error: {err}"),           // runs off the audio path
        None,
    )?;
    stream.play()?;
    spawn_decoder(producer);                               // worker thread fills the ring
    run_ui(shared);
    Ok(())
}

Notice what the callback does: one atomic load, a loop bounded by the buffer length, a wait-free pop and a counter increment. An underrun produces silence and a counted event instead of blocking. A fixed buffer size is a request; some backends ignore it, so read the size the callback actually receives and size your processing for the largest you observe. Popping one sample at a time is clear but slow for large blocks; rtrb also offers chunk APIs such as read_chunk that let you copy a contiguous region at once.

Moving samples: sizing the ring buffer

The ring buffer between the worker and the callback absorbs the worker's irregular timing. It is single-producer single-consumer, so one side only advances the write index and the other only advances the read index, and neither ever waits. rtrb's RingBuffer::new allocates once and returns a Producer and a Consumer that you move to their threads; push and pop return a Result instead of blocking, so a full or empty buffer is an ordinary branch.

Worked example. A worker decodes compressed audio in packets of 960 frames (20 ms at 48 kHz) and occasionally stalls for up to 15 ms when it reads from disk. The callback consumes 256 frames every 5.33 ms. To survive the stall without an underrun, the ring must hold at least 15 ms, 720 frames, of audio at the moment the stall begins, plus one packet of slack because the worker writes in 960-frame bursts. A target fill of about 2,000 frames (around 42 ms) is a sensible starting point; it adds that much latency, which is fine for playback and too much for live monitoring. In a live path you cannot hide a 15 ms stall with buffering at a 5 ms deadline, so the stall itself has to go: move the disk read to another thread or prefetch.

Keep the worker writing ahead by watching the fill level (rtrb's producer reports how many slots are free) and decoding more when it falls below target. If producer and consumer run on different clocks, for example a network stream and a local sound card, the fill level drifts over minutes; that needs adaptive resampling, covered in clock drift compensation and resampling.

Parameters and commands without locks

Small continuous parameters, such as gain or a filter cutoff, are best shared through atomics. Rust has no atomic f32, but you can store the bits of an f32 in an AtomicU32.

use std::sync::atomic::{AtomicU32, AtomicU64, Ordering};

#[derive(Default)]
struct Shared { gain_bits: AtomicU32, underruns: AtomicU64 }

impl Shared {
    fn set_gain(&self, g: f32) { self.gain_bits.store(g.to_bits(), Ordering::Relaxed) }
    fn gain(&self) -> f32 { f32::from_bits(self.gain_bits.load(Ordering::Relaxed)) }
    fn count_underrun(&self) { self.underruns.fetch_add(1, Ordering::Relaxed); }
}

// Inside the DSP: smooth towards the target so a jump does not click.
struct Smoothed { current: f32, coeff: f32 }
impl Smoothed {
    fn next(&mut self, target: f32) -> f32 {
        self.current += self.coeff * (target - self.current);   // one-pole low-pass
        self.current
    }
}

Relaxed ordering is enough for an independent value like gain, because nothing else depends on seeing it in a particular order. Smoothing matters: applying a new gain abruptly at a block boundary produces an audible step, known as zipper noise, so ramp towards the target over a few milliseconds.

Discrete events, such as 'start this note' or 'load preset 3', go through a second SPSC ring of a small Copy enum. The callback drains it at the start of each block, with a fixed maximum per block so a flood of commands cannot blow the deadline.

Swapping large objects safely

Some changes are too big for an atomic: a new impulse response for a convolution reverb, a rebuilt processing graph, a model's weights. Build the new object on the worker thread, where allocation is fine, and send a Box of it to the callback through a ring buffer. The callback swaps it in with std::mem::replace and pushes the old Box into a garbage ring instead of letting it drop. A reclaim thread pops the garbage ring periodically and drops the old objects there.

// In the callback, at the start of a block:
if let Ok(new_graph) = graph_in.pop() {
    let old = std::mem::replace(&mut graph, new_graph);
    if garbage_out.push(old).is_err() {
        // Garbage ring full: keep the old one alive rather than free it here.
        // Size the ring so this never happens, and count it if it does.
    }
}
graph.process(out);

The same rule applies to Arc: if both the callback and another thread hold an Arc, make sure the other thread is always the one that drops last. Crates exist that package this deferred-reclamation pattern, but the hand-written version above is short enough to own.

Inside the DSP loop

  • Allocate every scratch buffer at start-up, sized for the largest block the backend can deliver, and slice it per call.
  • Process in blocks rather than sample by sample where possible; it lets the compiler vectorise and amortises per-call overhead.
  • Guard against denormals: when filters decay towards silence, values can become subnormal floats that are very slow on some CPUs. Enable flush-to-zero for the audio thread where your platform allows it, or add a tiny offset in feedback paths.
  • Avoid panics: prefer iterators and chunks_exact over manual indexing, and never unwrap a ring-buffer result in the callback.
  • Keep anything with unpredictable cost, such as neural inference, on a worker and feed results through a ring, accepting one block of added latency.

Measuring: xruns and callback time

An xrun is an underrun or overrun at the device. You cannot fix what you do not count, so measure two things from inside the callback without logging there: how often the ring was empty, and how long each callback took relative to its deadline. Read a monotonic clock at entry and exit, and record the duration into a fixed array of histogram buckets with atomic increments. A control-thread timer reads the counters and logs them once a second.

Test under stress: run with a busy CPU, while the window is being resized, after the laptop wakes from sleep, and with Bluetooth output, which often uses larger and less predictable buffers. A pipeline that never glitches on an idle desktop proves little.

Thread priority depends on the backend. Callback threads created by platform audio APIs are typically given elevated scheduling, but this varies by platform; if you run your own time-critical threads, the audio_thread_priority crate requests real-time priority from each operating system.

Failure modes and trade-offs

SymptomLikely causeFix
Periodic clicks under loadWorker too slow or ring too smallRaise target fill; profile the worker; prefetch
Rare glitches, idle machineAllocation, lock or Arc drop in the callbackRun with assert_no_alloc; audit drops; move I/O out
Glitch when changing a settingUnsmoothed parameter or a big object freed in the callbackSmooth; use the garbage ring
Slow drift to underrun or overflow over minutesProducer and device clocks differAdaptive resampling on fill level
CPU spikes as sound fades outDenormals in feedback filtersFlush-to-zero or offset

The core trade-off is latency against robustness: every millisecond of buffering absorbs jitter and is heard as delay. The second is simplicity against throughput: per-sample code is easy to read, while block processing is faster. Lock-free structures are well covered in lock-free threading if you want to understand why SPSC rings need no compare-and-swap.

What to do next

  1. Pick a block size from the deadline table and set a callback budget of under half that deadline.
  2. Open a cpal stream that outputs a sine wave from a preallocated oscillator, and confirm the block sizes the callback really receives.
  3. Add an rtrb ring fed by a worker thread, size it with the worked example, and count underruns with an atomic.
  4. Move every parameter to atomics with smoothing, and every discrete event to a bounded command ring.
  5. Add the garbage ring before your first hot-swap feature, not after the first glitch report.
  6. Wrap the callback in assert_no_alloc in debug builds, add a callback-duration histogram, and stress-test under CPU load and device changes.
Key takeaway: A real-time audio pipeline in Rust is a deadline-driven design. Keep the cpal callback to bounded, allocation-free, lock-free work; feed it through SPSC ring buffers such as rtrb and atomics; build large objects elsewhere and send old ones back to be dropped off the audio thread; and size buffers from measured worst cases. Rust removes data races but not real-time hazards, so verify with allocation checks, xrun counters and callback-time histograms under stress.