Every audio system eventually has to change a sample rate. A podcast recorded at 44.1 kHz is mixed into a 48 kHz video timeline. A speech model was trained on 16 kHz audio, but the microphone delivers 48 kHz. A Bluetooth headset and a USB interface each run on their own crystal and disagree by a few parts per million. The right technique differs in each case, and the wrong one either adds audible artifacts or wastes CPU.
This article explains the techniques, puts numbers on their cost, and gives you code to measure the result. The internal architecture of a production polyphase converter is covered in sample-rate conversion architecture, and tracking two drifting clocks in clock drift compensation. Here the question is which technique to pick and how to prove it works.
What a resampler actually has to do
A digital signal sampled at rate fs can only represent frequencies below fs/2, the Nyquist frequency. Changing the rate means answering a question: what value would the original continuous signal have had at these new instants? For a band-limited signal the ideal answer is a sum of sinc functions centred on every input sample, and every practical resampler approximates that infinite sum.
Two distinct errors appear when the approximation is poor. When you raise the rate, the spectrum of the input repeats around multiples of the old sample rate. Those copies, called images, must be removed or they appear as high-frequency garbage above the old Nyquist frequency. When you lower the rate, any content above the new Nyquist frequency folds back into the audible band as aliases. A 21 kHz tone resampled from 48 kHz to 16 kHz without filtering reappears at 5 kHz, because 21 minus 16 is 5. Aliases land inside the band you care about and cannot be removed afterwards.
So a resampler is a low-pass filter, with its cutoff at the lower of the two Nyquist frequencies, evaluated at the instants the output needs.
The ladder of techniques
| Technique | How it works | Quality | Cost per output | Ratios |
|---|---|---|---|---|
| Drop or repeat samples | Keep every M-th sample, or repeat samples | Poor: full aliasing and imaging | None | Integer only |
| Linear interpolation | Straight line between neighbours | Weak image and alias rejection; dulls highs | 2 multiplies | Any |
| Cubic or Hermite | Polynomial through 4 neighbours | Better than linear, still far from transparent | About 4 to 8 multiplies | Any |
| Windowed sinc | Truncated, windowed sinc evaluated per output | Set by window and length; can be transparent | N multiplies plus kernel evaluation | Any |
| Rational polyphase | Precomputed sinc filter split into L phases | As good as the filter you design | K taps (filter length / L) | L/M with small integers |
| FFT resampling | Truncate or zero-pad the spectrum of the whole buffer | Exact for periodic signals; edge ringing otherwise | O(n log n) per buffer | Any length ratio, offline |
| Arbitrary ratio (Farrow, interpolated table) | Polyphase with many phases or a polynomial in the fractional delay | Near polyphase quality | K taps plus interpolation | Any, including time-varying |
Rational polyphase is the workhorse whenever both rates are fixed, which covers 44.1, 48, 96 and 16 kHz conversions. Arbitrary-ratio converters handle everything else, including asynchronous sample rate conversion (ASRC), where the ratio drifts slowly. Linear interpolation survives in prototypes, and is a common cause of training audio that differs from production.
How polyphase works and why it is cheap
Reduce the rate ratio fs_out / fs_in to lowest terms L/M. For 44.1 to 48 kHz the greatest common divisor of 44,100 and 48,000 is 300, so L = 160 and M = 147. The textbook procedure inserts L-1 zeros between samples, low-pass filters at L times the input rate, then keeps every M-th sample. Done literally, that runs a long filter at 7.056 MHz on mostly zeros and discards 146 of every 147 results.
For a kept output m, its position on the upsampled grid is t = mM. Only every L-th filter tap lines up with a real input sample, and which taps those are depends only on p = t mod L. Split the filter into L sub-filters, called phases, of K = length / L taps each. Each output is then one K-tap dot product between phase p and the K most recent input samples ending at index i = t div L.
import numpy as np
from math import gcd
from scipy.signal import firwin
def design(L, M, taps_per_phase=32, beta=8.0, rolloff=0.95):
# Cutoff as a fraction of the upsampled Nyquist: min(fs_in, fs_out)/2
# divided by L*fs_in/2 simplifies to 1/max(L, M).
h = firwin(L * taps_per_phase, rolloff / max(L, M), window=("kaiser", beta))
return h * L # restore the gain lost to zero stuffing
def polyphase_resample(x, fs_in, fs_out, h=None):
g = gcd(fs_in, fs_out)
L, M = fs_out // g, fs_in // g
if h is None:
h = design(L, M)
K = len(h) // L
phases = h[:K * L].reshape(K, L).T # phases[p][j] = h[p + j*L]
xp = np.concatenate([np.zeros(K - 1), x])
y = np.empty((len(x) * L) // M)
for m in range(len(y)):
i, p = divmod(m * M, L) # input index and phase
y[m] = phases[p] @ xp[i:i + K][::-1]
return yThis loop is written for clarity, not speed, and 32 taps per phase is only a demo default: size real filters with the estimate below. The factor of L restores the amplitude that zero stuffing removed, and the cutoff of 1/max(L, M) automatically picks the lower Nyquist frequency whether you are going up or down. The output is also delayed by half the filter length, which you must compensate if the audio has to stay aligned with video or with a second channel.
Sizing the filter: a worked example
Filter quality is set by the passband edge, the transition width and the stopband attenuation. Kaiser's estimate for the length of a windowed FIR is N = (A - 7.95) / (14.36 x df / fs), where A is stopband attenuation in dB and df is the transition width.
Take 44.1 to 48 kHz for music, with the passband kept flat to 20 kHz, the stopband starting at 22.05 kHz and 100 dB of attenuation. The filter runs at the intermediate rate of 7.056 MHz, so df / fs = 2,050 / 7,056,000, about 0.00029. That gives N of roughly 92 / 0.00417, or about 22,000 taps in total, but only about 138 taps per phase. Each output costs 138 multiply-adds, and 48,000 outputs per second cost about 6.6 million multiply-adds per second per channel. The latency is half the filter, about 11,000 samples at 7.056 MHz, or roughly 1.6 ms.
Now take 48 to 16 kHz for a speech model. Here L = 1 and M = 3, the cutoff is 8 kHz, and a passband to 7.2 kHz with 80 dB of rejection gives df / fs = 800 / 48,000. The estimate is about 300 taps, evaluated once per output at 16 kHz, so roughly 4.8 million multiply-adds per second. These are design estimates; the real check is the measurement below.
Using library resamplers
The common Python options each wrap one technique.
import torch, torchaudio
from scipy.signal import resample_poly, resample
# Rational polyphase; default window is ('kaiser', 5.0), zero padding at the edges
y1 = resample_poly(x, up=1, down=3)
# Windowed-sinc resampler; defaults lowpass_filter_width=6, rolloff=0.99,
# resampling_method="sinc_interp_hann". Runs on GPU tensors and batches.
y2 = torchaudio.functional.resample(torch.from_numpy(x).float(), 48000, 16000)
# A reusable transform caches its kernel; build it once, outside the data loop
to16k = torchaudio.transforms.Resample(orig_freq=48000, new_freq=16000)
# FFT method: treats the buffer as one period of a periodic signal
y3 = resample(x, num=len(x) // 3)resample_poly is polyphase with a filter designed from its window argument; raise the Kaiser beta for deeper rejection. The torchaudio resampler is a windowed sinc whose sharpness grows with lowpass_filter_width and whose cutoff sits at rolloff times the lower Nyquist frequency. scipy.signal.resample is the FFT method: fine for a whole periodic buffer, wrong for streaming, because it assumes the buffer wraps around and rings at both edges. In C and C++, libsoxr and libsamplerate are long-standing high-quality choices.
Arbitrary and drifting ratios
A ratio such as 48,000.37 / 47,999.81, from two real crystals, reduces to absurd integers. Arbitrary-ratio converters compute a fractional position for each output: the integer part selects input samples and the fractional part, mu, selects the filter.
There are two standard ways to turn mu into coefficients. An interpolated table stores a polyphase bank with many phases, such as 256 or 1,024, and linearly interpolates between the two phases that bracket mu. A Farrow structure instead represents each tap as a low-order polynomial in mu, so the output is a short polynomial evaluated with Horner's rule.
Asynchronous sample rate conversion closes the loop. A buffer sits between the producer clock and the consumer clock, a slow controller watches its fill level, and it nudges the ratio to hold the fill near its target. The controller must be slow, or its corrections become audible as wow. That loop is the subject of clock drift compensation,.
Measuring quality instead of trusting it
The most revealing downsampling test is a tone just above the new Nyquist frequency: an ideal converter removes it completely, so any residue is pure alias.
import numpy as np
from scipy.signal import resample_poly
fs = 48000
t = np.arange(2 * fs) / fs
x = 0.5 * np.sin(2 * np.pi * 21000 * t) # above the 8 kHz Nyquist of 16 kHz
def level_db(y, fs, f):
w = np.hanning(len(y))
Y = np.abs(np.fft.rfft(y * w)) / (w.sum() / 2)
k = round(f * len(y) / fs)
return 20 * np.log10(Y[k - 3:k + 4].max() + 1e-15)
naive = x[::3] # drop samples, no filter
good = resample_poly(x, 1, 3)
print("alias at 5 kHz, naive:", level_db(naive, 16000, 5000))
print("alias at 5 kHz, poly :", level_db(good, 16000, 5000))The naive version reproduces the tone at 5 kHz near its original level, about -6 dBFS; the filtered version should push it far down, and how far is what you are measuring.
- Sweep a tone across the whole input band and plot a spectrogram of the output. Aliases appear as lines moving in the wrong direction.
- Feed an impulse and inspect the output to read the delay and check for pre-ringing, which steep linear-phase filters produce before transients.
- Compare a long recording resampled in one call with the same recording resampled in streaming chunks. Any difference at chunk boundaries is a state-handling bug.
Resampling in ML audio pipelines
In machine learning, the model learns the resampler. If training data was converted with a high-quality windowed sinc and the production service uses linear interpolation, the serving audio has less energy near the band edge and more aliasing than anything the model saw in training. That distribution shift shows up as a stubborn accuracy gap.
The rules are short. Pin one resampler, with its parameters, in a shared function used by both training and serving. Resample each file once and cache it rather than resampling every epoch. When resampling must happen on the fly, construct the transform once so its kernel is cached, and batch it on the GPU. Record the source rate as metadata: a 44.1 kHz file read as 48 kHz silently shifts pitch and speed by about 9 percent. If the model consumes a codec's output, the interaction between resampling and coding is covered in audio codecs, in depth.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Metallic tones or whistles in speech | Decimation without an anti-alias filter | Filter before discarding samples; use polyphase |
| Dull, muffled music | Cutoff too low or transition band too wide | Raise rolloff; lengthen the filter |
| Clicks every buffer | Filter state reset per chunk, or FFT method on chunks | Keep history between calls; use a streaming resampler |
| Audio drifts out of sync with video | Filter delay not compensated, or wrong ratio | Subtract the group delay; assert the ratio from metadata |
| Model accuracy lower in production | Different resampler in training and serving | One shared resampling function, pinned versions |
Compute output lengths with integer arithmetic; float truncation in two places produces off-by-one lengths that break alignment. If you work with raw buffers, PCM audio fundamentals covers the sample formats these filters operate on.
Trade-offs
A longer filter rejects more aliasing but adds delay and compute. Linear-phase filters pre-ring before transients; minimum-phase designs avoid that and cut delay at the cost of phase distortion near the cutoff. Offline work can afford long filters and FFT methods; real-time work wants a short streaming converter. The cheapest resampler is the one you remove: pick one rate at the system edge.
What to do next
- List every point in your system where the sample rate changes, and the library and parameters used at each one.
- Run the 21 kHz alias test and a sine sweep through each resampler and record the alias level and passband edge.
- Pick one resampling function for training and serving, and make both import it.
- Compute output lengths with integer arithmetic and test chunked processing against one-shot processing.
- Measure and compensate the group delay wherever audio must stay aligned with video or another channel.
- If two devices have independent clocks, add an arbitrary-ratio converter driven by buffer fill, with a slow controller.