A single microphone hears everything around it equally. A beamformer combines several microphones so that sound from one direction adds up coherently while sound from other directions partly cancels. It is the first stage in almost every far-field voice device, from smart speakers to conference bars to laptops, and it decides how much noise and reverberation reach the echo canceller, the noise suppressor and the speech recogniser downstream.
Beamforming is often presented as a formula, but most of the engineering is in the choices around it: array geometry, what noise model the weights assume, how the noise statistics are estimated, how the beam is steered, how much the design trusts the microphones to be identical, and where echo cancellation goes. This article builds the architecture from first principles, runs a worked example with computed numbers, and ends with failure modes and a checklist.
What a beamformer can and cannot do
A beamformer is a spatial filter. In each frequency band it multiplies every microphone signal by a complex weight and adds them. The weights choose a pattern of gain over direction: a look direction with unity gain, and lower gain elsewhere. It works best against noise that arrives from directions other than the talker, and against diffuse noise such as room babble and air conditioning that arrives from all directions at once.
Two numbers describe a design. The directivity index (DI) is the gain against a diffuse noise field, in decibels: how much more of the look direction you hear than of everything else. The white noise gain (WNG) is the gain against noise that is independent at each microphone, such as the microphones' own self-noise. Delay-and-sum with N microphones has a WNG of 10 log10 N, the best possible. Designs that push DI higher do so by subtracting microphones from each other, and that amplifies independent noise, so WNG falls. Almost every design decision is a trade between those two numbers.
A beamformer cannot remove noise coming from the same direction as the talker, and it only partly removes reverberation, because reflections arrive from all directions including the look direction. Dedicated dereverberation and single-channel noise suppression handle what remains.
Geometry comes first
A plane wave from angle theta reaches microphone m with a delay tau_m that depends on its position. In the frequency domain, a delay is a phase rotation, so the array response to that direction at frequency f is the steering vector d(f), whose m-th entry is exp(-j 2 pi f tau_m). Every beamformer in this article is built from steering vectors.
Spacing sets the upper frequency limit. If adjacent microphones are more than half a wavelength apart, two different directions produce the same phase pattern and the beam grows copies of itself, called grating lobes. With sound at 343 m/s, a spacing of d gives an aliasing limit of 343 / (2d) Hz. For 2 cm spacing that is 8,575 Hz, just enough for 16 kHz audio; at 4 cm it drops to about 4.3 kHz. Aperture, the overall size of the array, sets the lower limit: at 300 Hz the wavelength is over a metre, and a 6 cm array sees almost no phase difference between microphones, so delay-and-sum has almost no directivity there.
Linear arrays are cheap and suit devices facing one way, such as a laptop or TV bar, but cannot tell front from back. Circular arrays give the same performance in every horizontal direction, which is why smart speakers use them, often with a centre microphone. Every microphone must be sampled by one clock; separate converters with drifting clocks destroy the phase relationships the beamformer relies on.
The processing architecture
Practical beamformers run in the short-time Fourier transform (STFT) domain: each microphone stream is cut into overlapping frames, transformed, processed independently in every frequency bin, and transformed back. That makes delays into multiplications and lets each bin have its own weights, which matters because array behaviour changes completely between 200 Hz and 6 kHz.
Around the core weight-and-sum stage sit three supporting blocks. A direction finder decides where to steer. A noise-statistics block estimates the spatial covariance of noise during frames without speech, using a voice activity detector or a speech-presence probability. A post-filter applies a per-bin gain to the beam output, because the beamformer leaves residual noise that a single-channel stage can reduce further.
Delay-and-sum and filter-and-sum
Delay-and-sum is the baseline: align every microphone to the look direction and average. Its weights are w = d / N. It never amplifies microphone self-noise, it is robust to small errors in position and gain, and its pattern is fixed by geometry. Its weakness is low frequencies, where the array is small compared with the wavelength and the pattern is almost omnidirectional.
Filter-and-sum generalises this by giving each microphone a filter rather than a pure delay, which in the STFT domain simply means arbitrary complex weights per bin. All the designs below are filter-and-sum beamformers that differ only in how the weights are chosen.
MVDR and superdirective weights
The minimum variance distortionless response (MVDR) beamformer chooses weights that minimise output noise power while keeping the look direction at exactly unity gain. If R is the noise spatial covariance matrix in a bin, the solution is w = R^-1 d / (d^H R^-1 d). The denominator enforces the distortionless constraint; the numerator steers nulls toward wherever the noise is strongest.
R can come from two places. A fixed design assumes a noise model. Assuming spherically diffuse noise gives a coherence between microphones at distance r of sinc(2 pi f r / c), and MVDR with that matrix is the superdirective beamformer, which achieves high DI from a small array. An adaptive design estimates R from the data during noise-only frames, so it can place nulls on a specific television or fan. Adaptive designs track real noise but suffer when speech leaks into the noise estimate: the beamformer then treats the talker as noise and cancels them, an effect called signal cancellation, which reverberation makes worse because reflections of the talker arrive from other directions.
Both need diagonal loading: adding a small multiple of the identity to R before inverting. It bounds the weights, keeps WNG from collapsing, and absorbs microphone mismatch. The amount of loading is the main tuning knob, as the next example shows.
import numpy as np
C = 343.0
def steering(pos_m, f, look=np.array([1.0, 0.0])):
tau = pos_m @ look / C # seconds per mic for a plane wave
return np.exp(-2j * np.pi * f * tau)
def mvdr_weights(R, d, loading):
R = R + loading * np.trace(R).real / len(d) * np.eye(len(d))
Rd = np.linalg.solve(R, d)
return Rd / (d.conj() @ Rd)
def diffuse_coherence(pos_m, f):
r = np.linalg.norm(pos_m[:, None] - pos_m[None, :], axis=-1)
return np.sinc(2 * f * r / C) # np.sinc(x) = sin(pi x) / (pi x)
def di_wng(w, d, gamma):
signal = abs(w.conj() @ d) ** 2
di = 10 * np.log10(signal / (w.conj() @ gamma @ w).real)
wng = 10 * np.log10(signal / (w.conj() @ w).real)
return di, wng
def update_noise_cov(R, X_bin, speech_prob, alpha=0.98):
# X_bin: complex vector of mic STFT values in one bin, one frame.
# Only learn from frames that are probably noise, to avoid signal cancellation.
a = alpha + (1 - alpha) * speech_prob
return a * R + (1 - a) * np.outer(X_bin, X_bin.conj())
pos = np.array([[i * 0.02, 0.0] for i in range(4)]) # 4 mics, 2 cm, endfire look
for f in (300, 1000, 3000):
d, gamma = steering(pos, f), diffuse_coherence(pos, f)
print(f, "das", di_wng(d / 4, d, gamma))
for load in (1e-4, 1e-2, 1e-1):
print(f, load, di_wng(mvdr_weights(gamma, d, load), d, gamma))
Worked example: four microphones, 2 cm apart
Take a linear array of four microphones with 2 cm spacing and steer along the array axis (endfire), the direction where small arrays work best. The code above computes DI and WNG for delay-and-sum and for the superdirective design with three amounts of diagonal loading. Because the coherence matrix has ones on its diagonal, the loading values are relative to the diagonal.
| Design | 300 Hz DI / WNG | 1 kHz DI / WNG | 3 kHz DI / WNG |
|---|---|---|---|
| Delay-and-sum | 0.1 / 6.0 dB | 0.9 / 6.0 dB | 4.7 / 6.0 dB |
| Superdirective, loading 0.0001 | 7.1 / -24.0 dB | 9.6 / -17.9 dB | 11.6 / -11.6 dB |
| Superdirective, loading 0.01 | 5.9 / -8.8 dB | 7.2 / -5.9 dB | 10.0 / -3.4 dB |
| Superdirective, loading 0.1 | 3.2 / -3.1 dB | 5.9 / 0.6 dB | 8.6 / 2.3 dB |
Delay-and-sum gives its full 6 dB WNG (10 log10 4) everywhere but almost no directivity below 1 kHz: 0.1 dB at 300 Hz means the array is effectively a single microphone there. The lightly loaded superdirective design reaches 7.1 dB DI at 300 Hz, but a WNG of -24 dB means each microphone's self-noise and any gain or phase mismatch are amplified about 250 times in power. In a quiet room, that hiss would be worse than the noise you removed. Moderate loading of 0.01 keeps most of the directivity and limits the penalty to under 9 dB; heavy loading of 0.1 trades a few dB of DI for a WNG near zero.
A sensible product design follows directly. Use frequency-dependent loading: heavy at low frequencies, where superdirectivity is expensive, and light at high frequencies, where it is cheap. Set a WNG floor, for example -10 dB, and increase loading per bin until it is met. The theoretical endfire maximum for four microphones is 20 log10 4, about 12 dB, which the lightly loaded design approaches at 3 kHz.
Steering: direction finding versus fixed beams
An adaptive or steered beamformer needs to know where the talker is. The common estimator is GCC-PHAT, which cross-correlates pairs of microphones after whitening their spectra so that the correlation peak marks the time difference of arrival. SRP-PHAT extends it by scanning candidate directions and summing the pairwise evidence. Both are cheap and work well for one dominant talker, but they lock onto loud noise sources and reflections, so they must be gated by speech presence and smoothed over time.
Many devices avoid the problem with fixed beams: compute, say, six or eight superdirective beams at design time covering all directions, run them all, and select the beam with the best score. The score can be the energy of the speech band, or the confidence of the wake word detector run on each beam. Fixed beams have predictable weights, no signal cancellation and easy testing; the cost is running several beams and slightly lower gain between beam centres.
Where echo cancellation goes
A device that plays audio also hears itself, and the echo is usually far louder than the talker. Acoustic echo cancellation subtracts an adaptive estimate of the loudspeaker signal, and its position relative to the beamformer is the most important architectural decision in the pipeline.
AEC first, per microphone, is the cleanest: each microphone gets its own canceller, the beamformer sees echo-free signals, and adaptive weights are free to move. It costs N cancellers. Beamformer first, AEC on the beam needs only one canceller, but every change of beam weights changes the echo path the canceller has learned, so it must re-converge and echo leaks through each time the beam moves. Fixed beams make this cheaper option work, because each beam has a stable echo path and can keep its own canceller state. A hybrid uses per-microphone cancellation on a subset and fixed beams after it.
Failure modes and operations
- Microphone mismatch. Sensitivity and phase tolerances between parts behave like independent noise and are amplified by superdirective designs. Measure the spread on production units, calibrate gains, and choose loading for the worst unit, not the best.
- Signal cancellation. An adaptive beamformer learns the talker as noise when speech leaks into the covariance. Gate updates by speech presence and cap adaptation speed.
- Clock and buffering errors. A dropped sample on one channel shifts every phase. Monitor inter-channel correlation at start-up and after underruns.
- Enclosure effects. Ports, gaskets and the device body change each microphone's response. Measure steering vectors on the finished product instead of trusting free-field geometry.
- Grating lobes. A spacing chosen for a mechanical reason quietly sets an aliasing limit inside the speech band.
- Evaluate on the task. Track wake-word false-reject rate and word error rate from recorded rooms at several distances and noise types, not only DI.
Trade-offs
| Design | Strength | Weakness |
|---|---|---|
| Delay-and-sum | Robust, best WNG, simple | Little low-frequency directivity |
| Superdirective (fixed) | High DI from a small array | Amplifies self-noise and mismatch |
| Adaptive MVDR | Nulls real interferers | Signal cancellation, needs speech gating |
| Fixed beams plus selection | Predictable, AEC-friendly | Several beams to run, lower gain between beams |
What to do next
- Compute the aliasing limit and the DI and WNG curves for your geometry with the code above before committing to a layout.
- Choose fixed beams or adaptive weights based on whether your noise is mostly diffuse or mostly a few loud sources.
- Set a WNG floor and derive per-bin diagonal loading from it, using measured microphone tolerances.
- Decide the AEC position early; if you pick beam-first AEC, use fixed beams with per-beam canceller state.
- Measure steering vectors on assembled units in a quiet room and use them instead of ideal geometry.
- Build an evaluation set of real rooms and distances, and gate every change on wake-word and recognition metrics.