On a handset call the loudspeaker is at the ear, quiet, and acoustically shielded from the microphone at the mouth. On a speakerphone the loudspeaker is loud, centimetres from the microphone, sometimes in the same plastic enclosure, and the room adds reflections that arrive hundreds of milliseconds later. Without echo cancellation, the far-end talker hears their own voice come back delayed; with poor echo cancellation they hear fragments of it, or hear the near-end talker chopped off whenever both speak at once.

The general architecture of acoustic echo cancellation (AEC), with reference alignment, an adaptive filter, double-talk handling and residual suppression, is covered in acoustic echo cancellation architecture. This article is about the hands-free case specifically: what makes it harder, which design choices change, how to use the AEC your platform already provides, and how to measure whether it works. It includes a working NLMS filter and an ERLE measurement in Python so the ideas can be tested on real recordings.

Advertisement

Why hands-free is the hard case

Four properties of a speakerphone push every part of the canceller harder than a handset does.

  • Echo louder than speech. With a loudspeaker near the microphone, the echo at the microphone can be as loud as the local talker or louder, especially when the talker sits a metre or two away. The echo-to-near-end ratio is often positive in decibels, which breaks detectors that assume echo is quieter than speech.
  • Long echo paths. A meeting room's reverberation time is often a few hundred milliseconds, so echo keeps arriving long after the direct sound. The adaptive filter needs to model that tail, or the uncancelled part must be suppressed.
  • Nonlinearity. Small loudspeakers driven loud distort, and enclosures buzz and rattle. Distortion products are echo that a linear filter cannot predict from the reference.
  • A changing path. People move, laptops get tilted, phones are picked up, and volume is changed mid-call. Each change makes the converged filter wrong until it re-adapts.
Speakerphone AEC: the reference must be the signal the loudspeaker actually playsFar-end decoderremote talkerLimiter / volumelast nonlinear stageDAC + amplifierspeaker driverx(n)Loudspeakermay distort when loudRoom + enclosureecho path h, tail 100s of msMicrophoney = h*x + speech + noiseDelay + drift alignreference to mic timingreference tapLinear adaptive filterestimate of h, subtracty(n)Residual echo suppressorper-band gainse(n)Noise suppression, AGCthen encoderecho + DT estimatesAnything nonlinear after the reference tap (OS effects, speaker DSP, clipping) is echo the linear filter cannot model
A speakerphone echo path. The reference is tapped after the last nonlinear processing stage, aligned to the microphone's timing, and fed to a linear filter whose residual is cleaned up by a suppressor.

The signal model

The microphone signal is y(n) = (h * x)(n) + d(n) + s(n) + v(n): the far-end reference x convolved with the echo path h, plus a distortion term d that a linear model cannot capture, plus near-end speech s, plus noise v. A linear adaptive filter estimates h from x and y and subtracts its estimate of the echo, leaving e(n) = y(n) - (h_hat * x)(n). In the ideal case e is just s + v. In practice e also contains the distortion d, the part of the tail the filter does not cover and any misadjustment while the filter tracks a changing path. A residual echo suppressor then attenuates the frequency bands where residual echo is likely to dominate.

The key point for speakerphones is the balance between those two stages. On a handset the linear filter does most of the work. On a loud, small, distorting speaker, the linear stage may only remove 15 to 25 dB of echo, and the suppressor has to make up the rest, which is where the speakerphone's sound quality is won or lost.

Advertisement

Get the reference right

The canceller can only remove what it is told was played. The reference must therefore be tapped after every stage that changes the signal: volume control, limiter, equaliser, spatial or loudness effects. If an OS audio enhancement, a Bluetooth speaker's DSP or a TV's sound processing modifies the audio after your tap, that change becomes part of the echo path, and a nonlinear one cannot be learned. Disable such enhancements in communication mode or tap the reference after them, for example from a hardware loopback channel.

Then the reference has to be aligned with the microphone. Output and input buffers add a delay between writing x and hearing it, often tens of milliseconds and varying by device and route. A delay estimator, typically cross-correlation between reference and microphone, or a search over filter positions, places the filter window over the echo. If the speaker and microphone run on different clocks, as with a USB speaker and a built-in microphone, the delay drifts continuously; clock-drift compensation must resample one stream, or the filter chases a moving target and never converges.

The linear filter: length, cost and NLMS

The filter must span the echo tail that matters. At a 16 kHz sample rate a 250 ms tail is 4,000 taps; at 48 kHz it is 12,000. A time-domain normalised least mean squares (NLMS) filter of length L costs about 2L multiply-adds per sample, so 4,000 taps at 16 kHz is about 128 million per second, too much for many embedded devices. Production cancellers therefore work in the frequency domain with partitioned blocks, or in subbands, where the cost grows roughly with the logarithm of the length and each band can adapt at its own rate. The time-domain version is still the clearest way to understand the algorithm:

import numpy as np

def nlms_aec(x, y, taps=4000, mu=0.3, eps=1e-6, adapt=None):
    # x: reference as played (aligned to the mic), y: microphone. Returns the error e.
    w = np.zeros(taps)
    buf = np.zeros(taps)                       # most recent reference samples, newest first
    e = np.zeros(len(y))
    for n in range(len(y)):
        buf = np.roll(buf, 1); buf[0] = x[n]
        echo_hat = w @ buf
        e[n] = y[n] - echo_hat
        if adapt is None or adapt[n]:          # freeze adaptation during double talk
            w += mu * e[n] * buf / (buf @ buf + eps)
    return e, w

The step size mu, between 0 and 2, trades speed for stability: larger values converge in fewer seconds and track a moving phone faster, but leave more residual misadjustment and diverge more readily when near-end speech leaks into the update. Values of 0.1 to 0.5 are common starting points. Normalising by the reference power is what makes NLMS behave the same at quiet and loud volumes; without it, a volume change alters the effective step size.

Double talk when the echo is louder than the talker

When both ends speak, the error e contains near-end speech, and an adaptive filter that keeps updating treats that speech as echo it failed to cancel. It diverges, and the result is a burst of echo just when people are interrupting each other. The classic guard is the Geigel detector, which declares double talk when the microphone level exceeds a fraction, often one half, of the recent peak reference level. That assumes the echo is at least 6 dB quieter than the reference, which is true for a telephone line hybrid and false for a speakerphone, where the coupling can have gain. On a speakerphone a Geigel detector either fires constantly or never.

Speakerphone cancellers use measures that do not depend on absolute levels. Coherence between reference and microphone is high when the microphone is dominated by echo and drops when another source is present. Comparing the energy of the error to the energy of the estimated echo indicates whether the filter is explaining the microphone. Many implementations also run two filters: a background filter that adapts aggressively and a foreground filter that produces the output, copying the background coefficients only when they cancel better. The two-filter structure turns divergence into a recoverable event rather than an audible one, and it needs no explicit detector to be right every time.

Residual echo suppression and nonlinearity

After the linear stage, the suppressor estimates residual echo power per frequency band, usually from the reference spectrum and an estimate of how much echo the linear stage leaves, and applies a gain below one where residual echo would be audible. Too little suppression lets echo fragments through; too much clips the near-end talker during double talk and makes the call feel half-duplex. Tuning that balance is most of the work on a speakerphone. Comfort noise must be inserted where the suppressor removes signal, otherwise the far end hears the background noise switch on and off, as described in comfort noise work.

For nonlinear echo, prevention beats cancellation. Cap the maximum volume below the point where the speaker distorts, put a limiter before the reference tap so that peaks are controlled where the canceller can see them, and fix enclosure rattles mechanically. Suppressors can model some nonlinearity by giving more weight to harmonics of loud reference bands. Neural residual echo suppressors that take the reference and the linear stage's output as inputs are now common in conferencing products and were the focus of public AEC challenges; they handle distortion better than rule-based gains but need careful testing for near-end speech damage.

Speakerphones that also play audio into the same room as other open microphones can howl rather than echo; that closed-loop problem is handled separately by howling suppression.

Use the platform AEC first

Most applications should not ship their own canceller. Browsers, Android and Apple platforms provide one that already knows the playback path, and running a second canceller on top of the first usually makes things worse.

// Web: the browser's canceller uses the audio the browser itself plays as its reference
const stream = await navigator.mediaDevices.getUserMedia({
  audio: { echoCancellation: true, noiseSuppression: true, autoGainControl: true },
});
console.log(stream.getAudioTracks()[0].getSettings().echoCancellation);   // confirm it applied
// Android: the VOICE_COMMUNICATION source enables the platform's voice processing path;
// play the far end with USAGE_VOICE_COMMUNICATION so it is used as the reference.
AudioRecord rec = new AudioRecord.Builder()
    .setAudioSource(MediaRecorder.AudioSource.VOICE_COMMUNICATION)
    .setAudioFormat(new AudioFormat.Builder().setSampleRate(16000)
        .setEncoding(AudioFormat.ENCODING_PCM_16BIT)
        .setChannelMask(AudioFormat.CHANNEL_IN_MONO).build())
    .build();
if (AcousticEchoCanceler.isAvailable()) {
    AcousticEchoCanceler aec = AcousticEchoCanceler.create(rec.getAudioSessionId());
    if (aec != null) aec.setEnabled(true);
}
// iOS: session in .voiceChat mode on the speaker route (AVAudioSession is iOS-only)
let session = AVAudioSession.sharedInstance()
try session.setCategory(.playAndRecord, mode: .voiceChat, options: [.defaultToSpeaker])
// iOS and macOS: voice processing on the engine's input node
try engine.inputNode.setVoiceProcessingEnabled(true)

Platform cancellers have limits. A browser's canceller generally cancels only audio the browser plays, so a meeting played through a separate app is not in its reference. On Android, quality varies a lot between device models because the canceller is supplied by the device maker. External speakers and Bluetooth devices add delay and processing the platform may not account for. Test your real device and route combinations, and only bring your own canceller, such as the one in the WebRTC library, if measurements show the platform one failing on hardware you must support.

Worked example: measuring a speakerphone

Echo return loss enhancement (ERLE) is the standard measure of a linear canceller: the ratio, in decibels, of microphone power to error power during periods when only the far end is talking. Record a test session with a reference file played through the device: thirty seconds of far-end speech alone, then thirty seconds of double talk with a near-end talker reading at normal level, then a volume change and a move of the device.

import numpy as np

def erle_db(y, e, frame=320):                  # 20 ms frames at 16 kHz
    n = len(y) // frame
    py = (y[:n*frame].reshape(n, frame) ** 2).mean(axis=1)
    pe = (e[:n*frame].reshape(n, frame) ** 2).mean(axis=1)
    return 10 * np.log10((py + 1e-12) / (pe + 1e-12))

curve = erle_db(mic[:30*16000], err[:30*16000])  # far-end-only section
print("converged ERLE dB:", np.median(curve[-250:]))
print("seconds to reach 15 dB:", np.argmax(curve > 15) * 0.02)

Illustrative readings for a small speakerphone at moderate volume: echo at the microphone around -20 dBFS, linear-stage output around -42 dBFS, an ERLE of about 22 dB once converged, reached in one to two seconds. The suppressor should then bring residual echo below the room's noise floor. At maximum volume ERLE may fall below 15 dB because of distortion, which is exactly the regime where suppression has to work hardest and where capping volume pays off.

TestPass criterion to setWhat a failure points to
Far end only, steadyNo audible echo at the far end; ERLE stableReference tap or delay alignment
Double talkNear-end speech intelligible, no clipped word startsSuppressor too aggressive; double-talk control
Volume step upEcho burst shorter than about a secondNormalisation or step-size control
Device movedQuick reconvergence, no howlTracking speed; background filter
Maximum volumeResidual echo below the noise floorNonlinearity; volume cap or limiter
External USB or Bluetooth speakerSame as built-inDelay estimation range and clock drift

Failure modes in the field

SymptomLikely causeFix
Echo only at the start of each callFilter starts from zero and needs seconds to convergeStore converged coefficients per device and route
Echo whenever both people talkFilter adapts during double talk and divergesCoherence-based control or a two-filter structure
Near-end voice choppedResidual suppressor gains too lowRetune suppression against double-talk recordings
Echo with one model of speaker onlyUnmodelled delay or drift on that routeWiden the delay search; compensate drift
Echo grows over a long callClock drift between capture and playbackResample to a common clock
Distorted, buzzing echo at high volumeLoudspeaker nonlinearityLower the volume cap; limiter before the reference tap

What to do next

  1. Start with the platform's voice processing path on every target and verify in logs that it is actually enabled.
  2. Confirm where the reference is tapped and disable or move any nonlinear processing after it.
  3. Record the six test scenarios above on each device and route you support, and compute ERLE and double-talk quality.
  4. If you run your own canceller, size the filter from the measured room tail, use a frequency-domain or subband design, and add a background filter.
  5. Cap speaker volume at the onset of audible distortion and place a limiter before the reference tap.
  6. Tune the residual suppressor with double-talk recordings, not only far-end-only ones, and add comfort noise.
  7. Rerun the tests whenever the OS, audio route handling or hardware revision changes.
Key takeaway: A speakerphone is the hardest acoustic echo case because the echo can be louder than the talker, the echo tail is long, the loudspeaker distorts and the path keeps changing. The fixes are structural: tap the reference after the last nonlinear stage, align and drift-compensate it, size the linear filter to the room, control adaptation with level-independent measures or a two-filter design, and let a carefully tuned residual suppressor handle what remains. Use the platform's canceller first, measure ERLE and double-talk quality on real devices, and cap volume where distortion begins.