Pulse-code modulation (PCM) is the plain, uncompressed representation of sound inside every computer: a list of numbers, each one the amplitude of the signal at an instant, taken at a fixed rate. Codecs such as Opus and AAC compress it, speech recognisers consume it, mixers and effects transform it, and sound cards turn it back into voltage. Because PCM is so simple, it is easy to assume there is nothing to get wrong. In practice most audio bugs, from chipmunk voices and white-noise bursts to inaudible-until-mastering distortion, come from two components disagreeing about one PCM detail.

This article builds PCM up from sampling theory to bytes on disk. It explains what sample rate and bit depth buy you, how samples are laid out in memory, how a WAV file really works, how to convert between integer and floating-point samples without off-by-one distortion, and when and how to dither. It finishes with failure modes, operational defaults and a checklist.

Advertisement

From pressure wave to numbers

Microphonecontinuous voltageAnti-alias filterremove above fs / 2ADCsample at fs, quantize to N bitsPCM streamframes: L R L R ...Container / bufferWAV chunks, ring bufferConvert to float32divide by 2^(N-1)Processgain, mix, resample, modelDither + requantizeonce, at the outputDACreconstruction filterSpeakercontinuous againPCM is the uncompressed middle of every audio system: numbers at a fixed rate, a fixed width and a fixed channel order.Most bugs come from disagreeing about one of those three facts, not from the signal processing.
A PCM pipeline: the converter samples and quantizes, software works on fixed-rate frames (usually in float), and the signal is requantized with dither once, on the way out.

Two operations turn a continuous signal into PCM. Sampling measures the signal every 1 / fs seconds, where fs is the sample rate. The sampling theorem says a signal containing no frequencies at or above fs / 2, the Nyquist frequency, can be reconstructed exactly from its samples. Anything above fs / 2 does not disappear: it folds back, or aliases, to a false lower frequency. That is why converters put a low-pass anti-aliasing filter before sampling, and why software must low-pass before reducing a sample rate; see audio resampling.

Quantization rounds each measured amplitude to one of 2N levels for an N-bit sample. The rounding error behaves, for busy signals, like a small added noise.

Common rates follow use. 44.1 kHz came from the CD and covers hearing up to about 20 kHz with room for a filter transition. 48 kHz is the standard for video, broadcast and most operating-system audio stacks. 16 kHz is the common rate for speech models and wideband telephony, and 8 kHz for narrowband telephony. Higher rates (96 and 192 kHz) give processing headroom in production but add little for playback.

Vocabulary matters because APIs mix it up. A sample is one number for one channel. A frame is one sample for every channel at the same instant, so one second of 48 kHz stereo is 48,000 frames and 96,000 samples. Buffer sizes in audio APIs are usually stated in frames; byte counts are frames x channels x bytes per sample.

Bit depth and dynamic range

Each extra bit halves the quantization step and so improves the signal-to-noise ratio by about 6 dB. For a full-scale sine wave and an ideal N-bit quantizer, SNR = 6.02 N + 1.76 dB. Sixteen bits gives about 98 dB, which comfortably covers a quiet room to loud playback. Twenty-four bits gives a theoretical 146 dB, more than any real converter or microphone achieves; its value is headroom, so recordings can peak well below full scale without losing the quiet detail.

32-bit floating point is the usual working format in software. Its 24-bit significand gives roughly 24-bit precision at every level, and its exponent means values above 1.0 do not clip while you process. Full scale is defined as 1.0 by convention; only the final conversion to integers has a hard ceiling.

FormatRangeBytes per sampleTypical use
Unsigned 8-bit0 to 255, silence at 1281Legacy WAV, old games
Signed 16-bit (s16le)-32,768 to 32,7672CD, telephony, most files and devices
Signed 24-bit-8,388,608 to 8,388,6073 packed, or 4 in a 32-bit containerRecording, production
Signed 32-bit-231 to 231 - 14Device interfaces, some DAWs
Float 32-bit (f32le)Nominal -1.0 to 1.04Processing, ML input, browser and OS mixers

The data rate is fs x channels x bits. CD audio is 44,100 x 2 x 16 = 1,411,200 bits per second, 176,400 bytes per second, about 635 MB per hour. 48 kHz stereo in float32 is 384,000 bytes per second. These numbers set disk, network and memory budgets before any compression.

Advertisement

Layout in memory: interleaving, endianness and buffers

Multichannel PCM is stored either interleaved, frame by frame (L R L R ...), or planar, with one contiguous array per channel. Files and most device APIs use interleaved; many DSP libraries and ML frameworks prefer planar because each channel is contiguous for vector processing. Converting between them is a transpose; forgetting to is the cause of the classic bug where stereo plays at half speed in one ear or turns into buzzing.

Channel order is a convention, not a law. WAV and most PC APIs order 5.1 as front left, front right, centre, LFE, back left, back right, while some film and codec formats differ. Always carry an explicit channel layout with multichannel audio.

Endianness applies to every multi-byte sample. WAV is little-endian; AIFF is big-endian; network and raw streams must say. Reading big-endian samples as little-endian produces loud, harsh noise, which at least makes the bug obvious. 24-bit samples are the other trap: packed 3-byte samples and 24 bits left- or right-justified in 4-byte containers all exist, and a mismatch gives either noise or a signal 48 dB too quiet.

Real-time systems move PCM in buffers of a fixed number of frames, and buffer size is latency: 256 frames at 48 kHz is 5.33 ms per buffer. Smaller buffers lower latency and raise the risk of underruns; buffer management for low latency covers that trade-off, and clock drift compensation covers what happens when two devices' sample clocks disagree slightly.

The WAV file, done properly

A WAV file is a RIFF container: the 12-byte header "RIFF", a 32-bit little-endian size, and "WAVE", followed by chunks. Each chunk has a four-character id, a 32-bit size and a body padded to an even length. Two chunks matter. The fmt chunk holds the format tag, channel count, sample rate, byte rate, block align (bytes per frame) and bits per sample. The data chunk holds the interleaved samples.

Format tag 1 is integer PCM, 3 is IEEE float, and 0xFFFE is WAVE_FORMAT_EXTENSIBLE, which adds a channel mask, valid bits per sample and a sub-format GUID whose first two bytes carry the real format code. Microsoft's guidance is to use the extensible form for more than two channels or more than 16 bits, so readers that only accept tag 1 reject many legitimate files. Eight-bit WAV samples are unsigned; every wider integer width is signed.

Four habits avoid most WAV bugs. Do not assume the canonical 44-byte header: LIST, fact, bext and other chunks can appear before data, so walk the chunks. Trust block align and the actual data length over the size fields, because streaming writers often leave sizes at zero or at 0xFFFFFFFF when they cannot seek back. Remember that 32-bit sizes cap a RIFF file at 4 GiB; longer recordings use RF64 or Wave64. And drop a trailing partial frame instead of misaligning every channel after it.

import struct
import numpy as np

def read_wav(path):
    """Walk RIFF chunks instead of assuming a 44-byte header."""
    with open(path, "rb") as f:
        riff, size, wave = struct.unpack("<4sI4s", f.read(12))
        if riff != b"RIFF" or wave != b"WAVE":
            raise ValueError("not a RIFF/WAVE file (RF64 and W64 need their own readers)")
        fmt = data = None
        while True:
            hdr = f.read(8)
            if len(hdr) < 8:
                break
            cid, clen = struct.unpack("<4sI", hdr)
            body = f.read(clen)
            if clen % 2:
                f.read(1)                              # chunks are padded to even length
            if cid == b"fmt ":
                tag, ch, rate, byte_rate, align, bits = struct.unpack("<HHIIHH", body[:16])
                if tag == 0xFFFE:                      # WAVE_FORMAT_EXTENSIBLE
                    tag = struct.unpack("<H", body[24:26])[0]   # first two bytes of SubFormat GUID
                fmt = (tag, ch, rate, align, bits)
            elif cid == b"data":
                data = body
        if fmt is None or data is None:
            raise ValueError("missing fmt or data chunk")
    tag, ch, rate, align, bits = fmt
    n = len(data) // align                             # drop a trailing partial frame
    data = data[: n * align]
    if tag == 3 and bits == 32:                        # IEEE float
        x = np.frombuffer(data, "<f4").astype(np.float32)
    elif tag == 1 and bits == 8:                       # 8-bit WAV is UNSIGNED
        x = (np.frombuffer(data, np.uint8).astype(np.float32) - 128) / 128
    elif tag == 1 and bits == 16:
        x = np.frombuffer(data, "<i2").astype(np.float32) / 32768
    elif tag == 1 and bits == 24:                      # packed 3-byte little-endian
        b = np.frombuffer(data, np.uint8).reshape(-1, 3).astype(np.int32)
        v = b[:, 0] | (b[:, 1] << 8) | (b[:, 2] << 16)
        v = np.where(v >= 1 << 23, v - (1 << 24), v)   # sign-extend
        x = v.astype(np.float32) / (1 << 23)
    elif tag == 1 and bits == 32:
        x = np.frombuffer(data, "<i4").astype(np.float32) / 2**31
    else:
        raise ValueError(f"unsupported format tag {tag} with {bits} bits")
    return x.reshape(-1, ch), rate                     # (frames, channels), interleaved source

def to_int16(x, rng=np.random.default_rng()):
    """Float [-1, 1) to int16 with TPDF dither and explicit clipping."""
    lsb = 1 / 32768
    tpdf = (rng.random(x.shape) - rng.random(x.shape)) * lsb   # triangular, +/- 1 LSB
    y = np.round((x + tpdf) * 32768)
    clipped = np.count_nonzero((y > 32767) | (y < -32768))
    if clipped:
        print(f"warning: {clipped} samples clipped")
    return np.clip(y, -32768, 32767).astype("<i2")

Converting between integer and float

Signed integers are asymmetric: 16-bit runs from -32,768 to 32,767. The most common convention, and the one the reader above uses, divides by 32,768, so integers map to [-1.0, 1.0) and -1.0 is reachable but +1.0 is not. Dividing by 32,767 instead makes positive full scale exactly 1.0 and pushes -32,768 slightly below -1.0. Neither is wrong; mixing them is. A library that reads with one and writes with the other introduces a tiny gain change on every round trip, and a few round trips through a pipeline can produce clipping or fail bit-exactness tests.

Going back to integers needs three steps in order: scale, round (with dither if the result is final), then clip explicitly to the integer range. Casting a float above 1.0 straight to int16 in C or NumPy wraps around rather than clipping, turning a slight overshoot into a full-scale click. Count clipped samples and log them; silent clipping is a quality bug that nobody hears in testing.

For ML pipelines, decide the convention once in the data loader. Speech models usually expect float32 mono at a fixed rate such as 16 kHz, scaled to [-1, 1). A model trained on audio divided by 32,768 and served audio that skipped the division sees input 32,768 times too loud and fails quietly with confident nonsense.

Gain, headroom and dither

Processing in float lets intermediate values exceed full scale, but the output still must fit. Leave headroom: mixing two full-scale signals can double the peak. Apply gain and limiting deliberately rather than relying on clipping; audio mixing covers summing and limiters.

Reducing bit depth, for example from a float mix to a 16-bit file, is requantization. On loud, busy signals the error is noise-like and harmless. On quiet or slowly fading signals the error correlates with the signal and becomes harmonic distortion and gritty "truncation" artifacts in reverb tails and fade-outs. Dither fixes this by adding a tiny random signal before rounding, so the error becomes independent of the signal. The standard choice is TPDF (triangular probability density) dither: the difference of two uniform random values, spanning plus or minus one least-significant bit. It raises the noise floor by about 4.8 dB compared with undithered rounding but removes the distortion entirely. Noise shaping goes further by pushing that noise towards frequencies where hearing is less sensitive.

Dither once, at the final reduction, and never before further processing or a lossy codec that will requantize anyway. Do not dither when converting losslessly (16-bit to float and back without changes), because the round trip should be bit-exact.

Failure modes and how to spot them

SymptomLikely causeCheck
Loud static instead of soundWrong endianness, or float read as integerFormat tag, byte order
Chipmunk or slowed voiceWrong sample rate or channel countfmt chunk versus what the player assumed
Buzz, one channel garbageInterleaved read as planar, or misaligned frameBlock align, trailing partial frame
DC offset, distortion on 8-bitUnsigned 8-bit treated as signedSubtract 128
Audio 48 dB too quiet24-bit in 32-bit container misreadJustification, valid bits
Clicks at loud passagesFloat to int cast without clippingClip count in the converter
Gritty fade-outsTruncation without ditherRequantization step
Slowly growing latency or gapsTwo clocks at slightly different ratesDrift compensation

Operational guidance

Pick one internal format and convert only at the edges: float32, interleaved or planar by library preference, at one fixed rate (48 kHz for media, 16 kHz for speech models). Carry sample rate, channel count and channel layout with every buffer or file rather than inferring them. Resample once, with a good filter, as close to the source as possible. Log format details at every boundary so a mismatch shows up in logs, not complaints. Keep a test corpus of awkward files: 8-bit, 24-bit packed, extensible, multichannel, extra chunks, zero-size headers, and odd lengths. For loudness targets rather than peak levels, see loudness normalization.

What to do next

  1. Write down your pipeline's internal format: sample type, rate, channel layout, interleaving and float scaling convention.
  2. Replace any fixed 44-byte WAV parsing with a chunk walker, and support tags 1, 3 and 0xFFFE.
  3. Make every float-to-integer conversion clip explicitly and count clipped samples.
  4. Add TPDF dither at the single final bit-depth reduction, and verify lossless round trips are bit-exact.
  5. Build a corpus of awkward test files and run every reader and writer against it in CI.
  6. Measure buffer latency in frames and milliseconds, and check for drift between capture and playback clocks.
Key takeaway: PCM is just numbers at a fixed rate, width and channel order, and nearly every PCM bug is two components disagreeing about one of those facts. Work in float32 at one rate, convert only at the edges with explicit scaling and clipping, walk WAV chunks instead of trusting a fixed header, dither once at the final reduction, and test against awkward files.