Fixed-point arithmetic stores a real number as an ordinary integer with an implied scale. The integer 3277 can mean 0.1 if everyone agrees that the scale is 1/32768. Additions are integer additions, multiplications are integer multiplications followed by a shift, and the hardware never sees a floating-point unit. That was the only option on early DSPs and microcontrollers, and it is still the right option whenever you need bit-exact results across machines, predictable latency, or cheap silicon: audio codecs, motor control, financial ledgers, deterministic game simulations, and the int8 kernels that run quantized neural networks.

This article builds fixed-point from first principles: the Q notation, conversion and its error, the double-width product, rounding and the negative-number trap, saturation, division, and the integer-only requantization step used by int8 inference runtimes. Every code fragment was run before publication, and the worked example uses real numbers you can reproduce.

Q formats from first principles

Q notation names a format by its bit split. In Qm.n, a signed integer has m integer bits, n fraction bits and a sign bit, and the represented value is the integer divided by 2^n. Two formats cover most uses. Q15 (also written Q1.15 or Q0.15 depending on whether the sign bit is counted) stores values in an int16 with 15 fraction bits, covering -1 to 1 - 2^-15 in steps of about 0.0000305. Q16.16 stores values in an int32 with 16 integer and 16 fraction bits, covering about -32768 to 32768 in steps of about 0.0000153.

The trade is fixed: every bit you give to the fraction is a bit taken from the range. Floating point moves that trade per value with an exponent; fixed point decides it once per variable at design time. That is why fixed-point design starts with knowing the range of every signal: the maximum amplitude of an audio sample, the largest balance in a ledger, the largest activation in a layer.

Converting and the error budget

Converting from a real number multiplies by 2^n, rounds and clamps. Converting back divides. With round-to-nearest, each conversion is off by at most half a step, 2^-(n+1). For 0.1 in Q15 the stored integer is 3277, which decodes to 0.100006: an error of about 6.1 x 10^-6, inside the half-step bound of 1.5 x 10^-5.

Q = 15
QMIN, QMAX = -(1 << 15), (1 << 15) - 1

def sat16(v):
    return max(QMIN, min(QMAX, v))

def to_q15(x):
    return sat16(round(x * (1 << Q)))

def from_q15(v):
    return v / (1 << Q)

to_q15(0.1)              # 3277
from_q15(3277) - 0.1     # 6.1e-06

Errors add up through a computation. A chain of k rounded operations can drift by up to k half steps in the worst case, and by roughly the square root of k half steps when errors are independent. Budget for this when you choose n: if a filter runs 64 multiply-adds per output, carry extra fraction bits in the accumulator and round once at the end.

Multiplication needs double width

Adding two numbers in the same format is plain integer addition. Multiplying is where formats change. Multiply a Qa.b value by a Qc.d value and the raw product is in Q(a+c).(b+d): two Q15 numbers give a Q30 product that needs 32 bits. To return to Q15 you shift right by 15. Shifting alone truncates, which rounds toward minus infinity; adding half a step first rounds to nearest.

#include <stdint.h>

static inline int16_t sat16(int32_t v) {
    return v > INT16_MAX ? INT16_MAX : v < INT16_MIN ? INT16_MIN : (int16_t)v;
}

/* Q15 x Q15 -> Q15, round half up, saturating */
static inline int16_t q15_mul(int16_t a, int16_t b) {
    int32_t prod = (int32_t)a * (int32_t)b;   /* Q30 */
    prod += 1 << 14;                            /* half of the 2^15 step */
    return sat16(prod >> 15);                   /* arithmetic shift assumed */
}

/* Q16.16 needs a 64-bit product */
static inline int32_t q16_mul(int32_t a, int32_t b) {
    int64_t prod = (int64_t)a * (int64_t)b;   /* Q32 */
    return (int32_t)((prod + (1 << 15)) >> 16);
}
Q15 multiply: two 16-bit values, a 32-bit product, then round, shift and saturatea: int16, Q151 sign + 15 fractionb: int16, Q15range -1 to 1 - 2^-15a x b: int32Q30, 30 fraction bits+ 2^14round to nearestshift right 15back to Q15saturateclamp to int16result: int16Q15Only -1 x -1 overflows: the true result +1.0 does not fit in Q15, so it saturates to 0x7FFF.Wider formats follow the same path: Q16.16 needs a 64-bit product before the shift.
Figure 1: the Q15 multiply data path. The product is computed at double width, rounded, shifted back and saturated.

The only Q15 product that overflows is -1 x -1. The true answer, +1.0, is one step above the largest Q15 value, so saturation clamps it to 0x7FFF. The Q16.16 version omits saturation for brevity; in real code, check the 64-bit value against the int32 range before the cast.

Rounding and the negative-number trap

Adding half a step and shifting right rounds ties toward plus infinity: 2.5 steps become 3, but -2.5 steps become -2. Over many operations that is a small positive bias, which matters in long-running accumulators such as integrators in control loops. Round-half-to-even removes the bias at the cost of a few more instructions. Decide which one your reference model uses, and use the same one everywhere.

The second trap is language semantics on negative numbers. In Python, >> and // both round toward minus infinity, so -7 >> 1 is -4. In C, integer division truncates toward zero, so -7 / 2 is -3. Right-shifting a negative signed integer in C is implementation-defined; mainstream compilers shift arithmetically, and C++20 made arithmetic shift the defined behaviour. A Python model of a C kernel that mixes // and / will disagree on exactly the negative inputs, and the mismatch will look like a rare rounding bug. Write the reference model with the same operator, or use an explicit truncating division helper.

Saturation versus wrap-around

When a fixed-point result leaves the representable range, integer hardware wraps: one step past the maximum becomes the minimum. In audio, a wrapped sample is a full-scale click; in control, it is a motor command that flips sign. Saturating arithmetic clamps to the nearest limit instead, which degrades gracefully. DSP instruction sets and SIMD extensions provide saturating add and multiply instructions, and compilers expose them as intrinsics.

Saturation is not free of trade-offs. It hides overflow, so a saturating pipeline can quietly produce clipped output for weeks. Count saturation events in debug builds and in production telemetry, and treat a non-zero count as a range design bug rather than normal operation. DSPs also provide guard bits, an accumulator wider than the operands, for exactly this reason: intermediate sums may exceed the range as long as the final result does not.

Division and other functions

Division of a Qn value by another Qn value needs the dividend pre-shifted: compute (a << n) / b in double width so the quotient keeps n fraction bits. Division is slow on small cores, so fixed-point code avoids it: divide by a constant by multiplying with its reciprocal in fixed point, normalise a value to a known range and refine a reciprocal with a few Newton iterations, or use a lookup table with interpolation. Square roots, sines and logarithms are handled the same way, or with CORDIC, an iterative shift-and-add method designed for hardware without multipliers.

Fixed point inside int8 inference

Quantized neural networks are fixed-point arithmetic with a per-tensor or per-channel scale. A real value r is stored as an int8 q with r = s x (q - z), where s is the scale and z the zero point. A matrix multiply accumulates int8 x int8 products into int32. Each product is at most 16384 in magnitude, so an int32 accumulator has room for about 2^17 products before overflow is possible, which comfortably covers normal layer widths.

The accumulator is in scale s_in x s_w, and the output must be in s_out, so it is multiplied by M = s_in x s_w / s_out, a real number usually below 1. Integer-only runtimes in the gemmlowp lineage, including TensorFlow Lite's reference kernels, store M as an int32 mantissa M0 in Q31 and a right-shift exponent. At run time they apply SaturatingRoundingDoublingHighMul (the high 32 bits of the doubled product, rounded) and then RoundingDivideByPOT (a rounding right shift), add the output zero point and clamp to int8.

Integer-only requantization in an int8 layerint8 x int8inputs and weightsint32 accumulatorplus int32 biasx M0, high halfQ31 multiplierrounding shiftby the exponent+ zero point, clampto int8offline: M = s_in x s_w / s_outstored as M0 and shiftNo floating point at inference time: the float scale becomes a fixed-point multiplier.
Figure 2: requantization turns a float scale into a Q31 multiplier and a shift, so the whole layer runs on integers.
import math

def quantize_multiplier(real):
    """0 < real < 1  ->  (m0, shift) with real ~= m0 / 2**31 / 2**shift."""
    mant, exp = math.frexp(real)           # real = mant * 2**exp, mant in [0.5, 1)
    m0 = round(mant * (1 << 31))
    if m0 == 1 << 31:
        m0, exp = m0 // 2, exp + 1
    return m0, -exp

def srdhm(a, b):                           # SaturatingRoundingDoublingHighMul
    if a == b == -(1 << 31):
        return (1 << 31) - 1
    ab = a * b
    nudge = (1 << 30) if ab >= 0 else 1 - (1 << 30)
    x = ab + nudge
    return abs(x) // (1 << 31) * (1 if x >= 0 else -1)   # C-style truncation

def rounding_divide_by_pot(x, e):          # RoundingDivideByPOT
    mask = (1 << e) - 1
    threshold = (mask >> 1) + (1 if x < 0 else 0)
    return (x >> e) + (1 if (x & mask) > threshold else 0)

def requantize(acc, m0, shift, zero_point):
    v = rounding_divide_by_pot(srdhm(acc, m0), shift) + zero_point
    return max(-128, min(127, v))

Note the explicit truncation inside srdhm: it mirrors the C kernel, which is the point made in the rounding section. For the design side of choosing scales, see quantization in depth and int8 versus int4.

Worked example: requantizing one layer

Take a layer with input scale 0.02, weight scale 0.005 and output scale 0.08, and output zero point 3. The real multiplier is 0.02 x 0.005 / 0.08 = 0.00125. frexp gives 0.00125 = 0.64 x 2^-9, so M0 = round(0.64 x 2^31) = round(1374389534.72) = 1374389535 and the shift is 9. Decoding them gives 0.0012500000003, an error in the eleventh significant digit.

An accumulator of 12345 should map to 12345 x 0.00125 = 15.43, rounded to 15, plus the zero point: 18. The integer pipeline produces 18. An accumulator of -12345 maps to -15.43, rounded to -15, plus 3: -12, and the integer pipeline produces -12. Testing the integer path against a float reference over every accumulator value in a realistic range, and requiring an exact match or at most one step of difference on exact ties, is how runtimes keep kernels honest.

Money and determinism

Money is fixed point with a decimal scale. Storing cents, or micro-units, as integers makes addition exact, which binary floating point cannot do: 0.1 + 0.2 is not 0.3 in IEEE doubles. Pick the scale from the smallest unit any rule needs, such as tax computed to a tenth of a cent, and pick the rounding rule from the regulation, often half-to-even for banking. SQL NUMERIC(p, s) columns are true decimal fixed point. Python's decimal module is decimal floating point, and behaves as fixed point only when you quantize() every result to a set number of places.

Determinism is the other reason to choose fixed point. Floating-point results can differ across compilers, instruction sets and fused multiply-add choices, which breaks lockstep multiplayer games and replicated state machines. Integer arithmetic gives the same bits everywhere, so simulations and ledgers that must agree across machines often use scaled integers on purpose.

Failure modes

  • Single-width products. Multiplying two Q16.16 values in 32 bits overflows as soon as the real product exceeds 0.5 in magnitude, because the raw product carries 32 fraction bits; always widen first.
  • Wrong reference semantics. Python floor division versus C truncation produces mismatches only on negative values.
  • Silent saturation. Clipping hides a range bug; count saturation events.
  • Mixed formats. Adding a Q15 to a Q12 value without aligning them is off by a factor of eight; encode formats in type names.
  • Biased rounding in integrators. Round-half-up drifts over millions of steps.

Trade-offs

ChoiceGainCost
Fixed point over floatBit-exact, cheap silicon, predictable timingRange analysis per variable
More fraction bitsSmaller rounding errorLess headroom, more overflow risk
Saturating opsGraceful overloadHidden range bugs
Round half to evenNo biasA few extra instructions
Wide accumulatorFewer roundings, guard bitsWider registers, more memory

Fixed point is also the basis of secure multi-party protocols that compute on shares of real values, where truncation after multiplication is a protocol step; see encrypted inference.

What to do next

  1. Write down the range and required precision of every variable before choosing a format.
  2. Name formats in types or variable suffixes, such as q15 and q16, and align before adding.
  3. Widen every product, round once, then shift and saturate.
  4. Make the reference model use the same rounding and division semantics as the target.
  5. Count saturation events in debug builds and in production telemetry.
  6. For int8 kernels, test requantization exhaustively against a float reference over the realistic accumulator range.
Key takeaway: Fixed point is integer arithmetic with an agreed scale. Choose formats from measured ranges, widen every product, round once, saturate visibly, and make your reference model match the target's rounding and division semantics. The same discipline powers DSP code, money and int8 inference.