SHA-3 is not a patched SHA-2. It is a different design, Keccak by Bertoni, Daemen, Peeters and Van Assche, chosen by NIST in 2012 after a public competition and standardised as FIPS 202 in August 2015. Where SHA-256 compresses fixed blocks into a 256-bit chaining value, Keccak keeps a 1,600-bit state, XORs message blocks into part of it, and scrambles the whole state with a fixed permutation. The design has no additions, so no carries, and its only nonlinear step is a 5-bit AND-and-XOR map that is cheap in hardware and easy to mask against side channels.

The SHA family article covers why hashes need a new construction at all: Merkle-Damgard, length extension and the fall of SHA-1. This page goes inside Keccak. It covers the state layout, the five step mappings of Keccak-f[1600], how rate and capacity set security, the padding and domain bits that separate SHA3-256 from SHAKE and from Ethereum's Keccak-256, and a from-scratch Python implementation that matches hashlib byte for byte.

The state: lanes, rate and capacity

The state is 1,600 bits arranged as a 5×5 array of 64-bit lanes, A[x][y] with x and y from 0 to 4. Bit z of lane (x, y) is state bit 64(5y + x) + z, so when bytes are loaded the lanes fill row by row: lane (0,0), (1,0) through (4,0), then (0,1) and so on, each lane read as a little-endian 64-bit word. A 64-bit CPU therefore holds each lane in one register, and every step below is XORs, ANDs, NOTs and rotations on 25 words.

The state splits into the rate r, the first r bits that message blocks are XORed into and output is read from, and the capacity c = 1600 - r, which the outside world never touches directly. Capacity is the security parameter. A generic attack on the sponge costs about 2^(c/2) work, so FIPS 202 sets c to twice the digest length for the fixed-output functions.

FunctionRate r (bytes)Capacity c (bits)OutputCollision / preimage
SHA3-224144448224 bits112 / 224
SHA3-256136512256 bits128 / 256
SHA3-384104768384 bits192 / 384
SHA3-512721024512 bits256 / 512
SHAKE128168256any dmin(d/2, 128) / min(d, 128)
SHAKE256136512any dmin(d/2, 256) / min(d, 256)

The rate column explains the speed ranking. Each permutation call absorbs r bytes, so SHA3-512 does nearly twice as many permutations per message byte as SHA3-256. SHAKE128 is the fastest member, and SHA3-512's 512-bit preimage resistance is far more than most systems need.

One round of Keccak-f[1600]: θ, ρ, π, χ, ι

Keccak-f[1600] is 24 rounds. Each round applies five steps, named by Greek letters, and each step does one job.

  • θ (theta), diffusion across columns. Compute the parity of each of the five columns, C[x] = A[x][0] ⊕ … ⊕ A[x][4]. Then XOR into every lane of column x the value C[x-1] ⊕ rot(C[x+1], 1). One bit of input reaches 11 bits after θ.
  • ρ (rho), diffusion within lanes. Rotate each lane by a fixed offset. The 24 offsets are the triangular numbers (t+1)(t+2)/2 mod 64, assigned along the walk (x, y) → (y, 2x + 3y), and lane (0,0) is not rotated.
  • π (pi), moving lanes. Move lane (x, y) to position (y, 2x + 3y mod 5). Together with ρ, this breaks the column structure that θ relies on.
  • χ (chi), the only nonlinear step. For each row, A[x] ← A[x] ⊕ (¬A[x+1] ∧ A[x+2]). It is a degree-2 map on 5 bits, invertible, and costs one AND, one NOT and one XOR per lane.
  • ι (iota), breaking symmetry. XOR a round constant into lane (0,0). Without it every round would be identical and the permutation would map shift-symmetric states to shift-symmetric states. The constants come from an 8-bit LFSR and have bits set only at positions 0, 1, 3, 7, 15, 31 and 63.

The permutation is public and invertible, which is fine because security does not come from hiding f. It comes from the attacker's inability to control the c capacity bits. Twenty-four rounds is a deliberate margin: the best published attacks on the full hash reach far fewer rounds, which is why the Keccak team could also propose the 12-round KangarooTwelve for speed.

Padding and domain separation

Before absorbing, the message gets a few domain separation bits appended, then pad10*1: a 1 bit, as many 0 bits as needed, and a final 1 bit that ends exactly on a rate boundary. Because Keccak numbers bits from the least significant end of each byte, the domain bits and the first padding 1 fold into a single byte.

FunctionDomain bitsFirst pad byteLast pad byte
SHA3-224/256/384/512010x06|= 0x80
SHAKE128/25611110x1F|= 0x80
cSHAKE (SP 800-185)000x04|= 0x80
Original Keccak (Ethereum keccak256)none0x01|= 0x80

Two edge cases matter in tests. If the message is one byte short of a block, the first and last padding bytes coincide: SHA3-256 then writes 0x86 = 0x06 | 0x80 into the 136th byte. If the message fills a block exactly, padding needs a whole new block. Test at lengths r - 1, r and r + 1 and both cases are covered.

The last row is a real trap. Ethereum adopted Keccak before FIPS 202 added the domain bits, so its keccak256 is not SHA3-256. The empty-string hashes differ: SHA3-256("") starts a7ffc6f8 while Keccak-256("") starts c5d24601. Libraries that call their function "sha3" sometimes mean either one.

A from-scratch implementation

The whole hash fits in about 50 lines of Python. It is far too slow for production, but it is exact and readable, so it makes a good test oracle and a good way to see the steps. Lanes are stored as A[x][y].

The sponge: absorb r-byte blocks into a 200-byte state, then squeezemessage Mpad10*1+ domain bitsP0 P1 P2r-byte blocksKeccak-f[1600]⊕ P0Keccak-f[1600]⊕ P1Keccak-f[1600]⊕ P2rate rcapacity cabsorbingsqueezingZ0c bits never leavethe state: no length extensionSHA3-256: r = 1088 bits (136 bytes), c = 512. Output is the first d bits squeezed from the rate.
Absorbing XORs each padded r-byte block into the rate and applies Keccak-f[1600]; squeezing reads r bytes per permutation. The capacity is never output, which is why SHA-3 has no length-extension attack.
MASK = (1 << 64) - 1
def rol(v, n):
    n %= 64
    return ((v << n) | (v >> (64 - n))) & MASK if n else v

def rc_bit(t):                          # LFSR x^8 + x^6 + x^5 + x^4 + 1
    if t % 255 == 0:
        return 1
    r = 1
    for _ in range(t % 255):
        r <<= 1
        if r & 0x100:
            r ^= 0x171
    return r & 1

RC = [sum(rc_bit(j + 7 * i) << ((1 << j) - 1) for j in range(7)) for i in range(24)]
ROT = [[0] * 5 for _ in range(5)]
x, y = 1, 0
for t in range(24):
    ROT[x][y] = ((t + 1) * (t + 2) // 2) % 64
    x, y = y, (2 * x + 3 * y) % 5

def keccak_f(A):
    for rnd in range(24):
        C = [A[x][0] ^ A[x][1] ^ A[x][2] ^ A[x][3] ^ A[x][4] for x in range(5)]
        D = [C[(x - 1) % 5] ^ rol(C[(x + 1) % 5], 1) for x in range(5)]
        A = [[A[x][y] ^ D[x] for y in range(5)] for x in range(5)]          # theta
        B = [[0] * 5 for _ in range(5)]
        for x in range(5):
            for y in range(5):
                B[y][(2 * x + 3 * y) % 5] = rol(A[x][y], ROT[x][y])        # rho, pi
        A = [[B[x][y] ^ (~B[(x + 1) % 5][y] & B[(x + 2) % 5][y])
              for y in range(5)] for x in range(5)]                         # chi
        A[0][0] ^= RC[rnd]                                                  # iota
    return A

def sponge(msg, rate, suffix, out_len):
    p = bytearray(msg) + bytes([suffix])      # domain bits + first pad bit
    p += bytes(-len(p) % rate)
    p[-1] |= 0x80                             # final pad bit
    A = [[0] * 5 for _ in range(5)]
    for off in range(0, len(p), rate):        # absorb
        for i in range(rate // 8):
            A[i % 5][i // 5] ^= int.from_bytes(p[off + 8*i: off + 8*i + 8], "little")
        A = keccak_f(A)
    out = bytearray()
    while True:                               # squeeze
        for i in range(rate // 8):
            out += A[i % 5][i // 5].to_bytes(8, "little")
        if len(out) >= out_len:
            return bytes(out[:out_len])
        A = keccak_f(A)

sha3_256  = lambda m: sponge(m, 136, 0x06, 32)
shake128  = lambda m, n: sponge(m, 168, 0x1F, n)
keccak256 = lambda m: sponge(m, 136, 0x01, 32)

We checked it against hashlib.sha3_256, sha3_512 and shake_128 with 500 bytes of output, at message lengths 0, 1, 135, 136, 137, 300 and 1,000. All matched. The derived constants also match the published tables: RC[0] = 0x1, RC[1] = 0x8082 and RC[23] = 0x8000000080008008. As a worked value, SHA3-256("abc") is 3a985da74fe225b2045c172d6bd390bd855f086e3e9d525b46bfe24511431532.

SHAKE, cSHAKE and KMAC

The sponge naturally produces as much output as you ask for, so FIPS 202 also defines two extendable-output functions, SHAKE128 and SHAKE256. Use them for key derivation, for deterministic random streams such as sampling matrices in ML-KEM (Kyber) and ML-DSA (Dilithium), which use SHAKE internally, and for any digest of non-standard length. One subtlety: SHAKE128(M, 32) is a prefix of SHAKE128(M, 64). The output length is not an input, so never treat two lengths of the same XOF as independent keys. Put the length or a label into the message instead.

NIST SP 800-185 builds on this. cSHAKE adds a function name and a customisation string, so two protocols using the same message get unrelated outputs. KMAC is a MAC built directly on cSHAKE. Because the sponge does not leak its capacity, KMAC can simply absorb the key and then the message; it does not need HMAC's two-pass nesting. TupleHash hashes a sequence of strings unambiguously, and ParallelHash splits long inputs into independently hashed blocks for multicore speed.

Operational guidance

On commodity x86 servers SHA-256 with the SHA extensions is usually faster than SHA3-256 in software. SHA-3 only wins clearly where hardware support exists (ARMv8.2 includes optional SHA-3 instructions such as EOR3 and RAX1) or in hardware designs, where Keccak is small and fast. Measure on your own target before choosing for throughput.

  • Pick by property, not fashion. SHA-256 is fine and ubiquitous. Choose SHA-3 when you need resistance to length extension without HMAC, an XOF, domain separation via cSHAKE, or a second, unrelated design family as a hedge.
  • Name the exact function in formats. Write "SHA3-256 (FIPS 202)" or "Keccak-256 (pre-FIPS padding)" in specs and wire formats, never just "sha3".
  • Pin a known-answer test. Ship the empty-string and "abc" digests in your test suite so that swapping libraries cannot silently swap padding.
  • Passwords still need a slow KDF. SHA-3 is fast by design. For passwords, use a memory-hard or iterated function.

Failure modes

  • Keccak/SHA-3 confusion. Signatures and addresses computed with one fail to verify with the other. The symptom is an every-time mismatch on otherwise correct code.
  • Wrong lane order. Loading bytes big-endian, or filling lanes column-first, produces a hash that is self-consistent but wrong. Only a known-answer test catches it.
  • Missing the full-block padding case. Implementations that pad into the existing block when the message is exactly r bytes long pass short tests and fail at 136 bytes.
  • Truncating XOF output for separation. Using SHAKE output prefixes as distinct keys gives related keys. Use cSHAKE customisation strings instead.
  • Assuming a reduced-round variant is SHA-3. KangarooTwelve and TurboSHAKE are useful and well studied, but they are different functions with different outputs.

Trade-offs

SHA-3's strengths are structural: no length extension, a clean security argument from the capacity, native variable-length output, and a design unrelated to SHA-2, so a breakthrough against one is unlikely to hit the other. The costs are a 200-byte state instead of 32 bytes, slower software speed on CPUs without Keccak instructions, and a padding difference from early Keccak that keeps causing interoperability bugs. For a tree of hashes, such as a Merkle tree, or for hash-based signatures, the XOF and domain-separation features often matter more than raw speed.

What to do next

  1. Run the reference implementation above against your platform's library at lengths 0, r - 1, r and r + 1 for each variant you use.
  2. Grep your codebase for "sha3" and "keccak" and confirm, for each call site, which padding it uses and which one the other party expects.
  3. Replace any ad-hoc "hash(key || message)" MAC with KMAC or HMAC, and any home-made key separation with cSHAKE customisation strings.
  4. Benchmark SHA-256, SHA3-256 and SHAKE128 on your actual hardware before picking for throughput.
  5. For password storage, read the password hashing article instead of reaching for any fast hash.
Key takeaway: SHA-3 is a sponge built on Keccak-f[1600], a 24-round permutation of 25 64-bit lanes. θ and ρ/π diffuse, χ is the only nonlinear step and ι breaks symmetry. Message blocks are XORed into the r-bit rate. The c-bit capacity is never exposed, which gives about c/2 bits of security and rules out length extension. The padding byte decides the function: 0x06 for SHA3, 0x1F for SHAKE and 0x01 for Ethereum's Keccak-256. Pin known-answer tests, name the exact variant, and use SHAKE, cSHAKE and KMAC where they remove ad-hoc constructions.