Almost every modern protocol ends a key exchange holding a secret that is not yet a key. An X25519 shared secret is 32 bytes, but it is a point coordinate, not 32 uniformly random bytes. A pre-shared key may be reused across many sessions. A hardware seed may be biased. And even a perfectly uniform secret is only one value, while a session needs several: an encryption key per direction, a MAC key, nonce bases, perhaps a key for resumption. HKDF, the HMAC-based Extract-and-Expand Key Derivation Function of Krawczyk and Eronen (RFC 5869, 2010), is the standard answer to both problems.
This article builds HKDF from first principles: what a key derivation function must guarantee, why HKDF splits the job into two steps, a tested implementation that reproduces the RFC test vectors, how TLS 1.3 and HPKE wrap it, and the mistakes that quietly break it. The one rule to carry away: HKDF turns entropy into keys; it cannot create entropy, so it is the wrong tool for passwords.
Why a raw secret is not a key
A cryptographic key is assumed to be uniformly random. Ciphers and MACs are proven secure under that assumption, and nothing else. Real secrets fall short in two ways. They can be non-uniform: a Diffie-Hellman output has plenty of entropy, but it is structured, and some bit patterns are more likely than others. And they can be the wrong size or count: you hold one secret and need four keys of different lengths.
The tempting shortcut is key = SHA256(secret), then SHA256(secret + 'enc') for the next key. It works often enough to survive code review, but it has no proof behind it, it mixes up the two jobs, and it leads to ad hoc encodings that collide ('a' + 'bc' equals 'ab' + 'c'). Krawczyk's insight was to separate the jobs: first an extractor that concentrates whatever entropy the input has into one short, pseudorandom key (PRK); then a pseudorandom function in expansion mode that stretches the PRK into as many independent keys as you like, each bound to a label. Each step has its own security argument, and HMAC serves as both.
Extract, then expand
With HashLen the hash output size (32 for SHA-256), RFC 5869 defines:
HKDF-Extract(salt, IKM) -> PRK
if salt is empty: salt = HashLen zero bytes
PRK = HMAC-Hash(key = salt, msg = IKM)
HKDF-Expand(PRK, info, L) -> OKM # requires L <= 255 * HashLen
N = ceil(L / HashLen)
T(0) = empty string
T(i) = HMAC-Hash(key = PRK, msg = T(i-1) || info || byte(i)) for i = 1..N
OKM = first L bytes of T(1) || T(2) || ... || T(N)Notice the role swap. In Extract the salt is the HMAC key and the secret is the message; in Expand the PRK is the key. The counter is a single byte, so Expand can produce at most 255 blocks: 8,160 bytes for SHA-256 and 16,320 for SHA-512. Asking for more is an error, not a wrap-around.
A tested implementation
The whole construction fits in a few lines of standard-library Python. This version was run against the RFC 5869 test vectors and against the cryptography package's HKDF class, and matched both.
import hmac, hashlib
def hkdf_extract(salt: bytes, ikm: bytes, hash_name="sha256") -> bytes:
if not salt:
salt = bytes(hashlib.new(hash_name).digest_size)
return hmac.new(salt, ikm, hash_name).digest()
def hkdf_expand(prk: bytes, info: bytes, length: int, hash_name="sha256") -> bytes:
hash_len = hashlib.new(hash_name).digest_size
if len(prk) < hash_len:
raise ValueError("PRK shorter than HashLen")
if length > 255 * hash_len:
raise ValueError("length exceeds 255 * HashLen")
okm, t = b"", b""
for i in range(1, -(-length // hash_len) + 1):
t = hmac.new(prk, t + info + bytes([i]), hash_name).digest()
okm += t
return okm[:length]
def hkdf(ikm, salt, info, length, hash_name="sha256"):
return hkdf_expand(hkdf_extract(salt, ikm, hash_name), info, length, hash_name)In production, call a vetted library rather than this code. The point of writing it is to see that there is nothing hidden: two HMAC calls in a loop, and two length checks that the RFC requires.
Worked example: RFC 5869 test case 1
RFC 5869 test case 1 uses SHA-256 with an IKM of 22 bytes of 0x0b, a 13-byte salt 0x000102...0c, a 10-byte info 0xf0f1...f9 and L = 42. Running the code above gives:
PRK = 077709362c2e32df0ddc3f0dc47bba6390b6c73bb50f9c3122ec844ad7c2b3e5
OKM = 3cb25f25faacd57a90434f64d0362f2a2d2d0a90cf1a5a4c5db02d56ecc4c5bf
34007208d5b887185865Both values match the RFC. Trace the work: Extract made one HMAC call and produced a 32-byte PRK. Expand needed ceil(42 / 32) = 2 blocks: T(1) is HMAC(PRK, info || 0x01), T(2) is HMAC(PRK, T(1) || info || 0x02), and the output is the 32 bytes of T(1) plus the first 10 of T(2). Test case 3 repeats this with an empty salt and empty info; the empty salt becomes 32 zero bytes, and the code reproduces that vector too (8da4e775a563c18f...).
Two more runs show properties worth knowing. Deriving 16 bytes and 64 bytes with the same inputs gives outputs where the 16-byte key is exactly the first 16 bytes of the 64-byte one. And changing only info, from app v1 enc to app v1 mac, gave unrelated outputs (013f6c58... versus 4447e3c1...). That second fact is the whole point of info; the first is a trap covered below.
Salt, info and domain separation
Salt is a non-secret, ideally random value that turns HMAC into a good extractor. It is optional, and an all-zero salt still works under stronger assumptions about HMAC, but a random salt gives the better proof and also separates the outputs of two systems that happen to share the same IKM. It can be public, sent in the clear, and reused across many derivations. What it should not be is chosen by an attacker after seeing the secret.
Info is the context string, and it is your domain separation. Put everything that distinguishes one derived key from another into it: the protocol name and version, the key's role, the direction (client-to-server or the reverse), the algorithm it will feed, and the output length. Encode it unambiguously: fixed fields or length-prefixed strings, never bare concatenation of variable-length values.
The prefix property. Because OKM is a prefix of a longer output, length alone does not separate keys. If one code path derives a 16-byte AES-128 key and another derives a 32-byte key with the same info, the AES-128 key is half of the other key. Including the length in info, as TLS 1.3 does, closes this.
Skipping Extract. RFC 5869 allows calling Expand directly when the input is already a uniformly random key of at least HashLen bytes, for example a key fresh from a CSPRNG or the PRK of an earlier step. Do not skip it for Diffie-Hellman outputs, passwords or anything with structure.
HKDF inside TLS 1.3 and HPKE
TLS 1.3 (RFC 8446) runs its whole key schedule on HKDF. It chains Extract calls: an early secret from the PSK (or zeros), a handshake secret from the (EC)DHE output, then a master secret. Between them it calls a wrapper, HKDF-Expand-Label(Secret, Label, Context, Length), whose info is a structure holding the 2-byte output length, the label prefixed with "tls13 ", and a context, usually a transcript hash. Derive-Secret is that wrapper with the transcript hash as context. Length, protocol and purpose all land in info, which is exactly the discipline described above.
HPKE (RFC 9180) defines LabeledExtract and LabeledExpand, which prefix every input with "HPKE-v1", a suite identifier and a label, and prefix expand info with the 2-byte length. The Signal protocol's ratchets, Noise handshakes and many storage-encryption schemes use HKDF in the same role. The TLS 1.3 internals article shows where each derived secret is used on the wire.
Library APIs
Every mainstream platform ships HKDF. Use these rather than hand-rolled code:
# Python (cryptography): instances are single-use
from cryptography.hazmat.primitives import hashes
from cryptography.hazmat.primitives.kdf.hkdf import HKDF, HKDFExpand
key = HKDF(algorithm=hashes.SHA256(), length=32,
salt=salt, info=b"myapp v2 | aes-256-gcm | c2s").derive(shared_secret)
// Go 1.24+: standard library crypto/hkdf
key, err := hkdf.Key(sha256.New, secret, salt, "myapp v2 | aes-256-gcm | c2s", 32)
// Node.js
const key = crypto.hkdfSync("sha256", secret, salt, info, 32); // ArrayBuffer
// Browser WebCrypto
const base = await crypto.subtle.importKey("raw", secret, "HKDF", false, ["deriveBits"]);
const bits = await crypto.subtle.deriveBits(
{ name: "HKDF", hash: "SHA-256", salt, info }, base, 256);Go's package also exposes hkdf.Extract and hkdf.Expand separately, and before Go 1.24 the same code lived in golang.org/x/crypto/hkdf. In Python, HKDFExpand is the expand-only variant for inputs that are already uniform keys.
Failure modes
- Using HKDF on passwords. HKDF is fast by design, so an attacker can try billions of guesses per second against a derived key. Passwords need a slow, memory-hard function such as Argon2, or PBKDF2 where compliance demands it.
- Same info for two purposes. Two keys derived with identical inputs are identical. Reusing an AES-GCM key in both directions with independent nonce counters leads to nonce reuse, which is catastrophic for GCM.
- Ambiguous info encoding. Concatenating a user name and a purpose without delimiters or lengths lets two contexts encode to the same bytes.
- Length as separation. See the prefix property: keep the length in info.
- HMAC key quirks in the salt. HMAC hashes keys longer than the block size (64 bytes for SHA-256) and zero-pads shorter ones, so salts
abcandabc\x00give the same PRK. Never use the salt to distinguish contexts; that is info's job. - Overlong output. Requests beyond 255 * HashLen raise an error. If you need more, you need a stream cipher or a DRBG seeded from HKDF output, not a longer Expand.
- Logging the PRK. The PRK is as sensitive as the original secret. Zeroize it after use where the language allows, and keep it out of debug output.
Trade-offs
| Option | Use it for | Avoid it for |
|---|---|---|
| HKDF (RFC 5869) | High-entropy secrets: DH outputs, PSKs, master keys | Passwords |
| Expand only | Inputs that are already uniform keys | Structured or biased inputs |
| Raw hash of the secret | Nothing new; legacy compatibility only | Anything you design today |
| SP 800-108 counter-mode KDF | Deriving from an existing key in NIST-certified stacks | Non-uniform input (it has no extract step) |
| PBKDF2, scrypt, Argon2 | Low-entropy passwords | Session keys: the cost buys nothing there |
HKDF also fits the extract-then-expand pattern in NIST SP 800-56C for key establishment, which is why it appears in FIPS-oriented stacks as well as modern protocols. Its cost is negligible: two HMAC calls plus one per output block.
What to do next
- Run the code above and confirm RFC 5869 test cases 1 and 3 byte for byte.
- List every key your system derives and write down its info string. If two keys share inputs, add a role, direction or version field.
- Make sure each info string includes the output length and is encoded with fixed fields or length prefixes.
- Audit callers for passwords flowing into HKDF and replace them with Argon2.
- Switch to your platform's library:
cryptographyin Python,crypto/hkdfin Go 1.24 and later,crypto.hkdfSyncin Node, WebCrypto in browsers. - Read the Diffie-Hellman article to see where the IKM usually comes from, then trace HKDF through a real TLS 1.3 handshake.