Password hashing exists for one bad day: the day an attacker copies your user table. From then on the attacker can guess offline, as fast as their hardware allows, with no rate limit and no alerting. The only thing you control is how expensive each guess is. A general-purpose hash such as SHA-256 is designed to be fast, and a single modern GPU computes billions of them per second, so an unsalted or merely salted SHA-256 table falls to dictionary attacks within hours. Password hashing functions are deliberately slow, salted and tunable.
This article focuses on bcrypt, still the most widely deployed of them, and on the engineering around it: how the algorithm works, what its 60-character string encodes, the 72-byte limit that caused a real authentication bypass, how to choose parameters, how to migrate from weak legacy hashes without forcing a reset, and how to keep the login path from becoming a denial-of-service target. The comparison with PBKDF2, scrypt and Argon2id from the derivation side is in key derivation functions.
What a password hash must resist
Three properties matter. A unique random salt per password means identical passwords produce different hashes, so the attacker cannot crack all users at once or use precomputed tables. A tunable work factor lets you raise the cost of each guess as hardware improves. And resistance to parallel hardware stops a GPU or ASIC from turning a large budget into an equally large speedup. PBKDF2 has the first two. bcrypt adds a modest form of the third, because each guess needs 4 KB of rapidly and unpredictably accessed S-box state, which suits CPUs better than the narrow per-thread memory of GPUs. scrypt and Argon2id go further by requiring megabytes per guess, which is why they are called memory-hard.
None of this protects a weak password for long. A slow hash turns a password that falls in a second into one that falls in a few days; it cannot save password123. That is why breached-password screening at registration, for example with a Bloom filter over a breach corpus, belongs in the same design.
Inside bcrypt
bcrypt, published by Provos and Mazieres in 1999, is built on the Blowfish cipher, whose key schedule is unusually expensive: it fills an 18-word P-array and four 256-entry S-boxes by repeatedly encrypting with the cipher itself. bcrypt makes that schedule tunable and salt-dependent in a construction called EksBlowfish, for expensive key schedule.
def bcrypt(cost, salt16, password):
state = init_from_pi_digits() # P-array and S-boxes from hex digits of pi
state = expand_key(state, salt16, password) # mix both salt and password in
for _ in range(2 ** cost): # the tunable, expensive part
state = expand_key(state, 0, password)
state = expand_key(state, 0, salt16)
ctext = b"OrpheanBeholderScryDoubt" # 24 bytes = three 64-bit blocks
for _ in range(64):
ctext = blowfish_encrypt_ecb(state, ctext)
return encode(cost, salt16, ctext[:23]) # standard format keeps 23 of 24 bytesEach increment of the cost doubles the loop, so cost 12 does 4,096 iterations and twice the work of cost 11. The result is serialised as a self-describing string, shown in Figure 1: a variant prefix, a two-digit cost, a 22-character salt and a 31-character hash, both in bcrypt's own base64 alphabet. Because the string carries its parameters, a verifier reads them back from the stored value, which is what makes gradual upgrades possible.
The 72-byte limit and the variant prefixes
The password is cycled into the 18 32-bit words of the P-array, so bcrypt uses at most 72 bytes of input. Most implementations silently ignore everything after byte 72, and some stop at a NUL byte. Two consequences follow. Long passphrases are not as strong as they look, and with multi-byte UTF-8 the limit can arrive well before 72 characters. More dangerously, if you ever feed bcrypt a concatenation such as user id plus username plus password, a long enough prefix pushes the password out of the hashed region entirely.
That is not hypothetical. In October 2024 Okta disclosed that its AD/LDAP delegated authentication cache generated keys by bcrypt-hashing a combination of user id, username and password; for usernames of 52 characters or more, a login could match a cached key without the correct password under certain conditions. The fix is to never put anything except the password into bcrypt's input, and to handle long passwords deliberately.
The variant prefixes record bug history. $2a$ is the original. In 2011 a sign-extension bug was found in the crypt_blowfish implementation for characters with the high bit set, which led to $2x$ (the buggy behaviour, for legacy hashes) and $2y$ (correct). In 2014 OpenBSD fixed a bug where the password length was stored in an 8-bit counter and wrapped for very long passwords, introducing $2b$. New hashes should use $2b$; a good library verifies the others.
If you must support passwords longer than 72 bytes with bcrypt, pre-hash with a keyed hash and encode the result so it contains no NUL bytes: bcrypt(base64(HMAC-SHA256(pepper, password))). The key matters. Pre-hashing with a plain, unsalted SHA-256 exposes users to password shucking, where an attacker matches your inner values against unsalted SHA-256 hashes leaked from other sites. Alternatively, reject inputs over a sane maximum or switch new hashes to Argon2id, which has no such limit.
Choosing an algorithm and its cost
Algorithm choice is now fairly settled. The OWASP Password Storage Cheat Sheet recommends Argon2id first, scrypt if Argon2id is unavailable, bcrypt for legacy systems, and PBKDF2 where FIPS-140 validated primitives are required. Its current minimums are summarised below; treat them as floors and check the sheet for revisions.
| Function | OWASP minimum configuration | Notes |
|---|---|---|
| Argon2id | m = 19 MiB, t = 2, p = 1 | Equivalent trade-offs allow more memory and fewer passes |
| scrypt | N = 2^17, r = 8, p = 1 | About 128 MiB per hash |
| bcrypt | cost 10 or more | 72-byte input limit |
| PBKDF2-HMAC-SHA256 | 600,000 iterations | Not memory-hard; use where FIPS is required |
Minimums are not targets. Choose the highest cost your login path can afford, by measurement on production hardware. A common budget is a few hundred milliseconds of one core per verification for interactive logins; machine-to-machine authentication should use random API keys, which need only a fast hash, instead.
import time
import bcrypt
def pick_bcrypt_cost(target_ms=250, lo=10, hi=16):
pw = b"correct horse battery staple"
best = lo
for cost in range(lo, hi + 1):
salt = bcrypt.gensalt(rounds=cost)
t0 = time.perf_counter()
bcrypt.hashpw(pw, salt)
ms = (time.perf_counter() - t0) * 1000
print(f"cost {cost}: {ms:.0f} ms")
if ms > target_ms:
break
best = cost
return best
Verify, then upgrade
The application code should be short and should never compare hashes itself; libraries do the constant-time comparison. With the argon2-cffi library, new hashes use Argon2id, and the same function verifies existing bcrypt hashes and upgrades them on successful login.
import bcrypt
from argon2 import PasswordHasher
from argon2.exceptions import VerifyMismatchError, InvalidHashError
ph = PasswordHasher(time_cost=2, memory_cost=19456, parallelism=1) # KiB: 19 MiB
MAX_PASSWORD_BYTES = 1024
DUMMY_HASH = ph.hash("not-a-real-password")
def hash_new(password: str) -> str:
return ph.hash(password)
def verify_and_upgrade(user, password: str) -> bool:
if len(password.encode("utf-8")) > MAX_PASSWORD_BYTES:
return False
stored = user.password_hash if user else DUMMY_HASH # same cost for unknown users
try:
if stored.startswith("$2"):
pw = password.encode("utf-8")
ok = len(pw) <= 72 and bcrypt.checkpw(pw, stored.encode("ascii"))
needs_upgrade = ok
else:
ok = ph.verify(stored, password)
needs_upgrade = ok and ph.check_needs_rehash(stored)
except (VerifyMismatchError, InvalidHashError, ValueError): # ValueError: malformed bcrypt row
return False
if user is None or not ok:
return False
if needs_upgrade:
user.password_hash = hash_new(password)
user.save()
return TrueThe dummy hash prevents user enumeration by timing: an unknown username costs the same as a wrong password. Rejecting bcrypt inputs over 72 bytes, rather than letting them truncate, keeps behaviour explicit, and the input cap stops a megabyte-long password from being used to burn CPU.
Migrating legacy hashes without a reset
Suppose you inherit a table of unsalted MD5 hashes. Waiting for each user to log in leaves the inactive majority exposed indefinitely. Instead, wrap every hash immediately: compute bcrypt(md5_hex) for each row offline and store it with a marker such as legacy-md5$. At login, compute MD5 of the submitted password, verify it against the wrapped bcrypt value, and on success replace the record with a fresh Argon2id hash of the plaintext. After the batch job, no fast hash remains on disk, and active users migrate to the clean scheme naturally. Delete any backups that still hold the old column, because a migration that leaves MD5 in last month's snapshot has not finished. Note that the MD5 hex digest is 32 bytes, well inside bcrypt's 72-byte limit.
Worked example: raising the cost of a live user base
A service has stored bcrypt hashes at cost 10 since 2017 and wants cost 12. Because the cost is written into every string, the first step is to measure where you stand, straight from the database, with no plaintext involved.
-- characters 5-6 of a bcrypt string are the two-digit cost
SELECT substring(password_hash FROM 5 FOR 2) AS cost, count(*)
FROM users
WHERE password_hash LIKE '$2%'
GROUP BY 1 ORDER BY 1;Benchmarking on the application servers shows 70 ms per verification at cost 11 and about 140 ms at cost 12, the doubling the algorithm predicts, so each login at the new cost uses roughly four times the CPU of a cost-10 login. The team raises the cost in configuration, and the verify function's upgrade branch re-hashes each user at their next successful login. That login pays twice, one cost-10 verify plus one cost-12 hash, so CPU on the hashing pool spikes in the first days as the most active users migrate and then settles at the new, higher baseline. Watching the SQL histogram weekly shows the cost-10 share falling quickly for the first two weeks and then flattening: the remainder are dormant accounts.
Those dormant hashes are still at cost 10, and they are the ones an attacker with a leaked table will try first. Two options remain: wrap them, as in the migration above, or expire them and require a reset through email at the next login. Either way, the rollout is not finished until the histogram has a single bar.
Operating the login path
- Pepper. An HMAC key held in a KMS or HSM, outside the database, means a stolen table alone cannot be attacked at all. Plan rotation up front: a pepper cannot be changed without the plaintext, so store a pepper id with each hash.
- Denial of service. Slow hashing is a lever for attackers too. Rate limit before hashing, cap input length, and isolate hashing in a bounded pool.
- Raise cost over time. Review parameters yearly; check_needs_rehash and the cost field in bcrypt strings let upgrades roll out on login.
- Never log passwords. Scrub request bodies on the login route from access logs, error trackers and tracing spans.
- Integrity of the comparison. Use the library's verify, never
==on strings you built, and never compare only a prefix.
Failure modes
| Mistake | Consequence |
|---|---|
| Fast hash (MD5, SHA-1, SHA-256), even salted | Dictionary attack at GPU speed |
| Concatenating other fields into bcrypt input | 72-byte truncation can drop the password |
| Unkeyed SHA-256 pre-hash before bcrypt | Password shucking with other breaches |
| Cost chosen once in 2015 and never revisited | Work factor eroded by hardware |
| No dummy hash for unknown users | Username enumeration by response time |
| Unbounded password length with Argon2 or PBKDF2 | Cheap CPU exhaustion |
| Encrypting passwords instead of hashing | One key leak exposes every password |
For background on why SHA-256 is the wrong tool here even though it is a sound hash, see the SHA family.
What to do next
- Inventory every place a password or password-equivalent is stored, including caches, backups and legacy tables.
- Grep for any bcrypt input that is not exactly the password, and fix it.
- Run the cost benchmark on production hardware and record the chosen parameters.
- Use Argon2id for new hashes with at least the OWASP floor, and verify old bcrypt hashes with upgrade-on-login.
- Wrap any remaining fast hashes in bcrypt or Argon2id today, then delete old backups.
- Add a dummy hash for unknown users, input-length caps, rate limits and a bounded hashing pool.
- Put a yearly review of parameters in the security calendar.