Password hashing exists for one bad day: the day an attacker copies your user table. From then on the attacker can guess offline, as fast as their hardware allows, with no rate limit and no alerting. The only thing you control is how expensive each guess is. A general-purpose hash such as SHA-256 is designed to be fast, and a single modern GPU computes billions of them per second, so an unsalted or merely salted SHA-256 table falls to dictionary attacks within hours. Password hashing functions are deliberately slow, salted and tunable.

This article focuses on bcrypt, still the most widely deployed of them, and on the engineering around it: how the algorithm works, what its 60-character string encodes, the 72-byte limit that caused a real authentication bypass, how to choose parameters, how to migrate from weak legacy hashes without forcing a reset, and how to keep the login path from becoming a denial-of-service target. The comparison with PBKDF2, scrypt and Argon2id from the derivation side is in key derivation functions.

What a password hash must resist

Three properties matter. A unique random salt per password means identical passwords produce different hashes, so the attacker cannot crack all users at once or use precomputed tables. A tunable work factor lets you raise the cost of each guess as hardware improves. And resistance to parallel hardware stops a GPU or ASIC from turning a large budget into an equally large speedup. PBKDF2 has the first two. bcrypt adds a modest form of the third, because each guess needs 4 KB of rapidly and unpredictably accessed S-box state, which suits CPUs better than the narrow per-thread memory of GPUs. scrypt and Argon2id go further by requiring megabytes per guess, which is why they are called memory-hard.

None of this protects a weak password for long. A slow hash turns a password that falls in a second into one that falls in a few days; it cannot save password123. That is why breached-password screening at registration, for example with a Bloom filter over a breach corpus, belongs in the same design.

Inside bcrypt

bcrypt, published by Provos and Mazieres in 1999, is built on the Blowfish cipher, whose key schedule is unusually expensive: it fills an 18-word P-array and four 256-entry S-boxes by repeatedly encrypting with the cipher itself. bcrypt makes that schedule tunable and salt-dependent in a construction called EksBlowfish, for expensive key schedule.

def bcrypt(cost, salt16, password):
    state = init_from_pi_digits()                 # P-array and S-boxes from hex digits of pi
    state = expand_key(state, salt16, password)   # mix both salt and password in
    for _ in range(2 ** cost):                    # the tunable, expensive part
        state = expand_key(state, 0, password)
        state = expand_key(state, 0, salt16)
    ctext = b"OrpheanBeholderScryDoubt"           # 24 bytes = three 64-bit blocks
    for _ in range(64):
        ctext = blowfish_encrypt_ecb(state, ctext)
    return encode(cost, salt16, ctext[:23])       # standard format keeps 23 of 24 bytes

Each increment of the cost doubles the loop, so cost 12 does 4,096 iterations and twice the work of cost 11. The result is serialised as a self-describing string, shown in Figure 1: a variant prefix, a two-digit cost, a 22-character salt and a 31-character hash, both in bcrypt's own base64 alphabet. Because the string carries its parameters, a verifier reads them back from the stored value, which is what makes gradual upgrades possible.

Anatomy of a bcrypt hash string (60 characters)$2b$variant12$cost: 2^1222 chars128-bit salt31 chars184-bit hash outputBase64 uses bcrypt's own alphabet ./A-Za-z0-9, not the standard one.EksBlowfishSetup(cost, salt, password)expand key with salt, then 2^cost rounds of ExpandKey(password), ExpandKey(salt)Encrypt 'OrpheanBeholderScryDoubt' 64 timesECB with the expensive key schedule; keep 23 of 24 bytesEverything needed to verify, variant, cost and salt, travels inside the string.
Figure 1. A bcrypt string is self-describing: variant, cost and salt are stored beside the output, so verification needs nothing else.

The 72-byte limit and the variant prefixes

The password is cycled into the 18 32-bit words of the P-array, so bcrypt uses at most 72 bytes of input. Most implementations silently ignore everything after byte 72, and some stop at a NUL byte. Two consequences follow. Long passphrases are not as strong as they look, and with multi-byte UTF-8 the limit can arrive well before 72 characters. More dangerously, if you ever feed bcrypt a concatenation such as user id plus username plus password, a long enough prefix pushes the password out of the hashed region entirely.

That is not hypothetical. In October 2024 Okta disclosed that its AD/LDAP delegated authentication cache generated keys by bcrypt-hashing a combination of user id, username and password; for usernames of 52 characters or more, a login could match a cached key without the correct password under certain conditions. The fix is to never put anything except the password into bcrypt's input, and to handle long passwords deliberately.

The variant prefixes record bug history. $2a$ is the original. In 2011 a sign-extension bug was found in the crypt_blowfish implementation for characters with the high bit set, which led to $2x$ (the buggy behaviour, for legacy hashes) and $2y$ (correct). In 2014 OpenBSD fixed a bug where the password length was stored in an 8-bit counter and wrapped for very long passwords, introducing $2b$. New hashes should use $2b$; a good library verifies the others.

If you must support passwords longer than 72 bytes with bcrypt, pre-hash with a keyed hash and encode the result so it contains no NUL bytes: bcrypt(base64(HMAC-SHA256(pepper, password))). The key matters. Pre-hashing with a plain, unsalted SHA-256 exposes users to password shucking, where an attacker matches your inner values against unsalted SHA-256 hashes leaked from other sites. Alternatively, reject inputs over a sane maximum or switch new hashes to Argon2id, which has no such limit.

Choosing an algorithm and its cost

Algorithm choice is now fairly settled. The OWASP Password Storage Cheat Sheet recommends Argon2id first, scrypt if Argon2id is unavailable, bcrypt for legacy systems, and PBKDF2 where FIPS-140 validated primitives are required. Its current minimums are summarised below; treat them as floors and check the sheet for revisions.

FunctionOWASP minimum configurationNotes
Argon2idm = 19 MiB, t = 2, p = 1Equivalent trade-offs allow more memory and fewer passes
scryptN = 2^17, r = 8, p = 1About 128 MiB per hash
bcryptcost 10 or more72-byte input limit
PBKDF2-HMAC-SHA256600,000 iterationsNot memory-hard; use where FIPS is required

Minimums are not targets. Choose the highest cost your login path can afford, by measurement on production hardware. A common budget is a few hundred milliseconds of one core per verification for interactive logins; machine-to-machine authentication should use random API keys, which need only a fast hash, instead.

import time
import bcrypt

def pick_bcrypt_cost(target_ms=250, lo=10, hi=16):
    pw = b"correct horse battery staple"
    best = lo
    for cost in range(lo, hi + 1):
        salt = bcrypt.gensalt(rounds=cost)
        t0 = time.perf_counter()
        bcrypt.hashpw(pw, salt)
        ms = (time.perf_counter() - t0) * 1000
        print(f"cost {cost}: {ms:.0f} ms")
        if ms > target_ms:
            break
        best = cost
    return best

Verify, then upgrade

The application code should be short and should never compare hashes itself; libraries do the constant-time comparison. With the argon2-cffi library, new hashes use Argon2id, and the same function verifies existing bcrypt hashes and upgrades them on successful login.

import bcrypt
from argon2 import PasswordHasher
from argon2.exceptions import VerifyMismatchError, InvalidHashError

ph = PasswordHasher(time_cost=2, memory_cost=19456, parallelism=1)   # KiB: 19 MiB
MAX_PASSWORD_BYTES = 1024
DUMMY_HASH = ph.hash("not-a-real-password")

def hash_new(password: str) -> str:
    return ph.hash(password)

def verify_and_upgrade(user, password: str) -> bool:
    if len(password.encode("utf-8")) > MAX_PASSWORD_BYTES:
        return False
    stored = user.password_hash if user else DUMMY_HASH     # same cost for unknown users
    try:
        if stored.startswith("$2"):
            pw = password.encode("utf-8")
            ok = len(pw) <= 72 and bcrypt.checkpw(pw, stored.encode("ascii"))
            needs_upgrade = ok
        else:
            ok = ph.verify(stored, password)
            needs_upgrade = ok and ph.check_needs_rehash(stored)
    except (VerifyMismatchError, InvalidHashError, ValueError):   # ValueError: malformed bcrypt row
        return False
    if user is None or not ok:
        return False
    if needs_upgrade:
        user.password_hash = hash_new(password)
        user.save()
    return True

The dummy hash prevents user enumeration by timing: an unknown username costs the same as a wrong password. Rejecting bcrypt inputs over 72 bytes, rather than letting them truncate, keeps behaviour explicit, and the input cap stops a megabyte-long password from being used to burn CPU.

Login path with verify, rehash-on-login and legacy wrappingLogin requestrate limitedLoad recordor dummy hashVerifyworker poolFailgeneric errornoNeeds rehash?old cost or schemeRe-hash and storecurrent parametersSessionsuccessyesyesnoThe plaintext exists only during a successful login, so that is the only moment an upgrade can happen.
Figure 2. Verification, a dummy hash for unknown users, and upgrade-on-login, which is the only point where the plaintext is available to re-hash.

Migrating legacy hashes without a reset

Suppose you inherit a table of unsalted MD5 hashes. Waiting for each user to log in leaves the inactive majority exposed indefinitely. Instead, wrap every hash immediately: compute bcrypt(md5_hex) for each row offline and store it with a marker such as legacy-md5$. At login, compute MD5 of the submitted password, verify it against the wrapped bcrypt value, and on success replace the record with a fresh Argon2id hash of the plaintext. After the batch job, no fast hash remains on disk, and active users migrate to the clean scheme naturally. Delete any backups that still hold the old column, because a migration that leaves MD5 in last month's snapshot has not finished. Note that the MD5 hex digest is 32 bytes, well inside bcrypt's 72-byte limit.

Worked example: raising the cost of a live user base

A service has stored bcrypt hashes at cost 10 since 2017 and wants cost 12. Because the cost is written into every string, the first step is to measure where you stand, straight from the database, with no plaintext involved.

-- characters 5-6 of a bcrypt string are the two-digit cost
SELECT substring(password_hash FROM 5 FOR 2) AS cost, count(*)
FROM users
WHERE password_hash LIKE '$2%'
GROUP BY 1 ORDER BY 1;

Benchmarking on the application servers shows 70 ms per verification at cost 11 and about 140 ms at cost 12, the doubling the algorithm predicts, so each login at the new cost uses roughly four times the CPU of a cost-10 login. The team raises the cost in configuration, and the verify function's upgrade branch re-hashes each user at their next successful login. That login pays twice, one cost-10 verify plus one cost-12 hash, so CPU on the hashing pool spikes in the first days as the most active users migrate and then settles at the new, higher baseline. Watching the SQL histogram weekly shows the cost-10 share falling quickly for the first two weeks and then flattening: the remainder are dormant accounts.

Those dormant hashes are still at cost 10, and they are the ones an attacker with a leaked table will try first. Two options remain: wrap them, as in the migration above, or expire them and require a reset through email at the next login. Either way, the rollout is not finished until the histogram has a single bar.

Operating the login path

  • Pepper. An HMAC key held in a KMS or HSM, outside the database, means a stolen table alone cannot be attacked at all. Plan rotation up front: a pepper cannot be changed without the plaintext, so store a pepper id with each hash.
  • Denial of service. Slow hashing is a lever for attackers too. Rate limit before hashing, cap input length, and isolate hashing in a bounded pool.
  • Raise cost over time. Review parameters yearly; check_needs_rehash and the cost field in bcrypt strings let upgrades roll out on login.
  • Never log passwords. Scrub request bodies on the login route from access logs, error trackers and tracing spans.
  • Integrity of the comparison. Use the library's verify, never == on strings you built, and never compare only a prefix.

Failure modes

MistakeConsequence
Fast hash (MD5, SHA-1, SHA-256), even saltedDictionary attack at GPU speed
Concatenating other fields into bcrypt input72-byte truncation can drop the password
Unkeyed SHA-256 pre-hash before bcryptPassword shucking with other breaches
Cost chosen once in 2015 and never revisitedWork factor eroded by hardware
No dummy hash for unknown usersUsername enumeration by response time
Unbounded password length with Argon2 or PBKDF2Cheap CPU exhaustion
Encrypting passwords instead of hashingOne key leak exposes every password

For background on why SHA-256 is the wrong tool here even though it is a sound hash, see the SHA family.

What to do next

  1. Inventory every place a password or password-equivalent is stored, including caches, backups and legacy tables.
  2. Grep for any bcrypt input that is not exactly the password, and fix it.
  3. Run the cost benchmark on production hardware and record the chosen parameters.
  4. Use Argon2id for new hashes with at least the OWASP floor, and verify old bcrypt hashes with upgrade-on-login.
  5. Wrap any remaining fast hashes in bcrypt or Argon2id today, then delete old backups.
  6. Add a dummy hash for unknown users, input-length caps, rate limits and a bounded hashing pool.
  7. Put a yearly review of parameters in the security calendar.
Key takeaway: Password hashing buys time after a breach by making every offline guess salted and expensive. bcrypt does that well inside its 72-byte input limit; feed it only the password, tune its cost by measurement, and prefer Argon2id for new hashes. Wrap legacy fast hashes immediately, upgrade on login, and guard the login path with dummy hashes, input caps and rate limits.