Authentication & Password Hashing

How a backend verifies who's making a request — never storing raw passwords, and remembering that someone is logged in across future requests.

What is it?

A backend regularly needs to know who's actually asking: "who is this request from, and are they allowed to do this?" That's authentication — proving identity, usually by checking a password at login.

The single most important rule: never store a user's actual password. Instead, store a hash — the output of a one-way function (like bcrypt or argon2) that scrambles the password so it can't be reversed back into the original, even if the entire database leaks. At login, you hash the freshly submitted password the same way and compare the two hashes — the plaintext password itself is never stored or compared directly.

HTTP itself doesn't remember anything between requests, so once someone logs in, the backend needs a way to keep recognizing them on later requests too — either a session (the server remembers who's logged in, and gives the browser a cookie holding just an id to look it up) or a token like a JWT (a small, signed packet of data the client holds and resends, which the server can verify without storing anything itself).

Explain like I'm 10

Hashing a password is like feeding it through a paper shredder: you get scrambled confetti out, and there's no way to feed the confetti back in and reconstruct the original page. To check a password later, you shred the newly typed one the same way and compare the confetti — you never keep the original page around to compare against directly.

Examples

Hashing on signup, comparing on login (bcrypt)

const bcrypt = require("bcrypt");

// Signup: hash before storing anything
const passwordHash = await bcrypt.hash(plainPassword, 10);
await pool.query(
  "INSERT INTO users (email, password_hash) VALUES ($1, $2)",
  [email, passwordHash]
);

// Login: compare against the stored hash
const user = await findUserByEmail(email);
const isValid = user && (await bcrypt.compare(submittedPassword, user.password_hash));
if (!isValid) return res.status(401).json({ error: "invalid credentials" });

The plaintext password only ever exists briefly in memory — the database stores password_hash, never the password itself, and login works by comparing hashes, not by 'unhashing' anything.

Protecting a route with an auth middleware

function requireAuth(req, res, next) {
  const token = req.headers.authorization?.split(" ")[1];
  if (!token || !isValidToken(token)) {
    return res.status(401).json({ error: "unauthorized" });
  }
  req.user = decodeToken(token);
  next();
}

app.get("/profile", requireAuth, (req, res) => {
  res.json(req.user);
});

This is the requireAuth middleware referenced back in the middleware topic — it runs before the route handler and only calls next() once it has confirmed who the caller actually is.

The same idea in FastAPI — a dependency instead of middleware

from fastapi import Depends, FastAPI, HTTPException
from fastapi.security import OAuth2PasswordBearer
import jwt

app = FastAPI()
oauth2_scheme = OAuth2PasswordBearer(tokenUrl="login")

def get_current_user(token: str = Depends(oauth2_scheme)):
    try:
        payload = jwt.decode(token, SECRET_KEY, algorithms=["HS256"])
    except jwt.InvalidTokenError:
        raise HTTPException(status_code=401, detail="invalid token")
    return payload

@app.get("/profile")
async def profile(user: dict = Depends(get_current_user)):
    return user

OAuth2PasswordBearer tells FastAPI how to find the token (an Authorization: Bearer <token> header) and feeds it into get_current_user, which decodes and verifies the JWT's signature — the same protection as requireAuth, expressed as a dependency rather than a step in a middleware chain.

How it works

Password-hashing algorithms (bcrypt, scrypt, argon2) are deliberately slow, and mix in a random salt for every password, so that two users with the same password get different-looking hashes and an attacker can't precompute one shared table of hashes to crack many accounts at once (a "rainbow table"). After a successful login, staying recognized on future requests works one of two ways: a session stores the logged-in state server-side and hands the browser a cookie with just an id to look it up, while a token like a JWT carries the identity data itself, signed so the server can verify it wasn't tampered with — without needing to store anything server-side at all.

Why does it exist?

If a database stored passwords directly, any breach, backup leak, or careless insider access would hand over every user's actual password immediately — and because people reuse passwords across sites, that damage doesn't stay contained to just this one app. Hashing exists to make stored passwords useless to an attacker even if the entire database is exposed. Authentication overall exists because a server otherwise has no way to tell a returning, legitimate user apart from anyone else simply sending a similar-looking request.

When to use it

Any system with real user accounts needs authentication once there's a "you" to log in as — hashing wherever a password is stored, and a session or token wherever the backend needs to keep recognizing a logged-in user across requests.

When not to use it

Don't hand-roll password hashing or session handling for anything real — use a well-audited library (bcrypt, argon2) or a full auth framework/service rather than inventing your own scheme, since subtle cryptographic mistakes are easy to make and extremely costly. Internal tools with no real user accounts to protect may not need full authentication at all.

Common mistakes

  • Storing passwords in plaintext, or hashing them with a fast general-purpose hash (like MD5 or SHA-256) instead of a slow, purpose-built one like bcrypt.

  • Comparing passwords with a simple equality check after hashing manually without a salt, letting identical passwords produce identical hashes and enabling rainbow-table attacks.

  • Confusing authentication ('who are you?') with authorization ('what are you allowed to do?') — being logged in doesn't automatically mean access to every resource should be granted.

Practice exercises

  1. Easy:

    Explain why a database breach is far less damaging if passwords were hashed with bcrypt instead of stored as plaintext.

  2. Medium:

    Write a signup and login flow using bcrypt.hash and bcrypt.compare, including the SQL to store and look up a user's password_hash.

  3. Hard:

    Explain the difference between session-based and token-based (JWT) authentication, including exactly where the 'you're logged in' state lives in each approach.

Interview questions

Why shouldn't passwords ever be stored in plaintext?

Because anyone who gains access to the database — through a breach, an insider, or a leaked backup — would get every user's actual password immediately, and since people reuse passwords across sites, that damage spreads beyond just this app.

What's the difference between hashing and encryption for storing passwords?

Encryption is reversible with the right key; hashing is designed to be irreversible — which is exactly what's wanted for passwords, since you only ever need to verify a match, never recover the original.

What's the difference between authentication and authorization?

Authentication confirms who a user is; authorization determines what that already-authenticated user is allowed to do.

What is a salt, and what specific attack does it defeat that hashing alone doesn't?

Random data mixed into a password before hashing, unique per user; without it, two users sharing the same password would produce identical hashes, letting an attacker precompute one shared table of hashes for common passwords, a rainbow table, and crack every matching account at once — a per-user salt makes each stored hash unique even for an identical password, so no precomputed table applies.

With bcrypt, where is the salt actually stored, and why doesn't that make it useless?

It's embedded as part of the resulting hash string itself, so there's no need for a separate column — this is safe because a salt's purpose isn't to be secret, only to be unique per password, forcing an attacker to redo the expensive hashing work for every single user individually instead of reusing one precomputed table.

What does the cost factor passed to bcrypt, like the 10 in `bcrypt.hash(password, 10)`, actually control?

How many times the underlying hashing algorithm iterates internally — a higher number makes each hash deliberately slower to compute, roughly doubling in time with each increment; a legitimate login only pays that cost once, while an attacker trying to brute-force guesses pays it for every single one, making large-scale guessing impractical.

Why can't you simply use a fast, general-purpose hash like SHA-256 for passwords, even though it's cryptographically secure?

SHA-256 is deliberately optimized to be fast, which is exactly wrong for password hashing — its speed lets an attacker with a stolen hash try billions of guesses per second on cheap hardware, especially GPUs; bcrypt, scrypt, and argon2 are deliberately slow and tunable specifically to make brute-forcing many guesses computationally expensive.

What's a pepper, and how does it differ from a salt?

A single secret value, unlike a salt which is per-user, applied to every password before hashing but kept outside the database entirely, e.g. in an environment variable or secrets manager — so a full database leak alone doesn't hand over enough to brute-force any password, since the attacker would also need the pepper, which was never stored with the hashes.

How does `bcrypt.compare(submittedPassword, storedHash)` work without ever reversing the stored hash back into the original password?

It extracts the salt embedded in storedHash, hashes submittedPassword using that same salt and cost factor, and compares the two resulting hash strings for equality — the comparison is between two freshly computed hashes, never a decryption of the stored one.

What's a timing attack in the context of comparing a submitted password's hash to a stored one, and how do hashing libraries defend against it?

A naive character-by-character comparison returns as soon as it finds the first mismatched character, so the time a comparison takes can leak how many characters at the start were correct; well-built hash-comparison functions use a constant-time comparison that always takes the same amount of time regardless of where or whether a mismatch occurs, so timing reveals nothing.

What are the three parts of a JWT, separated by dots, and what does each contain?

A header, declaring the token type and signing algorithm; a payload, the actual claims like a user id and expiration; and a signature, computed over the header and payload using a secret or private key — written as header.payload.signature, each part base64url-encoded.

Is the payload of a JWT encrypted — can anyone read it just by decoding it?

No — a standard JWT's payload is only base64url-encoded, not encrypted, so anyone holding the token can decode and read its contents directly; the signature only proves the payload wasn't tampered with, it doesn't hide the contents, so sensitive data shouldn't be placed in a JWT payload.

What does verifying a JWT's signature actually protect against, if the payload itself isn't secret?

It protects against tampering — if anyone modifies the payload, like changing a role claim from "user" to "admin", without knowing the server's secret key, recomputing a valid signature for the altered payload is infeasible, so verification (recomputing the expected signature and comparing) rejects any token whose payload was changed after issuing.

What's the `alg: none` JWT vulnerability, and how is it typically prevented?

Some poorly implemented verification code would trust a token's own declared alg header, including a value of "none" meaning no signature at all, and skip verification entirely — an attacker could craft a token with alg: none and arbitrary claims and have it accepted; the fix is for the verifying code to explicitly specify and enforce which algorithm it accepts, ignoring whatever the token itself claims to use.

Why is it a mistake to let JWT verification accept whatever algorithm the token specifies, rather than pinning to one expected algorithm?

A token signed asymmetrically, like RS256 verified with a public key, and one signed symmetrically, like HS256 verified with a shared secret, can be confused by an attacker — e.g. a public key mistakenly treated as an HMAC secret — letting a forged token be accepted as validly signed; explicitly requiring one known algorithm during verification closes this off.

What claim in a JWT controls when it expires, and what happens if verification code forgets to check it?

The exp claim, a Unix timestamp; if it isn't checked, a token remains 'valid' forever once issued, even long after it should have expired, defeating the purpose of having an expiration at all.

Why can't a JWT be 'logged out' or revoked the same simple way a session can?

A session is looked up server-side on every request, so deleting that session record immediately invalidates it; a JWT is self-contained and verified purely by its signature, with nothing to look up — the server has no built-in way to mark one already-issued token invalid before its natural expiration, without maintaining a separate server-side revocation list.

What's a common approach to work around the difficulty of revoking JWTs immediately?

Keep JWT expirations short, minutes rather than days, and pair them with a separate, longer-lived refresh token that is tracked server-side and can be revoked; revoking the refresh token stops new access tokens from being issued, even though any already-issued short-lived access token still works until it naturally expires soon after.

What's a refresh token, and how does its lifecycle typically differ from an access token's?

A longer-lived credential, stored server-side so it can be revoked, and used only to obtain a new short-lived access token when the old one expires — the access token is what's sent with every regular API request, while the refresh token is used rarely, just to mint new access tokens, reducing how often the more sensitive credential needs to be transmitted.

What's the security tradeoff of storing a JWT in localStorage versus in an httpOnly cookie?

A token in localStorage is readable by any JavaScript running on the page, so it's directly exposed to theft via an XSS vulnerability; an httpOnly cookie can't be read by JavaScript at all, protecting it from XSS, but cookies are automatically sent by the browser on matching requests, which opens a different risk, CSRF, needing its own defenses like a SameSite attribute or a CSRF token.

What is session fixation, and how does it differ from simply stealing a session cookie?

Getting a victim to authenticate using a session id the attacker already knows, rather than stealing an existing one — e.g. tricking them into using a pre-set session id before login — so the attacker can then use that same, now-authenticated session id themselves; the standard defense is generating a brand new session id at the moment of successful login, invalidating whatever id existed before.

Why should a login endpoint return the same generic error message whether the email doesn't exist or the password is wrong, rather than distinguishing the two?

Returning a different message for 'no such user' versus 'wrong password' lets an attacker enumerate which email addresses actually have accounts on the system, one guess at a time — a single generic message reveals nothing about which half of the pair was incorrect.

Why should login attempts be rate-limited per account or per IP?

Without a limit, an attacker can attempt an unlimited number of password guesses in an automated brute-force attack; rate-limiting, e.g. locking out or slowing down after several failed attempts, makes that kind of large-scale guessing impractical, on top of whatever protection the hashing algorithm's cost factor already provides for a single guess.

What's a secure way to implement a 'forgot password' flow, at a high level?

Generate a single-use, random, unguessable token unrelated to the user's actual password, store its hash server-side alongside a short expiration, email a link containing that token to the account's registered address, and only allow setting a new password if a matching, unexpired token is presented — never email the user's actual existing password, since that would mean it was recoverable, i.e. not properly hashed to begin with.

Why should a password-reset token be single-use and short-lived?

A token that keeps working after use, or that never expires, remains a valid way to take over the account indefinitely if it's ever intercepted, from an email account compromise or a logged browser history; single-use and short expiration both shrink the window an intercepted token would still work in.

What does multi-factor authentication add on top of a password, and why does it meaningfully raise security even if the password is compromised?

It requires a second, independent proof of identity, commonly a time-based one-time code from an authenticator app or a hardware key, so a leaked or guessed password alone isn't sufficient to log in; an attacker would additionally need to compromise that separate factor, which typically isn't exposed by the same breach that leaked the password.

Why is serving a login form and its endpoint only over HTTPS essential, regardless of how well passwords are hashed server-side?

Hashing only protects the password once it's stored; over a plain HTTP connection, the submitted plaintext password, and any session cookie or token sent afterward, travels the network unencrypted and can be read by anyone able to observe that traffic — HTTPS protects the credential in transit, which hashing at rest does nothing for.

What's the difference between authentication middleware reading a token from an Authorization: Bearer header versus from a cookie?

A Bearer header requires the client to explicitly attach the token to every request it makes, common for APIs consumed by non-browser clients or SPAs managing their own tokens; a cookie is automatically attached by the browser to every matching request without the client-side code doing anything, which is convenient but is exactly what opens the door to CSRF unless mitigated.

In a `requireAuth` middleware pattern, why does it call `next()` only after successfully decoding the token, rather than always calling it?

Calling next() unconditionally would let the route run regardless of whether the caller was actually authenticated; only calling it after the token check succeeds ensures the protected route's handler never executes for an unauthenticated or invalid request — the 401 response is returned instead, and the chain stops there.

Why must the same secret, or key pair, used to sign a JWT also be available wherever it needs to be verified?

The verifying code recomputes what it expects the signature to be, using that secret or key, and compares it to the token's actual signature — without access to the correct secret, or the matching public key for asymmetric signing, verification can't confirm authenticity at all, so the secret's confidentiality is exactly as important as a password's.

What's the practical effect of rotating the secret key used to sign JWTs?

Every previously issued token was signed with the old key, so once the app verifies only against the new key, all outstanding tokens signed under the old one immediately fail verification and are effectively invalidated — a useful way to force universal logout after a suspected key compromise, but one that logs out every currently authenticated user at once, not just a targeted few.

Why does the column name `password_hash` matter beyond just being descriptive?

It's a habit that helps prevent an easy, costly mistake — accidentally writing or logging a raw password field somewhere in code that assumes the column literally holds the password — by making the stored value's actual nature, a hash and not the real password, explicit everywhere it's referenced.

What's argon2's relationship to bcrypt, and why might a newer system choose it instead?

Argon2 is a newer password-hashing algorithm, winner of the 2015 Password Hashing Competition, that unlike bcrypt lets you tune not just computational cost but also memory usage, deliberately requiring a lot of RAM per hash — which specifically makes it more resistant to attacks using GPUs or custom hardware optimized for fast computation but not large memory use.

Where does the boundary sit between what 'authentication' implementation covers, like hashing and JWTs, and the broader conceptual question of sessions vs. tokens as an architectural choice?

The mechanics — how a password is safely hashed and verified, how a JWT is structured, signed, and checked — are implementation details of the login step itself; whether to keep users recognized afterward via server-side sessions or client-held tokens is a wider architectural tradeoff around statelessness, scaling, and revocation that applies beyond just password-based login, and is covered in more depth as its own system-design topic.