Authentication & Sessions

How a server remembers who you are across multiple requests, even though HTTP itself doesn't.

What is it?

HTTP is stateless — the server doesn't automatically remember anything about you between one request and the next. But most apps clearly do remember you: you log in once and stay logged in as you click around. That memory has to be built on top of HTTP deliberately, and there are two common ways to do it: sessions (the server keeps track of who's logged in, and gives your browser a small id to prove which session is yours) and tokens (the server gives your browser a self-contained piece of data that proves who you are, without the server needing to store anything).

Explain like I'm 10

A session is like getting a wristband at a venue — the venue keeps a list of which wristband numbers were given to which paying customers, and checks your wristband against that list at each door. A token is more like getting a signed, tamper-proof ticket that itself already contains everything needed to prove it's valid — nobody needs to check a list, they just verify the ticket's signature.

Examples

A simple session-based login flow

// 1. User logs in with correct credentials
POST /login  { username: "amara", password: "..." }

// 2. Server creates a session record and sends back a cookie
Set-Cookie: sessionId=abc123; HttpOnly

// 3. Every later request automatically includes that cookie
GET /dashboard
Cookie: sessionId=abc123

// 4. Server looks up "abc123" in its session store to know who's asking

How it works

With sessions, the server stores session data (who's logged in, since when) in memory or a database, keyed by a random session id, and gives that id to the browser as a cookie; every later request includes the cookie, and the server looks up the matching session. With tokens (commonly JWTs), the server instead signs a small piece of data containing the user's identity, so it can be verified without a lookup — the token itself is the proof, as long as its signature checks out.

Why does it exist?

Without some form of authentication and a session/token mechanism, a server would have no way to distinguish one visitor from another, or remember that you already proved who you are — every request would need to include your full credentials again, which is both impractical and insecure.

When to use it

Use session-based authentication when you want the server to retain full control — able to instantly revoke access by deleting a session record. Use token-based authentication when you want to avoid a centralized lookup on every request, especially across multiple independent services, at the cost of tokens being harder to revoke early.

When not to use it

Don't build authentication from scratch for anything beyond a learning exercise — subtle mistakes here (like weak session ids, or improperly verifying a token's signature) are a common source of serious security vulnerabilities; use well-established, audited libraries and frameworks instead.

Common mistakes

  • Storing session ids or tokens somewhere JavaScript can read them (making them vulnerable to theft via XSS) instead of an HttpOnly cookie.

  • Assuming a token can be instantly revoked the way a session can — a signed token generally stays valid until it expires, unless you build extra revocation logic.

  • Sending credentials or session identifiers over plain HTTP instead of HTTPS, exposing them to anyone on the network.

Practice exercises

  1. Easy:

    Explain, in your own words, why HTTP being stateless means a server needs a separate mechanism to 'remember' a logged-in user.

  2. Medium:

    Describe the difference between a session id and a token, and one advantage each has over the other.

  3. Hard:

    Explain how a service could quickly revoke a compromised session versus a compromised token, and why one is harder.

Interview questions

Why can't the server just 'remember' a user across requests using HTTP alone?

HTTP is stateless by design — each request is handled independently, so remembering a user requires an explicit mechanism like a session or token layered on top.

What's the difference between a session and a token?

A session stores the user's state on the server, identified by an id the client holds; a token holds the user's state itself (signed, so it can be verified), without the server needing to store anything.

Why is a token harder to revoke early than a session?

A session can be invalidated by simply deleting its record on the server; a signed token is self-contained and remains valid until it expires, unless additional revocation tracking is built.

What's the difference between authentication and authorization?

Authentication answers 'who are you?' — verifying identity, typically via credentials; authorization answers 'what are you allowed to do?' — deciding what an already-identified user can access or perform.

What is a cookie, and how does the browser know to send it automatically on future requests?

A small piece of data the server asks the browser to store via a Set-Cookie response header; the browser then automatically attaches it to every subsequent request to the same domain (matching the cookie's scope) via the Cookie header, without the application needing to attach it manually.

What does the HttpOnly flag on a cookie protect against?

It prevents JavaScript running on the page from reading the cookie's value, so even if an attacker manages to inject a script (XSS), they can't steal a session id or token stored in an HttpOnly cookie directly.

What does the Secure flag on a cookie do?

It tells the browser to only ever send that cookie over HTTPS connections, preventing it from being exposed in plaintext if a request is ever made over unencrypted HTTP.

What is the SameSite cookie attribute, and what problem does it help prevent?

It controls whether a cookie is sent along with requests originating from a different site than the one that set it; setting it to Strict or Lax helps prevent the browser from automatically attaching a user's session cookie to requests forged by a malicious third-party site (CSRF).

What is CSRF (Cross-Site Request Forgery), and how does SameSite help mitigate it?

CSRF tricks a logged-in user's browser into sending an unwanted authenticated request to a site they're logged into, by embedding that request on a malicious page — since browsers normally attach cookies automatically regardless of where a request originates. SameSite mitigates this by refusing to attach the cookie to requests that originate from another site.

What is XSS (Cross-Site Scripting), and why does it matter for deciding where to store an auth token?

XSS is an attack where malicious script gets injected into and executed within a trusted page (e.g. via an unescaped user input); it matters for token storage because anything readable by JavaScript — like localStorage — can be read and exfiltrated by that injected script, whereas an HttpOnly cookie cannot.

Why is storing a JWT in localStorage often discouraged compared to an HttpOnly cookie?

localStorage is fully readable by any JavaScript running on the page, so a successful XSS attack can steal the token directly; an HttpOnly cookie is invisible to JavaScript entirely, closing off that particular theft vector (though cookies bring their own CSRF considerations).

What is a JWT structurally made up of?

Three base64url-encoded parts separated by dots: a header (describing the algorithm used), a payload (the claims/data, like user id and expiry), and a signature (computed over the header and payload, used to verify they haven't been tampered with).

Does base64-encoding a JWT's payload make its contents secret? Why or why not?

No — base64 is just an encoding, not encryption; anyone can decode a JWT's payload and read its contents. The signature only proves the payload hasn't been altered, so a JWT should never be assumed to be confidential, and shouldn't carry sensitive data unless it's also encrypted.

How does a server verify that a JWT hasn't been tampered with?

It recomputes the signature over the token's header and payload using its own secret (or public key, for asymmetric signing) and checks that it matches the signature attached to the token — any change to the header or payload would produce a different signature and fail verification.

What's the difference between signing a JWT with a symmetric secret (HMAC) versus an asymmetric key pair (RSA/ECDSA)?

With HMAC, the same secret both signs and verifies the token, so any service that can verify a token could also forge one; with an asymmetric key pair, only the holder of the private key can sign tokens, while any number of services can safely verify using the public key without being able to create valid tokens themselves.

Why might a system use short-lived access tokens paired with longer-lived refresh tokens?

A short-lived access token limits how long a stolen token remains useful, while the refresh token (kept more securely and used less often) lets the client obtain new access tokens without forcing the user to log in again — balancing security against convenience.

What happens if a refresh token is stolen, and how do systems typically mitigate that risk?

A stolen refresh token could let an attacker mint new access tokens indefinitely; systems mitigate this with refresh token rotation (issuing a new refresh token on each use and invalidating the old one) and by detecting reuse of an already-rotated token as a signal of theft.

Why do session-based systems scale less trivially across multiple servers than token-based systems, and how is that usually solved?

A session lives on whichever server created it, so a later request landing on a different server wouldn't find it unless something shares that state — this is usually solved with a shared session store (like Redis) that every server can read from, or by routing a client's requests back to the same server (sticky sessions).

What is a sticky session, and what problem does it cause for load balancing?

A load balancer configuration that routes all of a given client's requests to the same backend server, so that server's local, in-memory session data stays available; it undermines even load distribution and means losing that one server disproportionately affects the clients pinned to it.

Why is storing passwords in plain text always wrong, regardless of any other security measures in place?

If the database is ever compromised (a leak, a misconfigured backup, an insider), every user's actual password is immediately exposed — and since people reuse passwords across services, that single breach can compromise accounts on other sites too.

What is password hashing, and why is a slow, purpose-built algorithm like bcrypt preferred over a fast general-purpose hash like SHA-256?

Hashing stores a one-way transformation of the password instead of the password itself; bcrypt (and similar algorithms) are deliberately slow and tunable, which makes brute-forcing a huge number of guesses computationally expensive, whereas a fast hash like SHA-256 lets an attacker try billions of guesses per second on stolen hashes.

What is a salt, and what attack does it protect against?

A salt is random data added to a password before hashing, unique per user; it protects against precomputed rainbow-table attacks and against instantly spotting that two users share the same password, since identical passwords produce different hashes once salted.

Why is sending credentials over plain HTTP a serious vulnerability, even if the password is hashed before storage?

Hashing protects the password at rest in the database, but if it's transmitted over plain HTTP, anyone able to observe the network traffic (e.g. on shared Wi-Fi) can capture the plaintext password as it's sent, before it's ever hashed on the server.

What is multi-factor authentication (MFA), and why does it improve security beyond a strong password alone?

MFA requires proving identity with more than one independent factor (something you know, like a password, plus something you have, like a phone, or something you are, like a fingerprint) — so a stolen or guessed password alone isn't enough to gain access, since the attacker would also need the second factor.

What is OAuth, and what problem does it solve that a plain username/password login doesn't?

OAuth is a protocol for delegated access — it lets a user grant one application limited access to their data on another service without ever sharing their actual password with the first application, solving the problem of third-party apps needing credentials they shouldn't be trusted with.

What's the difference between using OAuth for delegated access and using it for 'Sign in with Google/GitHub'-style login?

Delegated access uses OAuth to grant a specific scope of permission to act on a resource (like reading a user's calendar); 'Sign in with' login repurposes the same flow purely to confirm identity (via OpenID Connect on top of OAuth), without necessarily requesting access to any other data.

Why is it a mistake to treat a valid JWT signature as proof that a user's access hasn't been revoked?

A valid signature only proves the token wasn't tampered with and was genuinely issued by the server — it says nothing about whether that user's access has since been revoked, since a self-contained token remains 'valid' by that check until it naturally expires, unless the system does extra work to track revocations.

Describe a realistic scenario where you'd choose token-based auth over session-based auth.

A system with multiple independent backend services (a microservices architecture, or a public API consumed by many different clients) benefits from tokens because any service holding the public key (or shared secret) can verify a request's identity on its own, without a shared, centralized session store every service must query.

Describe a realistic scenario where you'd choose session-based auth over token-based auth.

A traditional server-rendered web app with a single backend (or a small, tightly coupled set of services) benefits from sessions because it gets easy, immediate revocation (just delete the session record) and doesn't need to solve token-specific problems like revocation lists or key distribution across services it doesn't actually have.

Why should session ids and tokens be generated using a cryptographically secure random generator rather than something predictable?

If an id can be guessed or predicted (e.g. generated from a simple counter or a weak random source), an attacker could forge a valid session id or token for another user without ever needing to steal it, defeating the entire mechanism.