Building with WebSockets

What actually changes on the server when a connection stays open — handling events, pushing messages, broadcasting, and what happens when a client drops.

What is it?

A WebSocket connection is a real-time, two-way upgrade on top of a normal HTTP request. You don't need a whole separate subject to work with one — just a clear picture of a few things that are genuinely different from a normal request handler: the connection sticks around instead of closing after one response, either side can send a message at any moment, and your server code needs to actively track who's currently connected if it wants to reach more than one client at once.

That means a WebSocket route isn't shaped like (req, res) => ... anymore — it's shaped like "a connection just opened, here's what to do each time a message arrives on it, and here's what to do when it closes."

Explain like I'm 10

A normal API route is a phone call where you dial, ask one question, get one answer, and hang up. A WebSocket connection is more like a walkie-talkie left on: it stays open, either person can key in and speak whenever they want, and if you want to reach a whole group at once, you have to actually keep a list of everyone whose walkie-talkie is currently on.

Examples

Connection lifecycle and messages (Node.js, the ws library)

const { WebSocketServer } = require("ws");
const wss = new WebSocketServer({ port: 8080 });

wss.on("connection", (socket) => {
  console.log("client connected");

  socket.on("message", (data) => {
    console.log("received:", data.toString());
    socket.send("ack: " + data.toString());
  });

  socket.on("close", () => {
    console.log("client disconnected");
  });
});

This is the whole lifecycle in one place: 'connection' fires once when a client connects, 'message' fires every time that specific client sends something, and 'close' fires once the connection ends — there's no single request/response pair to reason about anymore.

Broadcasting to every connected client

const clients = new Set();

wss.on("connection", (socket) => {
  clients.add(socket);

  socket.on("close", () => clients.delete(socket));

  socket.on("message", (data) => {
    // send this message out to everyone else connected
    for (const client of clients) {
      if (client !== socket && client.readyState === client.OPEN) {
        client.send(data.toString());
      }
    }
  });
});

A single socket only knows about itself — reaching every connected client (like a chat room) means the server has to keep its own list of open connections and loop over it, adding and removing entries as clients connect and disconnect.

The same lifecycle in FastAPI

from fastapi import FastAPI, WebSocket, WebSocketDisconnect

app = FastAPI()

@app.websocket("/ws")
async def chat(websocket: WebSocket):
    await websocket.accept()
    try:
        while True:
            data = await websocket.receive_text()
            await websocket.send_text(f"ack: {data}")
    except WebSocketDisconnect:
        print("client disconnected")

await websocket.accept() completes the upgrade, the while True loop plays the same role as the 'message' event in Node (it just runs until the client sends something, over and over), and WebSocketDisconnect is FastAPI's way of surfacing the equivalent of the 'close' event.

How it works

Once a WebSocket connection is accepted, the server holds it open and both sides communicate through events rather than a single request/response pair: a "connected" event, repeated "message" events in either direction, and a "closed" event. Because the connection has no built-in concept of "everyone in this chat room," broadcasting is entirely the server's own responsibility — it has to keep some in-memory collection of currently-open connections and iterate over it to reach more than one client.

Reconnection is the client's responsibility, not the server's: a dropped connection (a phone losing signal, a laptop sleeping) just closes the socket, and a well-behaved client detects that and calls new WebSocket(...) again, usually with a short backoff delay, to reconnect on its own. The server doesn't "resume" an old connection — it just sees a brand-new one arrive.

Scaling is where WebSockets get genuinely harder than a stateless HTTP API: if a server keeps its list of connected clients only in its own memory, and you run more than one server instance behind a load balancer, a message from a client on server A can't reach a client connected to server B just by looping over server A's local list. That turns broadcasting into a distribution problem — every server instance needs to hear about every message, typically by having each instance publish incoming messages to a shared message broker (like Redis pub/sub) that every other instance is also subscribed to, and re-send them to its own locally connected clients.

Why does it exist?

Polling an HTTP endpoint every few seconds to check "did anything change?" wastes requests on "no" answers and still isn't fast enough for things that genuinely need to feel instant. WebSockets exist so a server can push a message the moment something happens, instead of waiting for the client to ask again.

When to use it

Reach for a WebSocket for chat, live notifications, collaborative editing, live dashboards, or multiplayer interactions — anything where the server frequently has something to say before the client asks.

When not to use it

For data that only changes occasionally, or where a few seconds of staleness is fine, a normal request (or occasional polling) is far simpler to build, test, and scale than a WebSocket — don't reach for a persistent connection just because "real-time" sounds appealing.

Common mistakes

  • Forgetting to remove a client from the tracked connections list on 'close'/WebSocketDisconnect, leaking references to dead connections and eventually crashing when the server tries to send to one.

  • Assuming the server needs to handle reconnection — it doesn't; a dropped connection is just gone, and it's the client's job to detect that and open a new one.

  • Keeping the list of connected clients only in one server's memory and expecting broadcasts to reach clients connected to a different server instance behind a load balancer.

Practice exercises

  1. Easy:

    Using the ws library, write the 'connection' and 'close' handlers needed to log when clients connect and disconnect.

  2. Medium:

    Extend the broadcasting example so a message is echoed back to every connected client, including the sender.

  3. Hard:

    Explain, step by step, why running three instances of the same WebSocket server behind a load balancer breaks naive in-memory broadcasting, and what changes would be needed to fix it.

Interview questions

Why does broadcasting a message to all connected clients require extra work on the server, when a socket already exists for each client?

Each socket only knows about its own single connection — reaching every client requires the server to maintain its own collection of currently-open connections and iterate over it.

Whose responsibility is reconnection — the client's or the server's?

The client's — a dropped connection simply closes, and a well-behaved client detects that and opens a brand-new connection, typically with a short backoff delay.

Why does scaling WebSockets across multiple server instances complicate broadcasting?

Because each instance only knows about the clients connected directly to it, a message from a client on one instance can't reach a client on another instance without some shared mechanism (like a pub/sub message broker) that every instance subscribes to.

What actually happens during the handshake that upgrades a connection from HTTP to a WebSocket?

The client sends a normal HTTP request carrying an Upgrade: websocket header; if the server agrees, it responds with a 101 status code and both sides switch the same underlying TCP connection over to the WebSocket protocol, after which it stays open for two-way messages.

Why can't a WebSocket route be shaped like a typical `(req, res) => ...` handler?

There's no single request that produces one response — the connection instead triggers separate events over its lifetime (opened, each message, closed), so the code has to be organized around those events rather than a single call-and-return.

In the ws library example, what does the `'connection'` event handler run for?

It runs once per new client that connects, handing back a socket object scoped to that specific client — any 'message'/'close' listeners registered inside it apply only to that one connection.

In the broadcasting example, why does the loop check `client !== socket`?

So a message is echoed to every other connected client but not back to the very client that just sent it — without that check, a sender would receive its own message echoed back to itself.

Why does the broadcasting loop also check `client.readyState === client.OPEN`?

A client's socket can still sit in the tracked collection while it's closing or already closed. Sending to a socket that isn't actually open anymore would throw or silently fail, so this guards against writing to a stale connection.

What would happen if a server never removed a socket from its tracked clients collection on close?

The collection would keep growing with references to dead connections that will never receive anything again — a memory leak, and eventually attempting to send to one of those dead sockets could throw an error that needs handling.

Debugging: a client disconnects and reconnects, but messages meant for them stop arriving even though the new connection succeeded. What's a likely bug?

The server is probably still tracking the old, closed socket object as 'this client' somewhere, instead of updating its tracking to the new socket object created by the fresh connection — a reconnect is a brand-new connection, not a resumed old one.

Why is detecting a dropped connection and reconnecting the client's job, not the server's?

The server has no way to distinguish 'this client is still there but silent for a while' from 'this client's connection quietly died' without the client itself noticing its own socket closed and acting on it — the server just sees a close event, or nothing at all.

Why do reconnection strategies typically use a backoff delay instead of retrying immediately in a tight loop?

Retrying instantly and repeatedly, especially from many clients at once (e.g. right after a server restart), could overwhelm the server with a flood of reconnect attempts — a backoff, often increasing between tries, spreads that load out and gives the server room to recover.

What is a heartbeat (ping/pong) mechanism used for in long-lived WebSocket connections?

Periodically sending a small ping and expecting a pong back lets either side detect a connection that's silently dead — e.g. the network dropped without a clean close — rather than waiting indefinitely for an underlying TCP timeout.

Why can't a server's in-memory list of connected clients be used directly to reach clients connected to a different server instance?

Each instance's list only contains the socket objects for connections made directly to it — a socket object simply can't be reached from a different process or machine, so another instance has no way to look up or send to it.

How does a message broker like Redis pub/sub solve the multi-instance broadcasting problem?

Every instance publishes incoming messages to a shared channel and subscribes to that same channel. When any instance publishes, every subscribed instance — including itself — receives it and forwards it to whichever of its own locally-connected clients should get it.

Scenario: with three server instances sharing state via Redis pub/sub, a client on instance A sends a chat message. Trace how a client on instance C receives it.

Instance A publishes the message to the shared channel; instance C, subscribed to that channel, receives it; instance C then loops over its own locally-connected sockets and sends the message directly to the client in question, the same way it would for a purely local broadcast.

Why does a load balancer typically need sticky sessions (session affinity) for WebSocket traffic specifically?

A WebSocket connection is one long-lived TCP connection to one specific instance — unlike stateless HTTP requests that can be freely routed per-request, all of a WebSocket's messages need to keep reaching that same instance for the life of the connection, or it breaks.

What's a practical reason to prefer a normal HTTP endpoint over a WebSocket for data that only changes occasionally?

A WebSocket requires holding a persistent connection open per client plus code to track and clean up connection state — for data where a few seconds of staleness is fine, that ongoing cost and complexity isn't justified when a plain request already solves it more simply.

Trap: a developer assumes that because a connection is 'always open,' incoming messages don't need their own validation beyond the initial handshake. Why is this risky?

A connection staying open doesn't mean every message on it should be trusted by default — unexpected or malicious content can arrive in any message at any point after connecting, so messages generally still need their own validation and authorization checks, not just a one-time check at connect time.

Why does the FastAPI example use a `while True` loop instead of a `'message'` event listener like the Node.js version?

FastAPI's WebSocket handling is built around async/await rather than events — await websocket.receive_text() blocks that handler until the next message arrives, so looping is what lets it keep handling messages one after another, functionally equivalent to a repeatedly-firing 'message' event.

What does `WebSocketDisconnect` represent in the FastAPI example, and why catch it with try/except rather than checking a return value?

It's an exception FastAPI raises when the connection closes, e.g. because the client disconnected mid-receive_text() call. Since receive_text() would otherwise hang waiting for a message that'll never come, catching this exception is how the code detects the connection ended and runs cleanup.

How is a WebSocket connection's statefulness fundamentally different from a typical REST request, beyond just staying open longer?

A REST request is stateless — the server handling it can, in principle, be swapped or scaled per-request with no memory required. A WebSocket connection has real, memory-resident state tied to one specific server process for its entire duration, which is exactly what makes horizontal scaling structurally harder.

Why would broadcasting via a shared database table be a poor choice for real-time messages, compared to pub/sub?

A database is built for durable storage and querying, not pushing new data to subscribers instantly. Every instance would need to continuously poll the table for new rows, adding latency and load, rather than being pushed a message the moment it's published, which pub/sub is specifically designed to do.

Advanced: what additional problem does horizontal scaling introduce for tracking who's currently connected, beyond message delivery?

A client's connected status now only lives in the memory of whichever specific instance they're connected to — answering whether a given user is online anywhere in the system requires checking across all instances or maintaining a shared, cross-instance record, not just consulting one instance's local list.

Follow-up: if a WebSocket connection requires authentication, at what point should that be checked?

Once, during the initial connection handshake, e.g. via a token in the connection request — matching how the connection's identity persists for its whole lifetime, rather than re-authenticating on every single message, though the app still needs a plan for what happens if that identity's permissions change while the connection stays open.