Validation & Sanitization

Checking that incoming data is well-formed and expected before your code uses it, and cleaning up anything unsafe it might contain.

What is it?

A backend can never fully trust the data that arrives in a request — even from your own frontend, because a request can be sent by anyone, using anything, not just the app you built. A field you expect to be a number might arrive as text; a required field might be missing entirely; a text field might contain something malicious, like a chunk of HTML or a database command hidden inside a name field.

Validation is the process of checking that incoming data matches what your code expects — the right fields are present, and they're the right type and shape — before you act on it. Sanitization goes a step further: actively cleaning or transforming data to strip out anything unsafe (like stripping HTML tags from a comment field) rather than just rejecting it outright.

Explain like I'm 10

Think of a bouncer at a club checking IDs at the door (validation — rejecting anyone who doesn't meet the requirements) versus a coat check that removes anything dangerous from a bag before letting it inside (sanitization — cleaning up what's allowed to pass through, rather than turning it away entirely).

Examples

Manual validation before using data

app.post("/signup", (req, res) => {
  const { email, age } = req.body;

  if (typeof email !== "string" || !email.includes("@")) {
    return res.status(400).json({ error: "Invalid email" });
  }
  if (typeof age !== "number" || age < 13) {
    return res.status(400).json({ error: "Invalid age" });
  }

  createUser({ email, age });
  res.status(201).json({ ok: true });
});

The route refuses to even attempt to create a user until it's confirmed the incoming data looks the way it's supposed to.

Using a validation library for a declarative schema

const { z } = require("zod");

const signupSchema = z.object({
  email: z.string().email(),
  age: z.number().min(13),
});

app.post("/signup", (req, res) => {
  const result = signupSchema.safeParse(req.body);
  if (!result.success) {
    return res.status(400).json({ error: result.error.issues });
  }
  createUser(result.data);
  res.status(201).json({ ok: true });
});

Instead of hand-writing every check, a schema declares the expected shape once, and the library validates the incoming data against it in one call.

How it works

Validation typically runs early — often as middleware, before a route's main logic — comparing incoming data (the body, query string, or route parameters) against a set of rules: is this field present, is it the right type, does it fall within an allowed range or set of values. If the data fails, the request is rejected immediately with a clear error, before touching a database or running business logic. Sanitization similarly runs early, but transforms the data (trimming whitespace, escaping special characters, removing disallowed HTML) rather than outright rejecting it.

Why does it exist?

Trusting incoming data blindly leads to two classes of problems: ordinary bugs (a function crashes because a field it expected to be a number was actually text) and security vulnerabilities (an attacker deliberately sends specially crafted data to manipulate a database query or inject a malicious script that other users will later see). Validation and sanitization exist to catch both at the door, before bad data can do any damage deeper in the system.

When to use it

Validate and sanitize any data that arrives from outside your own trusted backend code — request bodies, query strings, route parameters, uploaded file names — especially before that data touches a database, gets rendered back into a webpage, or gets used to construct a file path.

When not to use it

Data your own backend code generated internally and never exposed to outside input doesn't need the same scrutiny — re-validating data you already fully control adds unnecessary overhead without any real safety benefit.

Common mistakes

  • Validating only on the frontend and assuming the backend never needs to check the same data again.

  • Checking that a field merely exists, without checking its type or shape, letting malformed data slip through.

  • Confusing validation (rejecting bad data) with sanitization (cleaning it up) and using only one when the situation calls for both.

Practice exercises

  1. Easy:

    Write validation for a POST /login route that requires both username and password to be non-empty strings.

  2. Medium:

    Add a check that rejects a signup request if the age field is present but not a positive number.

  3. Hard:

    Explain the difference between rejecting a comment containing HTML tags (validation) and stripping the HTML tags out before saving it (sanitization), and describe a situation where each is the right choice.

Interview questions

What's the difference between validation and sanitization?

Validation checks that data meets expected rules and rejects it if not; sanitization actively cleans or transforms data to remove unsafe parts, rather than outright rejecting it.

Why can't you rely solely on frontend validation?

A request can be sent by anything, not just your own frontend — a malicious or buggy client can bypass frontend checks entirely, so the backend must validate independently.

Why is validating data early in the request lifecycle useful?

It stops bad data before it reaches business logic or a database, preventing crashes and security issues rather than discovering them deeper in the system.

What does a schema-based validation library (like zod or Joi) buy you over hand-written `if` checks?

It lets you declare the expected shape once, in one place, and get consistent parsing, type coercion, and error messages generated automatically, instead of hand-writing and maintaining scattered conditional checks that easily drift out of sync across routes.

What's the difference between a validation library that throws on failure and one that returns a result object, like zod's `safeParse`?

A throwing API requires wrapping every call in try/catch to treat invalid input as control flow; a result-object API returning { success, data | error } lets you check .success directly without exceptions, which fits validation well since invalid input is a common, expected outcome rather than a truly exceptional one.

Why does type coercion, like converting a query-string "42" into the number 42, matter for validation?

Query strings and route parameters always arrive as raw strings over HTTP even when they represent numbers or booleans, so a schema declaring a field as a number needs to convert it, not just check typeof, or every numeric query param would fail validation despite being correct from the caller's point of view.

What's an object-injection risk validation should guard against beyond type checks, e.g. in a MongoDB query built from `req.body`?

Without validation, a field like { password: { "$ne": null } } sent as JSON can be interpreted as a MongoDB query operator instead of a literal value, potentially bypassing an intended equality check — validating that a field is a plain string, not an object, before using it in a query closes this off.

What is HTML/script sanitization specifically protecting against?

Cross-site scripting (XSS) — if user-supplied text containing <script> tags or event handlers is stored and later rendered back into a page without sanitizing or escaping it, that script runs in other users' browsers as if it were part of the trusted site.

What's the difference between escaping and stripping when sanitizing user input meant to be displayed as HTML?

Escaping converts special characters, like < to &lt;, so they display literally as text rather than being interpreted as markup, preserving the original content visually; stripping removes disallowed tags or characters entirely, changing the actual content.

Why doesn't `typeof age !== "number"` alone fully validate a numeric field?

It only rules out non-numeric types — NaN and Infinity are both technically of type "number" in JavaScript, so an extra check like Number.isFinite(age) is needed to also exclude those, if the intent is a genuinely usable numeric value.

What's mass assignment, and how does validation help prevent it?

Blindly passing an entire req.body into a database create or update call lets a client set fields it was never meant to control, like role: "admin"; validating against an explicit schema that only allows specific expected fields, and ignores or rejects the rest, prevents a client from smuggling in extra fields.

Why should validation happen before any database call, rather than letting the database reject bad data on its own?

A database constraint failing, like a NOT NULL violation, still costs a wasted round trip and often surfaces as an unhelpful, generic database error rather than a clear, specific message — validating up front avoids the unnecessary work and gives the caller an immediately actionable response.

What does it mean to validate at the boundary of a system?

Checking data as it enters the system, e.g. as it arrives on a request, so that everything past that boundary — business logic, database calls — can safely assume the data already has the shape it expects, instead of re-checking it everywhere it's used.

Why might you sanitize a filename before using it to write a file to disk?

Path traversal — a filename like ../../etc/passwd can escape an intended upload directory if used directly, so filenames need sanitizing, removing path separators and resolving to a safe base directory, before being used to construct an actual file path.

What's the risk of trusting a simple email-format check, like a string containing `@`, as proof an email address is real and reachable?

It only confirms the format looks plausible — it says nothing about whether the address actually exists or belongs to the person submitting it; confirming a real, owned address requires an actual step like sending a verification email, not just format validation.

How should a validation error response differ from a plain 500, in terms of what it tells the caller?

It should be specific and actionable — which field failed and why, e.g. "email must be a valid email address" — since the caller can fix their own request, unlike a 500, which reflects a server-side problem the caller can't do anything about.

Why is it important to validate the contents of array and object fields, not just that the field itself is an array or object?

A field being 'an array' says nothing about what's inside it — an endpoint expecting an array of numeric ids could still receive one containing strings, objects, or an unexpectedly huge number of items, so the schema needs to validate each element too, not just the outer shape.

What's a practical reason to enforce limits, like max string length or max array size, during validation beyond checking type and shape?

Without limits, a technically valid request — right types, right shape — could still submit an enormous string or array designed to consume excessive memory or processing time, a form of denial-of-service, so validation should also bound the size of what it accepts.