Web Performance Basics

What actually makes a page feel slow to use, and the basic techniques for making it feel fast.

What is it?

A page can feel slow for a few very concrete reasons, and most of them come down to the browser being kept busy or waiting before it can show or use something. A big stylesheet or script that has to fully download and run before the browser can display anything is called render-blocking — the page sits blank while the browser waits on it. Large, unoptimized images can take a long time to download, delaying everything on the page that depends on them. And a large amount of JavaScript, even after it's downloaded, still takes time for the browser to parse and run, which can leave a page looking finished but unresponsive to clicks and taps.

There are a handful of well-established techniques for addressing each of these. Lazy loading means not fetching something (like an image far down the page) until it's actually about to be needed, instead of loading everything up front. Minification strips unnecessary characters (whitespace, long variable names) out of CSS and JavaScript files before they're sent, so there's less data to transfer. And caching lets a browser reuse a file it already downloaded on a previous visit, instead of re-downloading something that hasn't changed.

Explain like I'm 10

It's like being handed a phone book at the door before you're allowed into a restaurant, when all you needed was the menu. Render-blocking resources make you wait through unrelated work before you get the one thing you actually came for.

Examples

Avoiding render-blocking scripts

<!-- Blocks rendering until fully downloaded and run -->
<script src="analytics.js"></script>

<!-- Downloads in the background, runs after parsing finishes -->
<script src="analytics.js" defer></script>

Without defer, the browser stops parsing the rest of the HTML to fetch and run the script immediately. With defer, the script downloads in parallel and only runs once the page has finished parsing, so it no longer delays the visible page.

Lazy loading images

<img src="hero.jpg" alt="Product hero image">
<img src="footer-banner.jpg" alt="Seasonal promotion" loading="lazy">

The hero image loads immediately since it's visible right away. The footer banner, likely far below the visible area when the page first loads, is marked loading="lazy" so the browser only fetches it once the user scrolls close to it.

How it works

The browser can only do so much at once: parsing HTML, downloading files, parsing and running CSS and JavaScript, and rendering the result all compete for its attention. Techniques like deferring scripts, lazily loading offscreen images, and minifying files all work by reducing or rescheduling that competing work — either by shrinking how much data has to move over the network, or by making the browser wait less before it can show something useful. Caching works differently: it avoids the network entirely for files the browser recognizes it already has an unchanged copy of.

Why does it exist?

Performance work exists because a slow page directly costs user attention and, for businesses, real measurable outcomes — people abandon slow-loading pages far more readily than fast ones. As the average webpage has grown heavier over time (more images, more scripts, more third-party embeds), these techniques became necessary just to keep pages usable rather than optional polish.

When to use it

Apply these techniques by default on any real-world website: defer non-critical scripts, lazy-load offscreen images, minify production CSS and JavaScript, and set caching headers on assets that don't change often. Pay closer attention to performance whenever users report a page feeling slow, or whenever measuring actual load times reveals a problem.

When not to use it

Extremely small or low-traffic internal tools may not need aggressive performance tuning — the cost of, say, setting up a whole build pipeline for minification might outweigh the benefit if load time is already imperceptible. Optimizing prematurely, before knowing where real time is actually being spent, can also waste effort on the wrong thing.

Common mistakes

  • Loading large, unoptimized images at full resolution when a much smaller version would look identical on screen.

  • Marking a critical, above-the-fold image as lazy-loaded, delaying content the user sees immediately.

  • Assuming minification alone fixes a performance problem caused by simply shipping too much JavaScript in the first place.

Practice exercises

  1. Easy:

    Add the loading="lazy" attribute to images that appear below the visible area on a sample page.

  2. Medium:

    Take a page with a blocking <script> tag in the head and change it to use defer, then explain the difference in page load behavior.

  3. Hard:

    Using your browser's Network panel, measure a real page's load time, identify the largest resource slowing it down, and propose one concrete fix.

Interview questions

What does 'render-blocking' mean?

A resource, like a script or stylesheet, that the browser must fully download (and sometimes run) before it can continue rendering the page, causing a visible delay.

What is lazy loading, and when is it appropriate?

Deferring the loading of a resource, commonly an image, until it's actually about to be needed, such as when it scrolls near the visible viewport. It's appropriate for offscreen content, but not for content visible immediately on page load.

How does caching improve web performance?

It lets a browser reuse a previously downloaded file instead of re-fetching it from the network, as long as the file hasn't changed, saving both time and bandwidth on repeat visits.

What are the Core Web Vitals, and which three metrics make up the current set?

They're a specific set of metrics Google uses to measure real-world user experience: Largest Contentful Paint (LCP, loading speed), Interaction to Next Paint (INP, responsiveness), and Cumulative Layout Shift (CLS, visual stability).

What does Largest Contentful Paint (LCP) measure, and what's considered a good score?

It measures how long it takes for the largest visible content element (often a hero image or a large block of text) to finish rendering after the page starts loading. A score of 2.5 seconds or less is generally considered good.

What does Cumulative Layout Shift (CLS) measure, and what commonly causes a poor score?

It measures how much visible content unexpectedly shifts position during page load, weighted by how much of the viewport moved and how far. Common causes include images or embeds without reserved width/height, content (like an ad or banner) injected above existing content after load, and web fonts swapping in with different metrics than their fallback.

What replaced First Input Delay (FID) as a Core Web Vital, and why?

Interaction to Next Paint (INP) replaced it. FID only measured the delay before the first interaction started being processed, missing responsiveness problems later in a page's life. INP instead measures responsiveness across every interaction throughout the page's lifecycle, giving a more complete picture.

What are the main stages of the critical rendering path?

The browser parses HTML into a DOM, parses CSS into a CSSOM, combines the two into a render tree of only the visible nodes with their computed styles, calculates each node's position and size (layout), and finally paints pixels to the screen.

Why is CSS treated as render-blocking by default?

The browser deliberately waits for the full CSSOM to be built before rendering anything, to avoid painting unstyled content and then immediately repainting it once styles arrive — a jarring flash of unstyled content (FOUC) that blocking avoids by holding off the first paint until styling is known.

What's the tradeoff between putting a `<link rel="stylesheet">` in the `<head>` versus at the end of `<body>`?

In the head, the browser blocks rendering until the CSS is loaded, delaying first paint but guaranteeing the page never appears unstyled. At the end of the body, the browser may paint unstyled content briefly before the stylesheet loads and reflows everything, which usually looks worse even though something technically appears on screen sooner.

What's the difference between `async` and `defer` on a `<script>` tag?

Both let the script download without blocking HTML parsing. async runs the script the instant it finishes downloading, which can interrupt parsing and doesn't guarantee execution order relative to other scripts. defer instead waits until the HTML is fully parsed, and multiple deferred scripts always run in their original document order — making defer the safer default for scripts that depend on the DOM or on each other.

Why is reflow (layout) generally more expensive than repaint?

A repaint just redraws pixels that changed appearance (like a color) without affecting geometry. A reflow recalculates the position and size of the changed element and potentially every element affected by that change, which can cascade through large parts of the page, making it considerably more computationally expensive.

What is 'layout thrashing', and what's a common way it happens by accident?

It's when code repeatedly forces the browser to synchronously recalculate layout many times in a tight loop, rather than once. A classic cause is reading a layout-dependent property (like element.offsetHeight) immediately after writing a style change, inside a loop — each read forces the browser to flush and recompute layout before it can answer, over and over.

Why can overusing `will-change` hurt performance instead of helping it?

will-change is a hint that tells the browser to preemptively promote an element onto its own compositing layer in anticipation of an animation. Applying it broadly or to elements that don't actually need it creates many extra layers, consuming more GPU memory and potentially slowing things down rather than speeding them up.

When would you still need an Intersection Observer-based lazy-loading approach instead of native `loading="lazy"`?

When you need more control than the browser default gives you — like lazy-loading something other than images/iframes (e.g. a heavy component), customizing how far in advance loading starts, or supporting older browsers that don't implement the loading attribute.

What is code splitting, and how does it help a large JavaScript application load faster?

It breaks a single large JavaScript bundle into smaller chunks that are loaded on demand — for example, per route — so a user's first visit only downloads and parses the code needed for the page they're actually viewing, instead of the entire application upfront.

What does tree shaking remove from a JavaScript bundle?

Dead code — exports from a module that are never actually imported or used anywhere in the final application — based on static analysis of import/export statements, so the shipped bundle doesn't include library code the app never calls.

What's the difference between `<link rel="preload">`, `rel="prefetch"`, and `rel="preconnect"`?

preload tells the browser to fetch a resource it will definitely need for the current page, with high priority, sooner than it would discover it naturally. prefetch fetches something likely needed for a future navigation, at low priority. preconnect just establishes the network connection (DNS, TCP, TLS) to another origin ahead of time, without fetching a specific file, to save that setup time once a request to it is actually made.

Your page's LCP element is a large hero image loading late. Name two concrete ways to improve it.

Preload the image with <link rel="preload" as="image"> so the browser fetches it earlier instead of discovering it only after parsing reaches the img tag, and serve it as an appropriately sized, modern-format (e.g. WebP/AVIF) file so there's simply less data to download before it can render.

Does minifying JavaScript reduce how long the browser takes to execute it?

Not meaningfully — minification shrinks file size (less to download and parse), but the actual logic still runs the same number of operations, so execution time stays roughly the same. This is why shipping too much JavaScript in the first place can't be fixed by minification alone.

What is `font-display: swap`, and what performance/visual tradeoff does it introduce?

It tells the browser to render text immediately in a fallback font while a custom web font is still downloading, then swap to the real font once it arrives, avoiding invisible text during the wait. The tradeoff is that if the fallback and web font have different metrics (character widths, line height), the swap itself can cause a visible layout shift, hurting CLS.

Why does compressing assets with gzip or Brotli improve load performance, and which one usually compresses better?

Both reduce the number of bytes that actually travel over the network for text-based assets like HTML, CSS, and JS, so they download faster. Brotli generally achieves a better compression ratio than gzip for the same content, though it can take slightly longer to compress, which matters more for build time than for the already-compressed file being served.

How did HTTP/2 multiplexing reduce the value of older performance tricks like bundling many files together or sharding assets across multiple domains?

HTTP/1.1 could only send a limited number of requests in parallel per connection, so bundling files (fewer requests) and domain sharding (more parallel connections) worked around that limit. HTTP/2 multiplexes many requests over a single connection simultaneously, removing much of that per-request overhead, so aggressively bundling everything into one giant file can actually hurt caching efficiency without the payoff it used to have.

Using browser DevTools, how would you tell whether a slow page load is bottlenecked by network transfer or by JavaScript execution?

The Network panel's waterfall shows how long each resource actually spent downloading versus waiting, which points to a network bottleneck if requests are large or slow. The Performance panel's main-thread flame chart instead shows time spent parsing/executing/compiling JavaScript and running layout/paint, which points to a CPU-bound bottleneck if that's where the time is concentrated.

An `<img>` has no `width`/`height` attributes set. What layout symptom can result once it finishes loading on a slow connection, and how do you prevent it?

Before the image loads, the browser doesn't know its dimensions and reserves no space for it, so surrounding content occupies that area temporarily; once the image arrives, everything below it jumps to make room — a layout shift counted against CLS. Setting explicit width/height attributes (or a CSS aspect-ratio) lets the browser reserve the correct space upfront, before the image has even started downloading.

What's the Core Web Vitals tradeoff between client-side rendering (CSR) and server-side rendering (SSR)?

SSR typically produces better LCP/FCP on first load, since the browser receives HTML with real content already in it rather than an empty shell waiting on JavaScript to fetch and render data. CSR often has a slower first load for that reason, but can feel faster on subsequent in-app navigations once the JavaScript bundle is already loaded and only data needs fetching.

What does Time to First Byte (TTFB) measure, and what does a slow TTFB usually indicate?

It measures the time from a request being sent until the first byte of the response arrives, which mostly reflects server-side work — routing, database queries, server-side rendering — rather than front-end asset delivery. A slow TTFB points to a backend or network-latency problem, not something fixable with front-end techniques like minification or lazy loading.

Mechanically, how does a CDN improve load performance?

It caches copies of static assets on servers geographically distributed closer to users, so a request is served from a nearby edge location instead of traveling all the way to a single origin server — reducing the network latency (round-trip time) involved in fetching the resource.

A returning visitor loads a page much faster than a first-time visitor loading the exact same assets. What's the likely mechanism, given proper `Cache-Control` headers?

The browser's disk/memory cache already holds a copy of assets like CSS, JS, and images from the previous visit, and Cache-Control headers (e.g. a max-age) tell it those files are still valid, so it reuses them straight from cache instead of re-requesting them from the network at all — the first-time visitor has no such cache and must download everything.

What's a common pitfall of giving a file like `app.js` a very long `Cache-Control: max-age`?

After deploying a new version, users with a cached copy keep using the stale old file until the cache expires, potentially running outdated (or broken, if it no longer matches a changed backend) code for a long time. The standard fix is to include a content hash in the filename (e.g. app.3f2a1c.js), so a new deploy produces a new filename and is fetched fresh immediately, while the long cache lifetime remains safe since a given filename's content never changes.