Network & Server Response Optimization: From Request to First Byte and Beyond

Every page load starts with a network round trip and a server response, and nothing on the page can happen until the first bytes of HTML arrive. Time to First Byte (TTFB) is the foundation of every other loading metric: First Contentful Paint and Largest Contentful Paint can never be faster than TTFB, and on many sites TTFB alone consumes 40% or more of the LCP budget on mobile. After the first byte, how the HTML and its subresources travel — compressed or not, over which protocol, in one piece or streamed, requested on demand or speculatively ahead of time — determines how quickly the rest of the page follows.

The network layer is also where some of the largest, cheapest improvements hide. Turning on Brotli for text responses can shave 15–25% off HTML, CSS and JavaScript compared with gzip. Removing a redirect from a campaign link removes a full round trip — often 100–300ms on mobile. Streaming HTML lets the browser start fetching CSS and the LCP image while the server is still rendering the body. Speculation rules can make the next navigation effectively instant by prerendering it before the user clicks.

This section is organised around the path of a request. Time to First Byte optimization covers everything before the first byte: redirects, connection setup, server and backend time, cold starts. Compression with Brotli and Zstandard covers how many bytes travel. HTTP/2, HTTP/3 and connection management covers how connections are set up and shared. Streaming SSR and early flush covers sending HTML progressively. Speculative loading and prefetching covers starting the next page before it is requested.

Anatomy of a navigation on a mobile connection Timeline of a typical mobile navigation from redirect and connection setup through server time to the first byte and first paint. Anatomy of a navigation on a mobile connection Network redirect DNS TCP + TLS Server server + backend time Browser HTML + CSS paint 0ms 300ms 600ms 900ms 1200ms 1500ms 1800ms 2100ms 2400ms TTFB FCP

Diagnostic Overview: Where Network Time Goes

Start with field data. TTFB at the 75th percentile, segmented by country, device class and page type, tells you how much of the loading budget the network and server consume before the browser can do anything. Google's guidance considers a TTFB of 800ms or less "good" for most sites, but for a fast LCP on mobile, aiming for well under 600ms at p75 leaves room for everything else. Then break TTFB into its parts using Navigation Timing in RUM or a WebPageTest run:

  • Redirect time. Any 3xx hops before the final URL. Each one is a full round trip, often more on mobile, and can include a new connection to a different host.
  • DNS, TCP and TLS. Connection setup to the origin or CDN. Usually a few round trips; much lower with TLS 1.3, HTTP/3 (QUIC) and connection reuse.
  • Request and server time. From the request leaving the browser to the first byte arriving: network latency plus everything the server does — routing, authentication, database queries, API calls, rendering.
  • Content transfer. After the first byte: how long the HTML takes to arrive, which depends on size, compression and whether the server streams.

Each part has a distinct fix. Redirects are removed at the source (updated links, canonical URLs). Connection time is reduced by CDNs close to users, modern TLS and HTTP/3, and fewer distinct origins. Server time is reduced by caching (at the CDN or in the application), faster backends, and streaming so that the first byte does not wait for the slowest query. Transfer time is reduced by compression and smaller HTML.

Field and lab disagree more for network metrics than for any other. Lab tests run from a few locations on simulated connections; real users connect from everywhere, on networks of every quality, often hitting cold caches and cold connections. A TTFB of 200ms in your CI runs and 1.2s at p75 in the field is common for sites without a CDN in front of HTML, or with low CDN cache hit ratios for documents. Always validate network improvements with field data segmented by geography.

TTFB p75 on mobile by delivery setup (same application) Bar chart of mobile TTFB at the 75th percentile for one application served from different delivery setups. TTFB p75 on mobile by delivery setup (same application) Single-region origin, no CDN for HTML 1380ms CDN proxy, HTML not cached 1020ms CDN-cached HTML + SWR 310ms + HTTP/3 and TLS 1.3 270ms good threshold

Architecture 1: Getting the First Byte Out Fast

TTFB is the sum of everything before the browser receives the first byte of the document. The largest component is usually server time — the application building the response — followed by connection setup and redirects. The most effective improvement is not to build the response at all for most requests: serve HTML from a CDN cache, using s-maxage with stale-while-revalidate so the cache answers instantly while refreshing in the background. For personalised pages, cache the shared shell and fill personal parts at the edge or on the client.

When the response must be built per request, measure where the time goes with the Server-Timing response header: the server reports durations for database queries, API calls, template rendering and cache lookups, and the browser exposes them in DevTools and to RUM scripts. That turns "TTFB is slow" into "the recommendations API call takes 420ms at p75". Fix the slowest dependency, run independent calls in parallel, add application-level caching, or move slow parts out of the critical path by streaming them later.

Serverless and edge platforms add a special case: cold starts. A function that has not run recently must be initialised — loading code, establishing database connections — before handling the request, adding hundreds of milliseconds or more. Bundle size, initialisation work, provisioned concurrency and connection pooling all affect it. Redirect chains are the other classic TTFB inflator: http to https, apex to www, adding a trailing slash, locale detection and campaign tracking redirects can stack into three or four hops. The guides in Time to First Byte optimization cover each.

Architecture 2: Sending Fewer Bytes

Text resources — HTML, CSS, JavaScript, JSON, SVG — compress extremely well. Gzip has been universal for decades; Brotli, supported by all modern browsers, typically produces files 15–25% smaller than gzip for the same content, and more at its highest levels for static assets. Zstandard (zstd), now supported as a Content-Encoding in Chromium-based browsers and Firefox, offers compression ratios between gzip and Brotli with much faster compression speeds, making it attractive for dynamic responses.

Static assets should be compressed once at build time at maximum settings (Brotli level 11), stored alongside the originals, and served directly — the server never compresses them on the fly. Dynamic responses such as HTML from SSR should be compressed on the fly at moderate levels (Brotli 4–6 or zstd 3–6), where the trade-off between CPU time and size is best. Compression dictionary transport goes further: with a shared dictionary (for example, the previous version of a JavaScript bundle), the browser and server can transfer only the differences, cutting repeat-visit downloads of updated bundles by 80–95%.

The guides in Compression with Brotli and Zstandard cover enabling Brotli across servers and CDNs, zstd support, precompression at build time and dictionary transport.

Compressed size of a 480KB JavaScript bundle Bar chart comparing the transferred size of one JavaScript bundle with different compression methods. Compressed size of a 480KB JavaScript bundle Uncompressed 480KB gzip -6 142KB zstd -19 124KB Brotli -11 118KB Brotli with previous-version dictionary 14KB

Architecture 3: Connections and Protocols

Before any bytes can be requested, the browser needs a connection: DNS lookup, a TCP handshake, and a TLS handshake — together two to three round trips on HTTP/1.1 and HTTP/2 over TLS 1.2, fewer with TLS 1.3, and as few as one with HTTP/3 over QUIC (or zero for resumed connections). On mobile networks with 100–200ms round-trip times, connection setup alone can take half a second.

HTTP/2 multiplexes many requests over one connection, which made old HTTP/1.1 workarounds — domain sharding, spriting, concatenating everything — unnecessary and often harmful: each extra domain needs its own connection setup and splits prioritisation. HTTP/3 replaces TCP with QUIC over UDP, which removes TCP-level head-of-line blocking (a lost packet no longer stalls all streams), handles network changes (Wi-Fi to cellular) without reconnecting, and reduces handshake round trips. Its benefits are largest on poor networks: high latency and packet loss.

Connection management matters as much as protocol choice. Each third-party origin on the critical path costs a new connection; preconnect can start it early, but removing or self-hosting critical third-party resources is better. Browsers can reuse (coalesce) an HTTP/2 or HTTP/3 connection across hostnames that share an IP address and a certificate covering both names, which turns separate asset domains into the same connection. The guides in HTTP/2, HTTP/3 and connection management cover measuring HTTP/3 gains, domain sharding, coalescing and head-of-line blocking.

Architecture 4: Streaming HTML

Traditional server rendering waits for all data, renders the whole page, then sends it. TTFB equals the slowest query plus render time, and the browser sits idle the whole time. Streaming changes the order: send the document head (with links to CSS, preloads for the LCP image and fonts) immediately, then stream the body as it is rendered, with slow sections arriving later. The browser starts downloading critical resources while the server is still working, and it can paint the first parts of the page before the last ones arrive.

Modern frameworks support streaming natively: React's streaming SSR with Suspense boundaries, Next.js App Router loading states, Nuxt and SvelteKit streaming, Astro's streaming responses. Even without a framework, flushing the head early — writing <head> with critical hints before running slow queries — gives much of the benefit. The main obstacles are intermediaries that buffer the response: some CDNs, proxies, compression middleware and serverless platforms collect the whole response before forwarding it, silently disabling streaming. The guides in Streaming SSR and early flush cover React Suspense streaming, early head flushing and keeping streams intact through CDNs.

Buffered vs streamed server rendering Timeline comparing a buffered server-rendered response with a streamed one that flushes the head early. Buffered vs streamed server rendering Buffered SSR queries + render (browser idle) CSS + LCP image paint Streamed SSR head CSS + LCP image body streams paint 0ms 200ms 400ms 600ms 800ms 1000ms 1200ms 1400ms 1600ms 1800ms LCP streamed

Architecture 5: Loading the Next Page Before It Is Requested

The fastest navigation is one that has already happened. Speculative loading uses idle time and user intent to fetch, or fully render, likely next pages. The Speculation Rules API lets a page declare which URLs to prefetch or prerender, either as a list or with document rules that match links on the page, and at what eagerness: immediately, when the user hovers or presses a link, or conservatively on pointer down. A prerendered page is loaded and rendered in a hidden tab; activating it takes only a few milliseconds, and LCP for that navigation is close to zero.

Speculation has costs. Every prefetch uses bandwidth and server capacity, and every prerender also uses CPU and memory on the user's device and runs the page's JavaScript — including analytics, which must not count a visit that never happened. Good speculation strategies target high-probability navigations (the next article in a series, the top search result, the product a user hovers over), respect data saver modes, and use eagerness settings that balance hit rate against waste. Libraries like Quicklink prefetch links as they enter the viewport during idle time, an approach that works across browsers without speculation rules support.

The guides in Speculative loading and prefetching cover speculation rules, the prerender versus prefetch trade-off, viewport prefetching and analytics correctness.

Prioritising Network Work

With so many levers, order matters. A practical sequence for most sites is: first remove redirects on key entry URLs (cheap, large win per affected visit); second, cache HTML at the CDN wherever content allows (the biggest TTFB win); third, make sure Brotli is on for every text response and static assets are precompressed (a configuration change); fourth, enable HTTP/3 and TLS 1.3 at the CDN (usually a toggle); fifth, stream or early-flush uncached responses; and finally, add speculation rules for the most common navigation paths. Each step is measurable in the field on its own, so roll them out one at a time and confirm the effect on TTFB, FCP and LCP before moving on. Teams that change everything at once rarely learn which change mattered, and cannot tell which one caused a regression. Keep a short changelog of infrastructure changes next to your RUM dashboards, so every step in the TTFB trend line has an explanation.

Monitoring and CI: Holding the Network Layer to Budget

Network performance regresses for reasons outside the codebase: a CDN configuration change disables HTML caching, a new marketing redirect is added, a backend dependency slows down, a proxy starts buffering responses. Monitoring must cover infrastructure, not just builds.

In the field, collect TTFB and its sub-parts (redirect, DNS, connect, request-to-response) from Navigation Timing in your RUM, along with Server-Timing entries and the nextHopProtocol (h2 or h3). Segment by country and connection type. Alert on p75 TTFB per key template. Track CDN cache hit ratio for HTML separately from assets — it is often the single best predictor of TTFB.

In CI and synthetic monitoring, assert on response headers: Content-Encoding for text resources, Cache-Control for HTML and assets, alt-svc for HTTP/3 advertisement, absence of redirects for key entry URLs, and streaming behaviour (time from first byte to last byte for the HTML should be longer than zero when streaming is expected). These checks are cheap and catch the most common regressions.

Reference Implementations

Server-Timing for backend breakdown

javascript
// Express middleware: report timings for each phase of the request.
app.use((req, res, next) => {
  const marks = [];
  res.locals.time = async (name, fn) => {
    const t0 = performance.now();
    try { return await fn(); } finally { marks.push(`${name};dur=${(performance.now() - t0).toFixed(1)}`); }
  };
  const writeHead = res.writeHead;
  res.writeHead = function (...args) { if (marks.length) res.setHeader('Server-Timing', marks.join(', ')); return writeHead.apply(this, args); };
  next();
});
// In a route: const user = await res.locals.time('db-user', () => db.user(id));
// trade-off: Server-Timing is visible to anyone who loads the page; avoid exposing
// sensitive internals, or only emit it for sampled or authenticated requests.

Caching HTML at the CDN with background revalidation

http
Cache-Control: public, max-age=0, s-maxage=300, stale-while-revalidate=86400

Browsers always revalidate (max-age=0), while the CDN serves cached HTML for five minutes and keeps serving the stale copy for up to a day while fetching a fresh one in the background.

Brotli for static and dynamic responses (nginx)

nginx
brotli on;
brotli_comp_level 5;                     # dynamic responses: balance CPU and size
brotli_types text/html text/css application/javascript application/json image/svg+xml;
brotli_static on;                        # serve prebuilt .br files when present
gzip on; gzip_static on;                 # fallback for clients without br
# trade-off: level 11 is too slow for on-the-fly compression; precompress static
# assets at 11 during the build and keep dynamic compression at 4-6.

Flushing the head early (Node.js)

javascript
app.get('/product/:id', async (req, res) => {
  res.setHeader('Content-Type', 'text/html; charset=utf-8');
  res.write(`<!doctype html><html><head><link rel="stylesheet" href="/css/app.css">
    <link rel="preload" as="image" href="/img/${req.params.id}-hero.avif" fetchpriority="high"></head><body>`);
  const product = await getProduct(req.params.id);       // slow work after the flush
  res.end(`${renderProduct(product)}</body></html>`);
});
// trade-off: once the head is sent, the status code and headers are fixed; errors
// discovered later must be rendered as in-page error states, not HTTP errors.

Speculation rules for likely navigations

html
<script type="speculationrules">
{ "prerender": [{ "where": { "href_matches": "/articles/*" }, "eagerness": "moderate" }],
  "prefetch":  [{ "where": { "href_matches": "/*" }, "eagerness": "conservative" }] }
</script>
<!-- trade-off: moderate eagerness prerenders on hover (~200ms), which wastes some
     work on links users hover but do not click; measure hit rates per pattern. -->

Common Pitfalls

  • Treating TTFB as a server-only metric. Redirects, DNS, TLS and CDN configuration are often larger contributors than application code.
  • Never caching HTML. The single biggest TTFB win for most content sites is CDN-cached HTML with background revalidation.
  • Compressing static assets on the fly at low levels. Precompress at build time at maximum levels instead.
  • Compressing already-compressed formats. JPEG, AVIF, WOFF2 and video gain nothing and waste CPU.
  • Domain sharding on HTTP/2. Extra connections and broken prioritisation.
  • Streaming that silently buffers. A proxy or middleware collects the whole response; verify with timing, not assumptions.
  • Prerendering without analytics guards. Inflated page views and skewed metrics.
  • Speculating on every link at high eagerness. Wasted bandwidth and server load for little gain.

FAQ

What is a good TTFB?

Google considers 800ms or less at the 75th percentile good. For a fast LCP on mobile, aim lower — under about 600ms — because TTFB consumes part of the 2.5-second LCP budget.

Is TTFB a Core Web Vital?

No, but it is a foundational diagnostic metric. LCP and FCP cannot be faster than TTFB, so slow TTFB caps every loading metric.

Should I use Brotli or Zstandard?

Use Brotli for precompressed static assets, where its higher ratio matters most. For dynamic responses, Brotli at moderate levels and zstd are both good choices; zstd compresses faster at similar ratios. Keep gzip as a fallback.

Does HTTP/3 make my site faster?

Mostly for users on high-latency or lossy networks — mobile users, distant regions. On fast, stable connections the difference is small. Measure by protocol in RUM.

Does streaming SSR help if my CDN caches HTML?

Cached HTML is already fast, so streaming matters most for uncached or personalised responses where the server must do real work for each request.

Is prerendering safe for all pages?

No. Avoid prerendering pages with side effects on load (adding items to carts, logging out, one-time tokens) and pages that are expensive to render. Guard analytics so prerendered pages only count when activated.

How much do redirects cost?

Each redirect costs at least one round trip, and more if it switches hosts and needs a new connection. On mobile networks, that is typically 100–300ms per hop.

Topics in This Section