Streaming SSR & Early Flush: Send the First Bytes Before the Last Query

This topic is part of Network & Server Response Optimization. A conventional server-rendered page is built completely before any of it is sent: the server authenticates the user, runs every query, calls every API, renders the full template, and only then writes the response. TTFB equals the slowest dependency plus render time, and during all of that the browser has nothing to do. It cannot fetch the stylesheet, the fonts or the hero image, because it has not seen the <head> that references them.

Streaming reorders the work. The server sends the document head immediately — with stylesheet links, preloads for the LCP image and fonts, and preconnects — then streams the body as each part is ready. The browser starts critical downloads during the server's think time and can paint the first parts of the page while later parts are still being rendered. Early flush is the simplest version: write the head (and perhaps a static shell) before running slow data fetching, then write the rest. Frameworks like React (with Suspense), Next.js, Nuxt, SvelteKit, Remix and Astro provide streaming built in.

Streaming is powerful but fragile. Buffering anywhere in the delivery path — compression middleware, a reverse proxy, a CDN, a serverless adapter — collects the whole response before forwarding it, silently undoing the benefit. And once the first bytes are sent, the status code and headers are fixed, which changes how errors and redirects must be handled.

Buffered SSR vs early flush Timeline comparing when the browser can start fetching critical resources with a buffered response and with an early-flushed head. Buffered SSR vs early flush Buffered server builds whole page CSS + fonts + hero paint Early flush head sent CSS + fonts + hero body rendering paint 0ms 200ms 400ms 600ms 800ms 1000ms 1200ms 1400ms 1600ms LCP flushed LCP buffered

The Metric Degradation This Topic Addresses

Buffered rendering delays TTFB by the full server time and delays FCP and LCP by TTFB plus the download time of critical resources, which could have overlapped with server work. In the LCP breakdown, a long TTFB sub-part followed by a long resource load delay (because the hero image was discovered only when the HTML arrived) is the signature. Streaming shortens TTFB (the first byte is the head) and moves resource discovery earlier, cutting LCP even when total server time is unchanged.

There is a nuance: with streaming, TTFB measures the arrival of the first chunk, which may be much earlier than the meaningful content. LCP and FCP are the metrics that show whether users benefit. A page that flushes a head in 50ms but streams its main content after 2 seconds has a great TTFB and a poor LCP.

Prerequisites

  • Server rendering you control (framework SSR or a custom server).
  • Knowledge of which data each part of the page needs and how long it takes (Server-Timing helps).
  • Visibility into every hop between the server and the browser (proxies, CDN, compression).

1. Environment Setup: Check Whether Responses Stream

bash
# Print elapsed time as each chunk arrives; buffered responses arrive in one burst.
curl -sN https://example.com/product/42 | while IFS= read -r line; do printf '%6.3f %s\n' "$(date +%s.%N | cut -c7-14)" "${line:0:60}"; done | head -20
# trade-off: curl may request uncompressed responses, while browsers request
# compressed ones; test with --compressed too, since compression can buffer.

In DevTools, select the document request and look at the Timing tab: a long "Content Download" relative to "Waiting for server response" indicates streaming; a short download after a long wait indicates buffering.

2. Capture a Baseline

Record TTFB, FCP, LCP and Server-Timing phases for uncached responses on key templates. Note when the LCP resource starts downloading relative to navigation start.

3. Isolate Slow Data From the Shell

Identify which parts of the page depend on slow data: recommendations, reviews, personalised prices, inventory. The head, layout shell, and often the main content can be rendered without them. Those slow parts become streaming boundaries (Suspense boundaries, deferred sections, or later chunks).

4. Apply: Flush the Head, Stream the Rest

Send the head with critical hints immediately, render the shell and fast content next, and stream slow sections as they resolve, with placeholders that reserve space to avoid layout shifts.

javascript
// Node.js: early flush of the head, then stream body sections.
app.get('/product/:id', async (req, res) => {
  res.writeHead(200, { 'Content-Type': 'text/html; charset=utf-8' });
  res.write(headHtml({ lcpImage: `/img/${req.params.id}-1200.avif` }));   // CSS, preload, preconnect
  const productP = getProduct(req.params.id);                               // start work in parallel
  const reviewsP = getReviews(req.params.id);
  res.write(shellHtml());                                                    // header, nav, layout
  res.write(productHtml(await productP));                                    // main content
  res.write(reviewsHtml(await reviewsP));                                    // slower, below the fold
  res.end('</body></html>');
});
// trade-off: if getProduct fails after the head is sent, you cannot return a 404
// or 500 status; render an in-page error and log it, or check existence first.

Streaming rollout workflow Workflow from checking whether responses stream to flushing the head and streaming slow sections. Streaming rollout workflow Check does it stream today? Baseline TTFB, LCP, phases Boundaries slow vs fast parts Flush head hints first Stream sections as ready Verify end-to-end, no buffering

Deconstructing How Browsers Handle Streamed HTML

Browsers parse HTML incrementally. As soon as the head arrives, the preload scanner discovers stylesheets, scripts and preloaded resources and starts fetching them. The parser builds the DOM as body chunks arrive, and the browser can render content once render-blocking CSS has loaded — it does not need the full document. So a streamed page can show its header and main content while the server is still rendering the footer.

Two details matter. First, render-blocking CSS and synchronous scripts still block: if the head references a slow stylesheet, nothing paints until it arrives, streaming or not. Second, frameworks that stream with out-of-order boundaries (React Suspense) send placeholder markup first and later send the real content plus a small inline script that swaps it into place. The content arrives in the same response, but its position in the DOM changes via script, which requires JavaScript and can cause layout shifts if placeholders do not reserve space.

Advanced Diagnostics and Edge Cases

Compression buffering. Gzip and Brotli compressors buffer output to compress efficiently; middleware must flush the compressor on each write. In Node.js, compression exposes res.flush().

Proxies and CDNs. nginx buffers proxied responses by default (proxy_buffering on); disable it or send X-Accel-Buffering: no for streamed routes. Some CDNs buffer responses unless configured for streaming.

Serverless adapters. Some platforms only support buffered responses or require specific streaming APIs.

Status codes and redirects. Decide status codes and redirects before the first flush. Check that the resource exists (a fast query) before sending the head.

Caching streamed responses. CDNs can cache streamed responses; the cached copy is served in full, which is fine — caching already removes server time.

Validation and Budgeting

Validate end to end: the time between the first and last byte of the HTML should be greater than zero in production (from the browser's point of view), the LCP resource request should start well before the HTML finishes, and LCP should improve. Watch CLS: streamed placeholders must reserve space.

CheckTarget
Head flush time (server)< 50 ms after request
LCP resource request startBefore HTML response ends
CLS from streamed sections0 (placeholders sized)
Buffering hopsNone between server and browser

Worked Example: A Product Detail Page

A retailer's product pages were server-rendered with five backend calls; uncached TTFB p75 was 950ms, mostly waiting on reviews (380ms) and recommendations (420ms), which ran after product data. The team flushed the head (with the hero image preload and stylesheet) immediately, ran all calls in parallel, streamed the main product section as soon as product data resolved (about 180ms), and streamed reviews and recommendations later inside reserved containers. TTFB p75 dropped to 90ms (head flush), the hero image started downloading 800ms earlier, and LCP p75 for uncached views fell from 2.9 to 1.9 seconds. Initially, nothing improved in production: the platform's nginx ingress had proxy_buffering on. Disabling buffering for HTML routes made the gains appear.

Streaming With Frameworks

React 18+ streams with renderToPipeableStream (Node) or renderToReadableStream (Web streams), sending HTML up to each Suspense boundary and filling boundaries as their data resolves. Next.js App Router streams by default, with loading.js files and <Suspense> defining boundaries. Nuxt streams server-rendered output with appropriate configuration. SvelteKit streams promises returned from load functions. Astro streams pages by default in server mode, flushing components in order as they resolve. In every case, the framework handles the mechanics; your job is placing boundaries around slow data, keeping the shell fast, and ensuring nothing buffers downstream.

Measuring Streaming in the Field

TTFB alone can mislead with streaming. Track additional timings: the time the LCP resource request starts (from Resource Timing requestStart for the LCP element's URL), responseEnd for the document (when the stream finished), and LCP itself. A good streamed page has early TTFB, an early LCP resource request, and an LCP that comes well before the document finishes. If LCP tracks document responseEnd, the LCP content is in a late chunk — move it earlier.

Timings for an uncached product page (p75) Bar chart of TTFB, LCP resource request start and LCP for buffered versus streamed rendering of the same page. Timings for an uncached product page (p75) Buffered - TTFB 950ms Buffered - LCP 2900ms Streamed - TTFB 90ms Streamed - LCP 1900ms

Streaming and Caching Together

Streaming and CDN caching solve the same problem from different directions, and they combine well. For pages that are the same for many users, caching at the edge removes server time entirely; a cached response arrives in one quick burst and streaming adds nothing. For pages that cannot be fully cached — personalised, real-time or rarely requested — streaming makes the unavoidable server time less visible. Many sites end up with both: a cached, static shell (sometimes called partial prerendering) streamed immediately from the edge, with dynamic sections streamed from the origin into it. Next.js partial prerendering and edge-side composition in other frameworks follow this pattern.

When deciding where to invest, look at the cache hit ratio for HTML per template. Templates with high hit ratios gain little from streaming; templates with low hit ratios and slow backends gain the most.

Designing Placeholders That Do Not Shift

Every streamed section that arrives after the initial paint needs a placeholder, and placeholders are a common source of layout shifts. A good placeholder reserves the final size of its content at each breakpoint: a fixed height for a reviews block based on the typical number of reviews shown, an aspect ratio for media, a fixed row count for recommendation carousels. Skeleton screens should match the layout of the content they stand in for, so the swap looks like content filling in rather than the page rearranging. If content size is genuinely unpredictable, place the section below the fold, where shifts do not affect users reading the top of the page and are less likely to count towards CLS.

Streaming and Third-Party Scripts

Scripts that run at DOMContentLoaded or immediately in the head assume the document is complete. With streaming, DOMContentLoaded fires only after the last chunk has been parsed, so scripts waiting for it start later than with a buffered response of the same total duration — but scripts that run early may find streamed sections missing. Tag managers, A/B testing tools and widgets that query the DOM for elements may need to use mutation observers or framework hooks. Test third-party behaviour on streamed pages specifically, especially anything that modifies content above the fold.

A Rollout Plan

  1. Verify whether responses stream end to end today.
  2. Remove buffering at every hop (compression, proxies, CDN, adapters).
  3. Flush the head with critical hints before any slow work.
  4. Parallelise data fetching and define boundaries around slow sections.
  5. Reserve space for streamed sections to avoid CLS.
  6. Decide status codes before flushing; render in-page errors after.
  7. Measure TTFB, LCP resource start, LCP and CLS in the field.

Common Pitfalls

  • Buffering proxies. Streaming in the app, buffered at the edge.
  • Putting the LCP content in a late boundary. Fast TTFB, slow LCP.
  • Unsized placeholders. Layout shifts when streamed content arrives.
  • Status code surprises. 200 sent before discovering a missing resource.
  • Blocking resources in the flushed head. A slow stylesheet still blocks rendering.

FAQ

What is early flush?

Sending the beginning of the HTML response — usually the head with critical resource hints — before the server finishes rendering the rest, so the browser can start fetching resources.

Does streaming improve TTFB?

Yes, because the first byte arrives as soon as the head is written. But LCP is the metric that shows whether users benefit.

Does streaming work with CDNs?

Most modern CDNs pass streamed responses through, but some buffer by default or for certain features. Test end to end.

Is streaming bad for SEO?

No. Crawlers receive the complete HTML once the stream finishes. Make sure important content is in the HTML rather than only in client-rendered fallbacks.

How do I handle 404s with streaming?

Check whether the resource exists before sending the head, then stream. Errors after the first flush must be shown in the page.

Does compression break streaming?

It can, if the compressor buffers output. Configure compression to flush per chunk.

Should cached pages stream?

Cached responses are already fast; streaming mainly benefits responses that require server work. Streaming does not hurt cached pages.

Can static sites benefit from streaming?

Static files served from a CDN arrive quickly in full, so streaming adds little. The benefit is for responses generated per request.

Does streaming help INP?

Indirectly. Frameworks that stream also tend to hydrate in pieces, which splits main-thread work into smaller tasks and lets early interactions be handled sooner.

Guides in This Topic