Streaming SSR with React Suspense

This guide is part of Streaming SSR & Early Flush, within Network & Server Response Optimization. React's streaming server renderer (renderToPipeableStream in Node.js, renderToReadableStream for Web streams) sends HTML progressively. Everything outside Suspense boundaries — the shell — is sent as soon as it can be rendered. Each <Suspense> boundary whose children are waiting for data sends its fallback first; when the data resolves, React streams the real HTML for that boundary along with a tiny inline script that replaces the fallback in place.

On the client, React hydrates boundaries independently (selective hydration): parts of the page become interactive as their code and HTML arrive, and React prioritises hydrating the boundary a user interacts with. Together, streaming and selective hydration improve both LCP (content arrives earlier) and INP (hydration work is split and prioritised). The main design decision is where to put boundaries: too few and slow data blocks the whole page; boundaries around LCP content delay LCP.

How a Suspense boundary streams Sequence showing the server sending the shell with a fallback, then the resolved boundary content and an inline swap script. How a Suspense boundary streams Server Browser shell HTML + fallback boundary HTML (hidden) inline swap script

Rapid Diagnosis

  • Check the server renderer: renderToString buffers everything; renderToPipeableStream or renderToReadableStream streams.
  • Look at the HTML stream: fallbacks followed later by <template>/hidden div chunks and $RC scripts indicate streaming boundaries.
  • Find the LCP element and check whether it is inside a boundary that waits for data.
  • Check CLS around fallbacks when content swaps in.

Root Cause Analysis

1. Buffered rendering. renderToString waits for all data.

2. No boundaries. Data fetched at the top of the tree blocks the whole render.

3. Boundaries around LCP content. The hero waits for slow data and arrives late.

4. Unsized fallbacks. Swapping content shifts the layout.

Step-by-Step Resolution

1. Use the streaming renderer

javascript
import { renderToPipeableStream } from 'react-dom/server';
app.get('*', (req, res) => {
  let didError = false;
  const { pipe } = renderToPipeableStream(<App url={req.url} />, {
    bootstrapScripts: ['/assets/client.js'],
    onShellReady() {
      res.statusCode = didError ? 500 : 200;
      res.setHeader('Content-Type', 'text/html; charset=utf-8');
      pipe(res);                                   // stream the shell now, boundaries later
    },
    onShellError() { res.statusCode = 500; res.end('<h1>Something went wrong</h1>'); },
    onError(err) { didError = true; console.error(err); },
  });
});
// trade-off: onShellReady streams as soon as the shell is ready (best for users);
// onAllReady waits for everything, which suits crawlers that need full HTML at once.

2. Keep LCP content in the shell

Fetch the data for the hero and main content before or during shell rendering (fast queries, cached data), and wrap only slow, secondary sections (reviews, recommendations, personalised widgets) in Suspense.

3. Size fallbacks to match content

javascript
<Suspense fallback={<div className="reviews-skeleton" style={{ minHeight: 480 }} aria-busy="true" />}>
  <Reviews productId={id} />
</Suspense>
// trade-off: fixed minimum heights avoid CLS but may leave empty space if content
// is shorter; use typical heights per breakpoint.

4. Split client code per boundary

Use React.lazy for components inside boundaries so their JavaScript loads separately, enabling selective hydration to make the shell interactive before slow sections' code arrives.

Where to place Suspense boundaries on a product page Layers of a product page and whether each should be in the streamed shell or behind a Suspense boundary. Where to place Suspense boundaries on a product page Head, nav, layout Shell — sent immediately Product title, price, hero image Shell — LCP content, never behind slow data Stock and delivery estimate Boundary — fast API, small fallback Reviews Boundary — slow, sized skeleton Recommendations Boundary — slowest, below the fold

Verification

View the raw response stream (curl with timing) and confirm the shell arrives quickly and boundaries arrive later. In the browser, LCP should occur from shell content, before slow boundaries resolve. CLS should stay near zero as boundaries swap in. In the Performance panel, hydration should appear as several smaller tasks rather than one large one.

Worked Example: A Marketplace Listing Page

A marketplace moved its listing pages from renderToString to renderToPipeableStream. Initially, it wrapped the entire page content in one Suspense boundary, so the stream sent a header and a big skeleton, then everything else 1.1 seconds later; LCP did not improve. The team restructured: listing title, price and primary image came from a fast cached endpoint and rendered in the shell; seller information, reviews and similar listings each got their own boundaries with sized skeletons. LCP p75 fell from 2.8 to 1.7 seconds, and INP p75 improved from 260ms to 180ms because hydration was split across boundaries and prioritised the area users tapped first.

LCP p75 on listing pages by boundary design Bar chart comparing LCP for buffered rendering, one large Suspense boundary, and targeted boundaries with LCP content in the shell. LCP p75 on listing pages by boundary design renderToString (buffered) 2800ms Streaming, one big boundary 2750ms Streaming, LCP in shell + 3 boundaries 1700ms LCP good

Crawlers and onAllReady

Some sites detect crawlers and use onAllReady so bots receive the complete HTML in one piece, without inline swap scripts. Modern search crawlers handle streamed HTML and run JavaScript, but sending full HTML to bots is a common, harmless choice. Make sure user-facing performance is not affected by this branch: detect bots by user agent only for this purpose, and keep the same content.

Data Fetching Patterns That Stream Well

Streaming works best when slow data is requested early and awaited late. Start fetches at the top of the request (or in route loaders) and pass promises down to the components that need them, which suspend only where the data is used. Avoid "waterfalls" where a child component starts its fetch only after its parent has rendered, which serialises latency across boundaries. In React Server Components, async components inside Suspense boundaries naturally follow this pattern; in client-rendered data libraries, prefetch on the server and hydrate the cache so the client does not refetch.

Common Mistakes

  • One giant boundary. Equivalent to buffering behind a skeleton.
  • LCP content behind a boundary. Streaming improves TTFB but not LCP.
  • Data fetching at the root. A top-level await blocks the shell.
  • Fallbacks without dimensions. CLS when content arrives.

Edge Cases

Error boundaries. Errors inside a streamed boundary render the nearest error boundary's fallback on the client; the status code is already sent.

Head tags from boundaries. Metadata decided inside slow boundaries cannot affect the already-sent head; resolve titles and meta in the shell.

Third-party scripts. Scripts that expect the full DOM at DOMContentLoaded may run before streamed content exists; use observers or framework hooks.

Buffering infrastructure. All hops must pass the stream through; see the CDN and proxy guide.

FAQ

What is the difference between renderToString and renderToPipeableStream?

renderToString renders the whole tree synchronously and returns a string; it cannot wait for Suspense data. renderToPipeableStream streams the shell and Suspense boundaries as they resolve.

What is selective hydration?

React hydrates Suspense boundaries independently and prioritises those the user interacts with, so parts of the page become interactive before the whole page has hydrated.

Does streaming require JavaScript on the client?

Out-of-order boundary content is swapped in by inline scripts. Without JavaScript, users see fallbacks for streamed boundaries; shell content is unaffected.

Where should Suspense boundaries go?

Around slow, secondary sections. Keep the LCP content and above-the-fold essentials in the shell, backed by fast or cached data.

Does Next.js do this automatically?

The App Router streams by default; loading.js and <Suspense> define boundaries. The same placement rules apply.

How do I avoid CLS from fallbacks?

Give fallbacks dimensions that match the typical content size, using min-height or aspect ratios.

Can I set cookies or headers after streaming starts?

No. Headers are sent with the shell. Decide cookies, status and caching headers before onShellReady pipes the response.

Does streaming increase total server time?

Not meaningfully. Total work is similar; it is reordered so that users see and can use parts of the page sooner.

How many Suspense boundaries should a page have?

Usually a handful, one per independent slow section. Very many tiny boundaries add overhead and visual churn; one huge boundary removes the benefit.

Does streaming work with React Server Components?

Yes. Server Components render on the server and stream into Suspense boundaries the same way, which is how the Next.js App Router streams pages.