Edge Compute & Dynamic Caching: Fast TTFB for Pages That Look Uncacheable

This topic extends Advanced Caching Strategies & CDN Architecture to the pages a traditional CDN cannot cache: those that vary by user, by experiment, by location or by session. A static marketing page is easy — cache it at the edge and TTFB drops to tens of milliseconds. A product page showing a logged-in user's name, a personalised recommendation rail and a price in their currency usually goes all the way to the origin for every request, and its TTFB is whatever the origin manages: often 400–1000ms at p75 once you add distance, cold caches and database queries.

Edge compute — Cloudflare Workers, Fastly Compute, Vercel and Netlify edge functions, CloudFront Functions and Lambda@Edge, Akamai EdgeWorkers — runs your code in the CDN's points of presence, next to the cache. That changes what is cacheable. The page can be split into a shared, cacheable shell and small personalised fragments; the edge can choose a cached variant by cookie or header without fragmenting the cache; and per-user data can be fetched in parallel with streaming the cached shell. Done well, pages that "had" to be dynamic get the same sub-200ms TTFB as static ones.

Origin-rendered vs edge-assembled personalised page Comparison of a personalised page rendered at the origin for every request with one assembled at the edge from a cached shell and small personal fragments. Origin-rendered vs edge-assembled personalised page Origin-rendered (no cache) • Every request travels to the origin • Full render + DB queries per request • TTFB 600-1000ms at p75 • Origin load scales with traffic Edge-assembled • Shared shell cached at every PoP • Personal fragment fetched in parallel • TTFB 60-150ms at p75 • Origin serves small fragments only

The Metric Degradation This Topic Addresses

TTFB is the first phase of every LCP. For server-rendered pages, slow TTFB pushes everything else later: if TTFB is 800ms at p75, the LCP budget of 2.5s has already lost a third. Field data typically shows a clear split:

  • Anonymous traffic on cached pages: TTFB 50–150ms.
  • Logged-in traffic, personalised pages, or pages with Vary: Cookie: TTFB 400–1200ms, because they bypass the cache entirely.
  • Pages with query-string variations or A/B test cookies: cache hit rates that look fine in aggregate but collapse for the variants that matter.

The goal of this topic is to move the second and third groups into the first, without serving anyone the wrong content.

TTFB p75 by request type before edge caching Bar chart of p75 time to first byte for anonymous cached pages, logged-in pages, A/B-tested pages and pages with currency variants. TTFB p75 by request type before edge caching Anonymous, cached 95ms Logged-in (bypass) 780ms A/B variant cookie 640ms Currency by geo 520ms 200ms target

Prerequisites

  • An edge compute platform in front of your origin, with access to its cache API (Cloudflare's caches.default and cf options, Fastly's surrogate controls, Vercel's edge cache headers, and so on).
  • A clear content inventory per template: which parts are the same for everyone, which vary by a small set of values (country, currency, experiment variant), and which are truly per user.
  • Origin endpoints for fragments — small JSON or HTML endpoints for the personalised parts.
  • Cache-key discipline: an understanding of what currently varies the cache key (cookies, headers, query parameters) — see normalizing cache keys to raise hit rate.

1. Environment Setup: Classify Every Byte of the Page

For each template, mark regions as shared (identical for all users), variant (one of a few versions, chosen by a low-cardinality key such as country or experiment arm), or personal (unique per user). Most pages are overwhelmingly shared: on a typical product page, everything except the header's account menu, the cart count and a recommendation rail.

2. Capture a Baseline

Record TTFB p75 and cache hit ratio per template and per request type (anonymous vs logged-in) from RUM with Server-Timing headers carrying cache status. This tells you how much traffic currently bypasses the cache and how slow it is.

3. Isolate What Forces Bypass

Usually one or two things push a whole page out of the cache: a session cookie read during rendering, a personalised header widget, or a variant decided at the origin. Find them in the origin code — every request.cookies read in a page render is a candidate.

Anatomy of an edge-cacheable product page Regions of a product page classified as shared, variant or personal, with how each is served at the edge. Anatomy of an edge-cacheable product page Shared shell layout, product details, images, reviews — cached for everyone, purged by tag Variant regions price in currency, experiment arm — cached per variant key (few values) Personal fragments account menu, cart count, recommendations — fetched per user, not cached Assembly edge function stitches or streams; client fetch as a fallback

4. Apply: Split, Key and Assemble

The fixes, in order of how much traffic they usually recover:

  1. Stop varying the whole page on cookies. Render the shared shell without reading session cookies; serve personal bits separately. This alone often moves logged-in traffic onto the cache.
  2. Key variants explicitly at the edge. Derive a small variant key (country → currency, experiment cookie → arm) in the edge function and add it to the cache key, instead of Vary: Cookie.
  3. Assemble personal fragments. Either at the edge — fetching fragments in parallel and streaming them into the cached shell (see edge side includes vs client-side fragments) — or on the client after the shell paints.
  4. Personalise at the edge only what must be in the first paint, such as a localised hero, as in personalizing cached pages at the edge.

Deconstructing TTFB for an Edge-Assembled Page

PhaseOrigin-renderedEdge-assembledLever
Client → edge20–80ms20–80msCDN coverage, HTTP/3
Edge → origin50–200ms0 (cache hit)cacheability
Origin render + data200–700ms0 for shellsplitting personal data out
Fragment fetch—30–150ms, in parallel with shell streamingfragment endpoint speed
First byte to clientsum of the aboveclient→edge + edge compute (~1–10ms)stream the shell immediately

First byte timing, origin-rendered vs edge-assembled Two timelines comparing time to first byte for a page rendered at the origin and a cached shell streamed from the edge while a personal fragment is fetched. First byte timing, origin-rendered vs edge-assembled Origin to edge to origin render + DB Edge shell to edge Fragment personal JSON 0ms 100ms 200ms 300ms 400ms 500ms 600ms 700ms 800ms 900ms 200ms TTFB The cached shell starts streaming at about 70ms; the personal fragment arrives in parallel and is inserted when ready.

Advanced Diagnostics and Edge Cases

Cache poisoning through personal data. The biggest risk of dynamic caching is serving one user's content to another. Any page cached at the edge must be rendered without reading per-user state; enforce it in code (render functions that do not receive the request's cookies) and test it (request as two users, compare responses).

Edge compute cold starts. Isolate-based platforms (Workers, Vercel/Netlify edge) start in milliseconds; container- or VM-based edge functions can have cold starts of hundreds of milliseconds. Measure p95 edge processing time, not just averages.

Data locality. Edge functions are close to users, but databases usually are not. An edge function that queries a single-region database for every request can be slower than origin rendering. Keep data access at the edge to cached reads or globally replicated stores.

Purging. Cached shells must be invalidated when content changes. Tag-based purging keeps this precise — see purging CDN cache by tag on deploy.

Streaming and buffering. Assembling at the edge only helps TTFB if the edge streams the shell immediately rather than buffering the full response. Some platforms buffer by default when you transform HTML; check with a timing waterfall.

Validation and Budgeting

Track TTFB p75 per template split by anonymous and logged-in traffic, the edge cache hit ratio for HTML, and a correctness check. Set budgets such as "logged-in product page TTFB p75 under 200ms" and "HTML hit ratio over 85%".

javascript
// ci/cache-isolation.test.mjs — make sure two users never receive each other's content.
const a = await fetch(url, { headers: { cookie: 'session=userA' } }).then((r) => r.text());
const b = await fetch(url, { headers: { cookie: 'session=userB' } }).then((r) => r.text());
if (a.includes('userA@example.com') && b.includes('userA@example.com')) {
  throw new Error('Personal data leaked across users via the edge cache');
}
// trade-off: string checks only catch the markers you look for. Render test
// users with unique, searchable markers in every personal region.

Worked Example: Moving Logged-In Traffic onto the Cache

A retailer's logged-in users — 40% of traffic and most of the revenue — bypassed the CDN entirely because the product page template read the session cookie to render a greeting and the cart count. TTFB p75 for them was 820ms versus 110ms for anonymous users. The team removed all session reads from the template, served the greeting and cart count from a small /api/me endpoint fetched by an edge function and streamed into a placeholder, and keyed the cache on a derived currency value instead of Vary: Cookie. Logged-in TTFB p75 fell to 140ms, LCP p75 for logged-in users improved by 560ms, and origin CPU dropped by more than half. A nightly isolation test confirmed no personal data was cached.

Choosing Fragment Granularity

How finely to split a page is a trade-off between cache efficiency and request overhead. One personal fragment per page — a single /api/me response carrying the greeting, cart count and any small personal flags — is usually the sweet spot: one request, fetched in parallel with the shell, small enough to be fast. Splitting every personal widget into its own fragment multiplies requests and complicates assembly; merging personal data back into the page defeats caching.

Recommendation rails and other expensive personal regions are the exception. They often depend on slower services, so give them their own fragment, load them after the main content, and reserve their space. Treat them as progressive enhancement: the page is complete without them, and a slow recommendation service never delays the first paint.

Variant regions follow a different rule: keep the number of variants small. Currency by country might produce ten variants; experiment arms two or three. Multiplying several variant dimensions together (currency × experiment × device class × language) quickly produces hundreds of cache keys per URL, most of which never warm. Collapse dimensions where you can — device class rarely needs to be in the key if the HTML is responsive — and move rare variants to client-side adjustment.

Measuring Edge Function Cost

Edge compute is not free in latency or money. Add your own Server-Timing entry for edge processing time and track its p95; isolate runtimes typically add 1–10ms, but code that awaits subrequests sequentially, parses large HTML with string operations, or calls a distant database can add hundreds. Use streaming HTML rewriters (such as Cloudflare's HTMLRewriter) instead of buffering and regex-replacing the full document; they transform chunks as they pass through and keep TTFB low. Watch request-count billing, too: an edge function on every request, including static assets, multiplies cost for no benefit. Route only HTML and fragment requests through compute, and let static assets hit the cache directly.

Observability for Dynamic Edge Caching

Once pages are assembled at the edge, a slow page can be slow for several new reasons: a cache miss on the shell, a slow fragment endpoint, an edge function waiting on a subrequest, or a purge storm after a deploy. Make each visible. Emit Server-Timing entries for shell cache status, edge processing time and fragment fetch time, carry them into RUM, and chart TTFB p75 segmented by shell hit or miss. Log fragment endpoint latency at p95 separately — fragments are on every request, so their tail latency matters more than their average. Finally, keep a per-template dashboard of HTML hit ratio over time with deploys annotated; a sharp dip after each deploy suggests purges are broader than necessary, and a slow decline suggests a new cache-key dimension is fragmenting the cache.

A Rollout Plan

Edge caching of dynamic pages is safe when introduced in stages. Start with one high-traffic template and anonymous traffic only, confirming hit ratio and correctness. Then remove cookie reads from that template's render path and enable caching for logged-in traffic behind a feature flag for a small percentage of users, with the isolation test running continuously. Expand to other templates once the pattern is proven, and only then consider edge-side personalisation of the first paint. Keep a kill switch that forces cache bypass per template; it is the fastest way to recover if a personal region slips into a cached shell.

FAQ

Is edge rendering the same as edge caching?

No. Edge rendering runs your full rendering code at the edge for every request; it helps when the render is cheap and data is available at the edge, but it does not avoid per-request work. Edge caching serves stored responses and is far cheaper. The patterns in this topic mostly use the edge to select and assemble cached content, with rendering at the edge reserved for small personalised pieces.

Does client-side fetching of personal data hurt LCP?

Usually not, because personal regions (account menus, cart counts, recommendation rails) are rarely the LCP element. The cached shell, including the LCP element, paints quickly; personal regions fill in afterwards. Reserve space for them to avoid layout shift.

How do A/B tests fit in?

Assign the experiment arm at the edge (from a cookie, or by bucketing a new visitor), add the arm to the cache key, and serve the cached variant. That keeps experiments flicker-free and cacheable, instead of rewriting the page client-side after load.

What about pages that are truly unique per user, like a dashboard?

Cache the application shell (layout, navigation, static UI) and fetch the user's data separately, either streamed from the edge or loaded on the client. The shell gives a fast first paint; the data determines when the page is useful, so optimise the data endpoint's latency as well.

Which platform should I choose?

Choose the edge platform that matches your CDN, since running compute next to the cache is the point. Prefer isolate-based runtimes for low cold-start latency, and check the platform's cache API, streaming support and HTML rewriting capabilities against the patterns you need.

Can static site generation replace edge caching?

For pages that are the same for everyone, yes — prebuilt HTML on a CDN is the simplest fast option. Static generation does not address personalisation, variants or content that changes faster than you can rebuild. Edge caching with tag-based purging covers those cases while keeping static-like TTFB.

How do I keep cached shells consistent with fragments after a deploy?

Version both. Include a build identifier in the shell and send it with fragment requests; if a fragment endpoint from a newer release returns data the old shell cannot render, the mismatch is detectable and the edge can fall back to fetching a fresh shell. Purging shells by tag on deploy keeps the window short.

Guides in This Topic