Time to First Byte (TTFB) Optimization: Getting the Response Started
This topic is part of Network & Server Response Optimization. Time to First Byte measures the time from the start of a navigation until the first byte of the HTML response arrives. It includes everything that happens before the browser has anything to work with: redirects, DNS lookup, connection setup, TLS negotiation, the request's trip to the server, all the work the server does to build the response, and the first bytes' trip back. Nothing visible can happen before it, so TTFB is a floor under First Contentful Paint and Largest Contentful Paint.
Google treats a TTFB of 800ms or less at the 75th percentile as good. That threshold is generous: on a mobile connection, if TTFB is 800ms, only 1.7 seconds remain for CSS, fonts, the LCP image and rendering before LCP crosses 2.5 seconds. Fast sites usually achieve 200–500ms at p75, mostly by serving HTML from a CDN cache close to users.
TTFB problems are rarely one thing. A typical slow TTFB is a stack: a tracking redirect (250ms), a connection to a distant origin without a CDN (300ms), and server time of 600ms because three backend calls run sequentially. Fixing it means measuring each component and attacking the biggest — which this topic organises into a repeatable workflow.
The Metric Degradation This Topic Addresses
High TTFB delays everything downstream: FCP, LCP, and indirectly INP (since hydration starts later). It shows up in the LCP breakdown as a large first sub-part. In field data, TTFB is unusually variable by geography: users far from a single-region origin see much higher values. It is also sensitive to cache state: a page served from CDN cache might have a 100ms TTFB, while a cache miss on the same page takes 1.5 seconds.
Two field patterns are common. A uniformly high TTFB across regions points at server time (slow backend, no caching). A TTFB that grows with distance from the origin points at missing edge caching or connection costs. A bimodal distribution (many fast, many slow) usually reflects cache hit versus miss.
Prerequisites
- Field TTFB at p75 by page type, country and device, from RUM or CrUX.
- Navigation Timing sub-parts from RUM (redirect, DNS, connect, TLS, request-to-response).
- Access to server logs or tracing for backend time.
- CDN analytics showing cache hit ratio for HTML.
1. Environment Setup: Collect TTFB Sub-Parts
// RUM: break TTFB into components from Navigation Timing.
const [nav] = performance.getEntriesByType('navigation');
const parts = {
redirect: nav.redirectEnd - nav.redirectStart,
dns: nav.domainLookupEnd - nav.domainLookupStart,
connect: nav.connectEnd - nav.connectStart, // includes TLS
tls: nav.secureConnectionStart ? nav.connectEnd - nav.secureConnectionStart : 0,
request: nav.responseStart - nav.requestStart, // server time + latency
ttfb: nav.responseStart,
protocol: nav.nextHopProtocol,
serverTiming: nav.serverTiming?.map((s) => `${s.name}:${Math.round(s.duration)}`).join(','),
};
navigator.sendBeacon('/rum', JSON.stringify(parts));
// trade-off: cross-origin redirects hide redirect timing (it reports 0); the gap
// between startTime and fetchStart still reveals that time was spent.
2. Capture a Baseline
Compute p75 for each sub-part by page type and region. The largest sub-part at p75 is where to start. Record CDN cache hit ratio for HTML and the share of navigations with redirects.
3. Isolate the Dominant Component
If redirects are large, find the redirecting URLs (logs, campaign links, internal links). If connect/TLS is large, check CDN coverage, TLS version, HTTP/3 and certificate chain size. If request (server time) dominates, add Server-Timing to see what the server does and look at cache hit ratios.
4. Apply the Matching Fix
Remove redirects at their source. Put a CDN in front of HTML and cache it where possible. Use TLS 1.3 and enable HTTP/3. Parallelise and cache backend calls. Stream or early-flush responses that cannot be cached. Address cold starts on serverless platforms.
// Parallelise independent backend calls instead of awaiting them one by one.
const [product, reviews, stock] = await Promise.all([
getProduct(id), getReviews(id), getStock(id),
]);
// trade-off: Promise.all fails fast if one call rejects; use Promise.allSettled
// when partial pages are acceptable, rendering fallbacks for failed sections.
Deconstructing Server Time
Server time is usually the largest and most variable component. It includes request routing (load balancers, API gateways), middleware (authentication, sessions, A/B assignment, logging), data fetching (databases, internal APIs, third-party APIs), rendering (templating or SSR), and response compression. Each is a candidate for optimization, but the pattern that matters most is sequential dependencies: three 150ms calls in sequence take 450ms; in parallel they take 150ms.
Server-Timing makes this visible. The server sends a header listing named durations (db;dur=45, api-reviews;dur=210, render;dur=60), which appear in DevTools' Timing tab and are available to RUM via PerformanceServerTiming. With RUM collection, you see the p75 of each phase in the field — often revealing that one backend dependency dominates.
Caching attacks server time at different levels: full-page caching at the CDN (eliminates server time for cache hits), application-level caching of rendered fragments or data, and database query caching. Full-page caching gives the largest gains; even short TTLs (60 seconds) with stale-while-revalidate can serve most requests from the edge for popular pages.
Advanced Diagnostics and Edge Cases
Personalised pages. Logged-in or personalised pages often bypass CDN caches. Separate the shared shell from personal data — cache the shell, fetch personal parts client-side or assemble at the edge.
Cookie-dependent caching. CDNs often bypass cache when requests carry cookies. Configure the cache key to ignore cookies that do not affect the response.
Geographic latency. For global audiences, a single-region origin adds 100–300ms of network latency for distant users on every uncached request. Edge caching or multi-region deployment addresses it.
Cold caches after deploys. Purging all HTML on deploy causes a burst of slow TTFBs. Use versioned assets so HTML can be revalidated gradually instead of purged.
Bot traffic. Crawlers hitting uncached low-traffic pages can load origins and slow everyone. Watch origin load and cache hit ratio together.
Validation and Budgeting
Validate in the field: TTFB p75 per template and region, before and after each change. In the lab, use WebPageTest from several locations to see connection and server components. Set budgets for TTFB p75 per key template and alert on regressions, plus a budget for HTML cache hit ratio.
| Budget | Target |
|---|---|
| TTFB p75 (mobile, all regions) | < 600 ms |
| HTML cache hit ratio (cacheable templates) | > 85% |
| Navigations with redirects (entry pages) | < 5% |
| Server time p75 (uncached) | < 300 ms |
Worked Example: A Regional News Site
A regional news site hosted on a single origin in one region, with no CDN caching for HTML (the CDN only served assets). Mobile TTFB p75 was 1.4 seconds; LCP p75 was 3.6 seconds. Navigation Timing showed: redirects 180ms (social links pointed at http:// URLs and a tracking domain), connect 260ms, server time 820ms (article rendering made sequential calls to a CMS API, a related-articles service and a comments count). The team updated social link templates to canonical HTTPS URLs, enabled CDN caching for article HTML with s-maxage=120, stale-while-revalidate=3600, and parallelised the three backend calls for cache misses. TTFB p75 fell to 380ms, LCP p75 to 2.3 seconds. Cache hit ratio for article HTML settled at 93%.
Edge Rendering and Regional Origins
For pages that cannot be cached, moving rendering closer to users reduces the network part of TTFB. Edge functions run near users, but if they must call a database in one region, each query pays the long round trip — sometimes making TTFB worse than rendering at the origin. The rule: compute should be close to data. Edge rendering helps when data is cached at the edge or replicated globally; otherwise, keep rendering near the data and cache what you can at the edge. Multi-region deployments with read replicas are the heavy-duty option for global, dynamic applications.
Measuring TTFB Correctly
Several measurement traps make TTFB look better or worse than users experience. Lab tools often run on fast connections from data centres near your origin. Synthetic tests may hit warm caches. Some RUM libraries measure TTFB from requestStart rather than navigation start, excluding redirects and connection setup. The web-vitals library reports TTFB from the start of navigation, which matches what users wait for. For prerendered or back/forward-cached pages, TTFB is effectively zero and should be segmented separately, otherwise improvements in speculation can mask regressions in server time.
Connection Setup: The Overlooked Share
For first-time visitors on mobile, connection setup can rival server time. A DNS lookup through a slow resolver, a TCP handshake, and a TLS 1.2 handshake with a long certificate chain can take 400–600ms before the request is even sent. Several changes reduce it. TLS 1.3 removes a round trip from the handshake, and session resumption makes repeat connections cheaper still. HTTP/3 combines the transport and TLS handshakes over QUIC, saving another round trip and coping better with packet loss. A CDN terminates connections close to the user, so each round trip is short even if the origin is far away. Shorter certificate chains and OCSP stapling reduce the bytes and checks in the handshake.
Check what your users actually negotiate: nextHopProtocol in Navigation Timing reports h2 or h3, and the connect and TLS sub-parts show how long setup took. If a large share of navigations still use HTTP/2 when HTTP/3 is enabled, the alt-svc advertisement may be missing or blocked, or users may be returning before the browser has learned that HTTP/3 is available.
Caching Strategies for Different Page Types
Not every page can be cached the same way, but most can be cached somehow. Static content pages (articles, docs, marketing) can be cached at the CDN for minutes to hours, with stale-while-revalidate so updates propagate without slow requests. Listing pages (categories, search results for common queries) can be cached briefly, keyed by the query parameters that change the result. Product pages with frequently changing prices or stock can cache the page shell and fetch volatile data client-side or at the edge. Logged-in pages can often cache a shared shell and personalise a small part. Only truly unique, per-request responses — checkout steps, account settings — need full server rendering on every request, and those are usually a small share of traffic.
Measuring cache hit ratio per template makes this concrete. A template with 40% hit ratio usually has a cache key problem (cookies, query parameters, Vary headers) rather than truly uncacheable content.
TTFB and Back/Forward Cache
Navigations restored from the back/forward cache (bfcache) have no network request at all: the page is restored from memory instantly. Making pages eligible for bfcache (avoiding unload handlers, Cache-Control: no-store on HTML where not needed, and open connections that block it) turns a large share of back and forward navigations into zero-TTFB events. In field data, segment bfcache restores separately so they do not hide regressions in real server response times.
A Rollout Plan
- Collect Navigation Timing sub-parts and Server-Timing in RUM.
- Remove redirects from high-traffic entry URLs.
- Enable CDN caching for HTML on cacheable templates, with stale-while-revalidate.
- Turn on TLS 1.3 and HTTP/3 at the CDN.
- Add Server-Timing and fix the slowest backend dependencies; parallelise calls and add timeouts with fallbacks.
- Stream or early-flush responses that must be rendered per request, so the head arrives first.
- Set budgets per template and region, and alert on regressions.
Common Pitfalls
- Optimising server code while ignoring redirects and connection time. Often the bigger share for mobile users.
- Never caching HTML. Leaves the largest TTFB win unused.
- Sequential backend calls. Turns several fast calls into one slow response.
- Edge rendering far from data. Adds round trips instead of removing them.
- Measuring TTFB only in the lab. Misses geography, cache misses and real networks.
FAQ
What is a good TTFB?
800ms or less at p75 is considered good. Aim for under 600ms on mobile for a comfortable LCP budget; fast sites achieve 200–500ms.
Does TTFB include redirects?
In the web-vitals definition, yes — TTFB is measured from the start of navigation, so redirect time is included.
How do I cache HTML safely?
Cache only responses that are the same for all users (or vary by a small set of keys), set s-maxage with stale-while-revalidate, and make sure cookies or personalisation do not leak into cached pages.
What is Server-Timing?
A response header in which the server reports named durations for its work. Browsers show it in DevTools and expose it to RUM scripts.
Does a CDN help TTFB for uncached pages?
Somewhat: CDNs terminate connections near users and often keep warm connections to the origin, which reduces connection setup time. The large gains come from caching.
Why is TTFB slower on mobile?
Higher network latency multiplies the cost of every round trip in redirects and connection setup. Server time is usually the same.
Can prerendering hide slow TTFB?
For prerendered navigations, users do not wait for TTFB. But entry pages and non-predicted navigations still pay it, so fix the underlying server time too.
Does HTML size affect TTFB?
Not directly — TTFB is the arrival of the first byte. But large HTML takes longer to transfer after that, delaying discovery of resources near the end of the document. Compression and streaming help with both.
Should I measure TTFB for API requests too?
Yes, for client-rendered pages that fetch data before rendering. Their response times behave like TTFB for the content users wait for, and Server-Timing works on API responses as well.
Guides in This Topic
- Diagnosing slow TTFB with Server-Timing — see where server time goes.
- Fixing TTFB from serverless cold starts — initialisation costs on serverless platforms.
- Eliminating redirect chains — remove round trips before the first byte.
Related
- Edge compute & dynamic caching — caching pages that look uncacheable.
- Streaming SSR & early flush — sending bytes before the server finishes.
- Cache-Control for HTML documents — header settings for cached HTML.