How to Diagnose Slow TTFB with Server-Timing

This guide is part of Time to First Byte Optimization, within Network & Server Response Optimization. When TTFB is slow and Navigation Timing shows that most of it is the time between sending the request and receiving the first byte, the problem is server time. But "the server is slow" is not actionable. Is it the database, a downstream API, authentication, template rendering, a cache miss, or a queue in front of the application?

The Server-Timing HTTP response header lets the server report named durations for its work, directly to the browser. Chrome DevTools displays them in the Network panel's Timing tab, and the PerformanceServerTiming API exposes them to JavaScript, so RUM scripts can collect them from real users. With a few lines of instrumentation, you can see the p75 of each server phase in the field, on the same dashboard as TTFB and LCP.

Server-Timing breakdown for a product page (p75, field) Bar chart of server-side phases for a product page at the 75th percentile, collected via Server-Timing in RUM. Server-Timing breakdown for a product page (p75, field) auth/session 35ms db-product 48ms api-pricing 72ms api-recommendations 410ms render 85ms

Rapid Diagnosis

  • Check Navigation Timing: responseStart - requestStart large compared with connection time means server time dominates.
  • Look for an existing Server-Timing header in DevTools (Network → select document → Timing).
  • Check CDN cache status headers on the document; a miss means the origin did the work.
  • Compare TTFB by template to find which routes are slow.

Root Cause Analysis

1. Slow dependencies. One downstream API or query dominates.

2. Sequential calls. Independent work awaited one after another.

3. Cache misses. Application or CDN caches missing more than expected.

4. Queueing. Requests waiting for a worker or database connection under load.

Step-by-Step Resolution

1. Instrument the server

javascript
// Generic helper: time async work and collect Server-Timing entries.
export function serverTimer() {
  const entries = [];
  return {
    async time(name, fn, desc) {
      const t0 = performance.now();
      try { return await fn(); }
      finally { entries.push(`${name};dur=${(performance.now() - t0).toFixed(1)}${desc ? `;desc="${desc}"` : ''}`); }
    },
    header: () => entries.join(', '),
  };
}
// Route: const t = serverTimer(); const p = await t.time('db-product', () => db.product(id));
// res.setHeader('Server-Timing', t.header());
// trade-off: headers must be set before the body is sent; with streaming, report
// only the phases completed before the first flush, or use trailers where supported.

2. Add cache status and total

Report cache hits as entries (cache;desc="hit") and a total (total;dur=...), so you can segment field data by cache outcome.

3. Read it in DevTools

In the Network panel, select the document request and open the Timing tab: Server-Timing entries appear under "Server Timing" with their durations.

4. Collect it in RUM

javascript
const [nav] = performance.getEntriesByType('navigation');
const st = Object.fromEntries((nav.serverTiming || []).map((e) => [e.name, Math.round(e.duration)]));
navigator.sendBeacon('/rum', JSON.stringify({ ttfb: Math.round(nav.responseStart), ...st, path: location.pathname }));
// trade-off: cross-origin documents only expose Server-Timing with a
// Timing-Allow-Origin header; same-origin navigations need nothing extra.

Server-Timing from server to dashboard Sequence showing the server adding Server-Timing, the browser exposing it, and a RUM script reporting it. Server-Timing from server to dashboard Server Browser RUM script Dashboard HTML + Server-Timing timing entries beacon with phases + TTFB p75 per phase

Verification

After fixing the slowest phase (for example, caching or parallelising a dependency), the corresponding Server-Timing entry should fall in the field at p75, along with TTFB for the affected template. Compare distributions, not just averages: tail latency of dependencies often matters more.

Worked Example: A Product Page With a Slow Recommendations API

An online retailer saw product page TTFB p75 of 1.1 seconds for cache misses. Server-Timing entries collected in RUM showed api-recommendations at 410ms p75 (and over 1.2 seconds at p95), called sequentially after the product and pricing lookups. The team moved recommendations out of the critical path: the page rendered without them and fetched them client-side after load (they were below the fold). Pricing and product calls were parallelised. Server time p75 fell from 650ms to 160ms, TTFB p75 for misses to 520ms, and the recommendations still appeared before most users scrolled to them.

From Server-Timing data to a fix Steps from collecting Server-Timing phases in the field to choosing and verifying a fix for the slowest phase. From Server-Timing data to a fix Instrument phases + cache Collect RUM beacons Rank p75 and p95 per phase Fix cache, parallelise, defer Verify phase + TTFB drop

Security and Privacy Considerations

Server-Timing is visible to anyone who can load the page, including competitors and attackers. Avoid descriptions that reveal internal hostnames, query text, user identifiers or infrastructure details. Durations themselves can, in rare cases, leak information (for example, timing differences between existing and non-existing accounts). A common approach is to emit detailed timings only for sampled requests or for authenticated internal users, and coarse phases (cache, app, total) for everyone else. For cross-origin resources, Timing-Allow-Origin controls which origins can read the entries.

Interpreting the Numbers

Server-Timing data is most useful when you compare distributions across phases and cache outcomes. A phase whose median is small but whose p95 is large is a tail-latency problem, usually caused by a dependency that sometimes times out or retries; adding a timeout with a fallback can remove the tail entirely. A phase that is consistently large is a throughput or design problem: add caching, precompute results, or move the work off the critical path. When the sum of reported phases is much smaller than responseStart - requestStart, the missing time is outside your instrumentation: network latency, load balancer queueing, or framework overhead before your handler runs. Add a phase that starts as early as possible in the request lifecycle to capture it.

Segment by cache status first. If cache hits are fast and misses slow, improving hit ratio is often cheaper than speeding up the miss path. If even hits are slow, the issue is in the edge or network path rather than the application.

Building a Dashboard

A useful dashboard shows, per key template: TTFB p75, the p75 of each Server-Timing phase, the cache hit ratio, and the share of cold starts if you run on serverless. Plot them over time with deploy markers. When TTFB regresses, the phase that moved points at the cause within minutes rather than after a long investigation.

Common Mistakes

  • Instrumenting only the total. You need phases to find the culprit.
  • Exposing sensitive descriptions. Keep names generic.
  • Measuring only in the lab. Field p75 and p95 reveal tail latency that a few test runs miss.
  • Ignoring queueing time. Time before your handler runs is invisible unless you measure from request arrival.

Edge Cases

CDN-added entries. Some CDNs add their own Server-Timing entries (cache status, edge time); keep names distinct to avoid confusion.

Streaming responses. Headers are sent with the first flush; later phases can be reported via a separate beacon or trailers where supported.

Serverless platforms. Report initialisation time as a separate entry to see cold start impact.

Multiple services. Propagate timings from downstream services through a tracing system; Server-Timing summarises the top-level view.

FAQ

Is Server-Timing supported in all browsers?

All major browsers support reading Server-Timing via the Performance APIs. DevTools display support varies, but the data is available to scripts.

Does Server-Timing add overhead?

Negligible — a short header and a few timer calls. Avoid very many entries per response.

Can I use Server-Timing for subresources?

Yes. API and asset responses can include it; cross-origin resources need Timing-Allow-Origin for scripts to read the entries.

How is this different from distributed tracing?

Tracing captures detailed spans across services for engineers. Server-Timing delivers a summary to the browser, where it can be correlated with user-facing metrics like LCP.

Should I include a desc attribute?

Optional. It is useful for categorical data like cache hit or miss, but keep it free of sensitive information.

What does a cache entry tell me?

Whether the response came from cache, so you can split TTFB into hit and miss distributions and see whether improving hit ratio or server time matters more.

Can CDNs strip Server-Timing headers?

Some configurations remove unknown headers or replace them. Check the header at the browser, not just at the origin.

/html>