How to Expose Backend Timing with Server-Timing Headers
This guide adds the server's side of the story to RUM Beacons & Field Data Collection, within Core Web Vitals & Measurement. Time to First Byte is the first phase of every LCP, and in RUM it is a single opaque number: the browser knows when it sent the request and when the first byte arrived, nothing about what happened in between. Was it a CDN cache miss? A slow database query? Server-side rendering of a heavy component tree? A cold serverless start?
The Server-Timing response header answers those questions in the field. The server lists named metrics with durations — cdn-cache;desc=MISS, db;dur=184, render;dur=96 — and the browser exposes them to JavaScript through PerformanceResourceTiming.serverTiming and on the navigation entry. Your RUM beacon reads them alongside LCP, and a slow TTFB becomes "slow because the product query took 400ms on cache misses for this template".
Rapid Diagnosis
- Check for the header today. In DevTools Network, select the document request and open the Timing tab; any Server-Timing entries appear at the bottom. Many CDNs and frameworks add some already.
- Correlate TTFB with LCP. In RUM, if TTFB p75 is a large fraction of LCP p75 (say over 30%), the backend deserves attribution.
- Check variance. A bimodal TTFB distribution (fast cluster plus slow cluster) usually means cache hits vs misses — exactly what a cache-status entry reveals.
- Check cross-origin access. Server-Timing on cross-origin resources is only exposed if the response includes
Timing-Allow-Origin.
Root Cause Analysis
1. TTFB is a black box. Without server attribution, field TTFB can only be compared over time, not explained.
2. Server logs are disconnected from user experience. Backend dashboards show query times per request; RUM shows TTFB per page view; joining them requires a shared identifier that is rarely in place.
3. Cache behaviour is invisible. Edge hit ratio dashboards report aggregates; they do not tell you which users and pages experienced misses and how slow those were.
4. Cold starts and queueing. Serverless cold starts and request queueing add latency that never appears in application timers measured inside the handler.
Step-by-Step Resolution
1. Emit Server-Timing from the origin
Time the phases you control and add them as entries. Keep names short and stable; dur is in milliseconds, desc is free text.
// Express/Node middleware — append phase timings to the response.
app.use((req, res, next) => {
const marks = [];
res.locals.time = async (name, fn) => {
const t0 = performance.now();
try { return await fn(); } finally { marks.push(`${name};dur=${(performance.now() - t0).toFixed(1)}`); }
};
const writeHead = res.writeHead;
res.writeHead = function (...args) {
if (marks.length) res.setHeader('Server-Timing', marks.join(', '));
return writeHead.apply(this, args);
};
next();
});
// in a handler: const product = await res.locals.time('db', () => getProduct(id));
// trade-off: Server-Timing is visible to anyone who opens DevTools. Do not
// include internal hostnames, query text or anything sensitive in desc values;
// strip or gate detailed entries for unauthenticated traffic if necessary.
Expected outcome: document responses carry db, render and other named durations.
2. Add edge timing and cache status
Most CDNs can append their own entries (cache status, edge processing time) via configuration or edge code, so you see where in the chain time was spent.
// Cloudflare Worker: forward origin timing and add cache status.
export default {
async fetch(req, env, ctx) {
const t0 = Date.now();
const res = await fetch(req, { cf: { cacheEverything: true } });
const out = new Response(res.body, res);
const status = res.headers.get('cf-cache-status') || 'NONE';
out.headers.append('Server-Timing', `edge;dur=${Date.now() - t0}, cdn-cache;desc=${status}`);
return out;
},
};
// trade-off: Date.now() inside a Worker is coarsened and does not advance
// during CPU work, so edge durations reflect I/O wait. That is the part you
// usually care about, but it is not a full CPU profile.
Expected outcome: every document response says whether it was a cache hit and how long the edge spent.
3. Collect entries in RUM
function serverTimings() {
const nav = performance.getEntriesByType('navigation')[0];
return Object.fromEntries((nav?.serverTiming || []).map((t) => [t.name, t.duration || t.description]));
}
onLCP((m) => send({ lcp: m.value, ttfb: performance.getEntriesByType('navigation')[0].responseStart,
server: serverTimings() }));
// trade-off: only entries on the navigation response are read here. For API
// calls that gate client rendering, also read serverTiming from their resource
// entries (requires Timing-Allow-Origin if cross-origin).
Expected outcome: each LCP beacon carries the server phases that produced its TTFB.
4. Analyse TTFB by server phase and cache status
Group field TTFB by cdn-cache status and template, and plot db and render durations at p75 for misses. That tells you whether to raise cache hit ratio (longer TTLs, better keys, stale-while-revalidate) or to make misses cheaper (query optimisation, rendering less on the server).
Expected outcome: a ranked list of backend causes of slow TTFB, weighted by how many real page views they affected.
Verification
Load a page with DevTools open and confirm the Timing tab shows your entries for the document. In RUM, confirm the share of beacons carrying server fields matches the share of traffic served by instrumented routes. After a backend fix — say, caching a query — the db p75 for misses and the MISS-segment TTFB p75 should both fall, and overall LCP p75 should move by roughly the TTFB change multiplied by the miss rate.
Joining Server-Timing with Traces
For deeper analysis, send a trace or request id as a Server-Timing entry (trace;desc=4bf92f3577b34da6) and include it in the RUM beacon. Your observability platform can then link a specific slow page view to the distributed trace that produced it — database spans, downstream service calls, queue time — without guesswork. Sample this: emitting trace ids on every response is cheap, but storing the link for every beacon may not be; keep it for slow page views (TTFB above a threshold) and a small random sample of the rest. Be aware that exposing a trace id is generally harmless but should not grant access to anything on its own.
Naming Conventions That Survive
Server-Timing entries become dashboard dimensions, so treat their names as an API. Use lowercase, short, stable names (db, render, auth, cdn-cache, edge, app), never per-request values in names, and put variable data in desc only when it is low-cardinality (cache status, region code). Document the list next to the middleware that emits it. When a new phase is added, add a new name rather than redefining an existing one — otherwise historical comparisons silently change meaning. Finally, keep the total header small: a dozen entries is plenty, and very long headers can be truncated by intermediaries.
FAQ
Does Server-Timing work with HTTP caching?
The header is cached with the response, so a cache hit replays the origin's original timings — misleading if you read them as "this request's" database time. Add a cache-status entry at the edge on every response, and interpret origin timings only when the status is a miss.
Can Server-Timing be sent after the body starts streaming?
As a trailer, in principle, but browser support for Server-Timing trailers has been limited. With streaming SSR, emit what you know before the first byte (cache status, routing, data fetch for the shell) in the header and record later phases server-side only.
Does it add measurable overhead?
A few dozen bytes per response and negligible CPU. The main cost is discipline: keeping names stable so dashboards do not break, and keeping sensitive details out of desc.
Can third-party APIs send Server-Timing that my RUM can read?
Only if they include the header and a Timing-Allow-Origin header that allows your origin. Without Timing-Allow-Origin, cross-origin serverTiming arrays are empty in the browser. For your own APIs on separate hostnames, add both headers; for vendors, ask — some already provide useful timings.
Related
- Diagnosing slow TTFB with Server-Timing — the server-side optimisation workflow.
- Diagnosing CDN cache misses from response headers — raising the hit ratio the cache-status entry measures.
- Segmenting RUM by device and connection — combining server and client dimensions.