How to Segment RUM Data by Device and Connection

This guide extends RUM Beacons & Field Data Collection within Core Web Vitals & Measurement. A p75 that fails tells you there is a problem; it does not tell you who has it. The slow quarter of page views is almost never a uniform slice of your audience — it is concentrated in particular devices, networks, countries and page templates. Segmentation turns "INP is 280ms" into "INP is 450ms on 4GB Android devices and 140ms everywhere else", which tells you to reduce main-thread work rather than bytes.

The difficulty is that the most useful dimensions are partly browser-specific, partly coarse by design (for privacy), and easy to collect in ways that inflate cardinality. This guide covers which signals to capture, how to bucket them into a small number of meaningful classes, and how to read the result.

Dimensions worth capturing with every beacon Layers of device, network, context and page dimensions that make RUM data segmentable. Dimensions worth capturing with every beacon Device deviceMemory, hardwareConcurrency, form factor (Client Hints or UA) Network effectiveType, rtt, downlink (Chromium), plus measured TTFB Context country, navigation type, first vs repeat visit, logged-in Page route template, release id, A/B variant

Rapid Diagnosis

  • Check what your beacon contains today. Many beacons carry only URL and metric values. Without device and network context, segmentation is impossible after the fact.
  • Look at your browser mix. navigator.deviceMemory and navigator.connection are Chromium-only. For Safari-heavy audiences you need derived signals (measured TTFB, resource timings).
  • Check cardinality. Raw rtt and downlink values, full user-agent strings and precise geo produce thousands of distinct values. Bucket them before storage.
  • Compare segment p75s. Even a crude split (mobile vs desktop × "low" vs "high" memory) usually reveals where the slow tail lives.

Root Cause Analysis

1. Missing context. Beacons without device or network fields cannot be segmented, so analysis falls back to averages that hide the problem.

2. Unbucketed values. Storing raw RTT in milliseconds or full UA strings creates high-cardinality columns that are expensive to query and impossible to read in a dashboard.

3. Browser-specific signals treated as universal. Segmenting on deviceMemory silently excludes Safari and Firefox users; conclusions drawn from it describe Chromium only.

4. Confounding. Slow devices are more common in some countries and on some templates. A device effect can actually be a template effect if not controlled for.

Mobile INP p75 by device memory class Bar chart of mobile p75 INP split by reported device memory, showing low-memory devices far above the 200 millisecond threshold. Mobile INP p75 by device memory class ≤ 2 GB 512ms 4 GB 318ms 8 GB 154ms Unknown (Safari, Firefox) 171ms 200ms INP

Step-by-Step Resolution

1. Collect device and network context once per page view

javascript
function context() {
  const c = navigator.connection || {};
  const nav = performance.getEntriesByType('navigation')[0];
  return {
    mem: navigator.deviceMemory ?? null,                       // 0.25..8, Chromium only
    cores: navigator.hardwareConcurrency ?? null,
    ect: c.effectiveType ?? null,                              // 'slow-2g'..'4g'
    rttBucket: c.rtt != null ? Math.min(Math.ceil(c.rtt / 100) * 100, 1000) : null,
    ttfbBucket: nav ? Math.min(Math.ceil(nav.responseStart / 200) * 200, 4000) : null,
    mobile: navigator.userAgentData?.mobile ?? /Mobi/.test(navigator.userAgent),
  };
}
// trade-off: deviceMemory and hardwareConcurrency are rounded and capped by the
// browser for privacy (memory tops out at 8). They separate low-end from
// mid-range well, but cannot distinguish a mid-range phone from a flagship.

Expected outcome: every beacon carries a handful of low-cardinality fields that cover device and network.

2. Derive a device class that works across browsers

Combine signals into a small set of classes, with a fallback for browsers that expose less.

javascript
function deviceClass(ctx) {
  if (ctx.mem != null) return ctx.mem <= 2 ? 'low' : ctx.mem <= 4 ? 'mid' : 'high';
  if (ctx.cores != null) return ctx.cores <= 4 ? 'low' : ctx.cores <= 6 ? 'mid' : 'high';
  return 'unknown';
}
// trade-off: core counts on iPhones are modest (6) while their per-core speed is
// high, so cores alone misclassify iOS devices as "mid". Keep "unknown" honest
// rather than forcing every device into a class.

Expected outcome: a three- or four-value deviceClass column that most dashboards can split by.

3. Use measured network timings for all browsers

effectiveType and rtt are unavailable outside Chromium, but every browser exposes Navigation Timing. Bucketed TTFB and the connection phases (connectEnd - connectStart) are cross-browser proxies for network quality.

Expected outcome: a network dimension that covers Safari users too.

4. Segment, then control for confounders

Compute p75 per segment, and before drawing conclusions, check that the comparison holds within a template and within a country.

sql
SELECT template, device_class, APPROX_QUANTILES(inp, 100)[OFFSET(75)] AS p75_inp, COUNT(*) AS n
FROM rum.final WHERE metric = 'INP' AND mobile
GROUP BY 1, 2 HAVING n > 1000 ORDER BY template, device_class;
-- trade-off: crossing many dimensions quickly runs out of samples. Cross two
-- at a time, and only add a third for segments with ample traffic.

Expected outcome: a defensible statement such as "low-memory devices fail INP on product and search templates, but not on articles", which points directly at JavaScript cost on interactive templates.

Reading a segmented result How the pattern of a failing metric across device and network segments points to a likely cause. Reading a segmented result Pattern in segments Likely cause Fix family Fails on low-memory devices only main-thread work Less JS, yield, workers Fails on high-RTT connections only round trips Preconnect, early hints, fewer hops Fails on one template everywhere template code Profile that template Fails everywhere equally shared asset or server TTFB, global CSS/JS

Verification

Check that segment counts add up to the total and that the "unknown" share matches your non-Chromium traffic. Confirm stability: a segment's p75 should not swing wildly day to day if it has enough samples. After shipping a fix targeted at a segment (say, reducing hydration work for low-memory devices), that segment's p75 should move most; if every segment moves equally, the fix was broader than you thought — or the change was something else.

Using Client Hints for Server-Side Segmentation

Client Hints let the server see device signals on the request itself, which is useful for joining RUM with server logs and for serving lighter experiences. Request them with Accept-CH: Sec-CH-UA-Mobile, Sec-CH-UA-Model, Device-Memory, ECT, RTT (subject to browser support and permissions policy), and log them alongside the request id that you also put in the page for the beacon. This gives you device context for page views where the beacon never arrived — the visits most likely to be slow — and lets you compare server-side counts with received beacons by segment to measure delivery bias. Respect privacy: request only the hints you use, and do not combine them into fingerprint-like identifiers.

Turning Segments into Lab Profiles

Segmentation pays off twice: once in explaining the field number, and again in making lab testing realistic. For each segment that matters — "low-memory Android on 4G", say — record the p75 RTT, bandwidth and an approximate CPU slowdown, and create a named throttling profile from them in your Lighthouse CI and Playwright configuration. Running scripted tests under the profile of the failing segment reproduces its problems far more reliably than a generic "Slow 4G" preset, and the profile can be refreshed quarterly from the same RUM queries as your audience changes.

FAQ

Is it acceptable to collect device memory under privacy rules?

The values are coarse by design and are commonly treated as non-identifying technical data, but your organisation's privacy review decides. Keep them bucketed, do not combine many dimensions into a high-entropy profile, and document the purpose. Avoid collecting precise model names unless you have a concrete need.

Why does effectiveType say 4g for users on slow networks?

effectiveType is derived from recently observed round-trip times and throughput, capped at 4g. Most modern connections report 4g even when they are congested or high-latency. Bucketed rtt and measured TTFB are more discriminating for typical web traffic.

Should segments drive different experiences, not just analysis?

They can. Serving fewer non-essential widgets, smaller images or simpler animations to low-memory devices is a legitimate optimisation, sometimes called adaptive loading. Make sure the core content and functionality are identical, and that the segment detection fails safe (towards the lighter experience) when signals are missing.

How many segments are too many?

When most segments have too few samples for a stable p75. A practical limit is two dimensions crossed at a time — device class by template, or connection by country — with a minimum of around a thousand page views per cell per reporting period. Beyond that, use segmentation to generate hypotheses and confirm them with targeted lab tests rather than ever-finer field slices.