How to Sample RUM Data Without Biasing Your p75

This guide sits within RUM Beacons & Field Data Collection in Core Web Vitals & Measurement. Real-user monitoring at scale is expensive: a site with 50 million page views a month sends 50 million beacons, each a few kilobytes, plus the storage and query cost of keeping them. Sampling is the obvious answer, and done carelessly it quietly changes the number you are trying to measure.

The p75 is a statistic of the distribution. Sampling preserves it only if the sampled visits are representative — if every kind of visit (slow device, fast device, first view, deep session) is equally likely to be kept. Common sampling shortcuts violate that: sampling by page view rather than by session, sampling only after a metric is known, sampling at different rates by page, or dropping beacons that fail to send on slow connections. Each skews the distribution, and the p75 moves with it.

Biased vs unbiased sampling Comparison of sampling practices that skew the 75th percentile with practices that preserve it. Biased vs unbiased sampling Skews the p75 • Random per page view (sessions split) • Decide after seeing the metric value • Different rates per template, unweighted • Beacons lost on unload not counted Preserves the p75 • Deterministic per session id • Decide before any metric is observed • Per-template rates with weights stored • Delivery measured and corrected

Rapid Diagnosis

  • Compare sampled and unsampled p75. Run 100% collection for a day on a subset of traffic and compare with the sampled pipeline. A gap of more than a few percent indicates bias.
  • Check where the sampling decision is made. If it happens inside a metric callback (onLCP(v => Math.random() < 0.1 && send(v))), sessions contribute some metrics and not others — inconsistent.
  • Check per-template rates. If product pages are sampled at 10% and checkout at 100%, an unweighted origin-level p75 over-represents checkout.
  • Check delivery rates by device. Compare beacons received against page views from server logs, by device class. Low-end devices often lose more beacons.

Root Cause Analysis

1. Page-view sampling breaks session-level metrics. INP and CLS are per page visit, but analysis often joins visits into sessions. Random per-page-view sampling keeps fragments of sessions, distorting anything computed across them.

2. Value-dependent sampling. Deciding to send based on the value ("only send slow LCPs to save cost") is no longer a sample of the distribution. Any percentile computed from it is meaningless.

3. Unweighted stratified sampling. Sampling different segments at different rates is efficient, but every aggregate must weight each beacon by the inverse of its sampling rate.

4. Survivorship bias in delivery. Beacons sent late (on pagehide or visibilitychange) are lost more often on slow, low-memory devices and poor networks — exactly the visits that sit above p75.

p75 LCP under different sampling schemes (same traffic) Bar chart of mobile p75 LCP computed from full data and from four sampling schemes, showing which schemes bias the result. p75 LCP under different sampling schemes (same traffic) Full data (truth) 2840ms Session-based 10% 2830ms Per-template rates unweighted 2610ms Send only if value over 2s 3420ms Unload-only beacons 2550ms

Step-by-Step Resolution

1. Sample by session, deterministically, before measuring

Derive the sampling decision from a stable session identifier, once, at page start. Every page view in the session gets the same decision; every metric in the page view is sent or not as a unit.

javascript
// Deterministic session sampling: the same session id always gets the same answer.
function sampled(sessionId, rate) {
  let h = 2166136261;                                   // FNV-1a hash
  for (let i = 0; i < sessionId.length; i++) h = Math.imul(h ^ sessionId.charCodeAt(i), 16777619);
  return (h >>> 0) / 2 ** 32 < rate;
}
const RATE = 0.1;
export const inSample = sampled(getSessionId(), RATE);
// trade-off: session-level sampling gives fewer independent units than page-
// view sampling at the same rate, so confidence intervals are wider for small
// segments. Raise the rate for low-traffic templates (and weight, step 2).

Expected outcome: sampled p75 matches full-data p75 within normal statistical noise.

2. Store the sampling rate on every beacon and weight aggregates

javascript
send({ ...metrics, sampleRate: RATE, template });
// SQL: weighted percentile — each row stands for 1/sampleRate page views.
// trade-off: most warehouses lack a native weighted percentile; approximate by
// expanding weights into histogram bins (step 4) rather than duplicating rows.

Expected outcome: you can sample rare templates at 100% and common ones at 5% while keeping origin-level aggregates correct.

3. Measure delivery and send early where possible

Send what you know as soon as it is final, rather than everything on page hide. LCP is final after the first interaction; CLS and INP must wait for hide. Track delivery rate per segment by comparing beacons with server-side page views.

javascript
import { onLCP, onCLS, onINP } from 'web-vitals';
onLCP((m) => inSample && sendNow(m));                     // final early — send immediately
onCLS((m) => inSample && queueForHide(m));
onINP((m) => inSample && queueForHide(m));
// trade-off: sending LCP separately costs an extra request per page view. On
// very high-traffic sites, batch it with the hide beacon but also flush on
// a 10s timer so long visits do not hold it until unload.

Expected outcome: fewer lost LCP values on slow devices; measured delivery gaps you can correct for or at least report.

4. Aggregate with histograms for weighted percentiles

Bucket values into fixed bins (for example, 100ms bins up to 10s for LCP), sum weights per bin, and read the percentile from the cumulative weighted histogram. This scales to billions of beacons and handles weights naturally.

sql
WITH bins AS (
  SELECT template, LEAST(FLOOR(lcp / 100) * 100, 10000) AS bin, SUM(1 / sample_rate) AS w
  FROM rum.beacons WHERE device = 'mobile' GROUP BY 1, 2
), cum AS (
  SELECT template, bin, SUM(w) OVER (PARTITION BY template ORDER BY bin) / SUM(w) OVER (PARTITION BY template) AS cdf
  FROM bins
)
SELECT template, MIN(bin) AS p75_lcp FROM cum WHERE cdf >= 0.75 GROUP BY template;
-- trade-off: 100ms bins cap precision at ±50ms. Use finer bins near the
-- threshold if you need to know whether p75 is 2480ms or 2520ms.

Expected outcome: correct weighted p75 per template and per origin at a fraction of the storage cost.

A sampling pipeline that keeps p75 honest Four stages of an unbiased RUM sampling pipeline, from the session-level decision to weighted percentile queries. A sampling pipeline that keeps p75 honest Decide per session at page start Hash of a stable session id against the rate Tag every beacon with its rate sampleRate travels with the data, per template if rates differ Send final values early where possible LCP immediately; CLS and INP on hide Aggregate with weighted histograms Each beacon counts as 1/sampleRate page views 1 2 3 4

Verification

Run an A/A comparison: for one week, collect a 100% sample on 5% of sessions alongside your production sampling. Compute p75 for LCP, INP and CLS per template from both. The difference should fall within the confidence interval for the sampled data; if it does not, check sampling decision placement and weighting. Repeat whenever you change rates.

How Much Sampling Is Safe?

The precision of a percentile estimate depends on the number of samples, not the fraction of traffic. As a rule of thumb, a few thousand page views per segment per period gives a p75 stable to within a few percent; below a few hundred, percentiles swing noticeably from day to day. Work backwards: decide which segments you need to report (template × device × week, say), estimate their traffic, and choose the lowest rate that keeps each above a few thousand samples. Where a segment is too small at any affordable rate, sample it at 100% and rely on the weights to keep origin-level numbers honest. Report sample counts alongside every percentile so readers can judge reliability.

Common Sampling Mistakes in Practice

  • Changing the rate mid-period without recording it. A dashboard computing a 28-day p75 across a rate change from 10% to 2% is fine only if every beacon carries its own rate and aggregates are weighted.
  • Sampling in the tag manager. A sampling trigger configured in a tag manager often evaluates per page view and after other tags have run, combining both of the worst choices.
  • Different rates per browser. Sampling Safari at a lower rate because "it has fewer metrics" skews any cross-browser comparison unless weighted.
  • Dropping beacons server-side under load. Ingestion pipelines that shed load randomly during traffic spikes under-represent peak periods — when pages are often slowest. Prefer client-side sampling decided in advance.

FAQ

Is CrUX itself sampled?

CrUX collects from opted-in Chrome users only, which is a different kind of selection — it is not a random sample of all your visitors, and it excludes other browsers entirely. That is one reason CrUX and your own RUM differ. Your own sampling should aim to represent your traffic, not to mimic CrUX.

Can I sample attribution data at a lower rate than metrics?

Yes, if the attribution sample is itself drawn session-deterministically from the metric sample — for example, a second hash threshold. Attribution then describes a representative subset of the same visits. Never select attribution sessions based on whether their metric was bad.

What about bots and synthetic traffic?

Exclude them before sampling, using user-agent and behavioural filters, so they neither consume sample budget nor distort percentiles. Synthetic monitoring traffic from your own tools should be tagged and excluded explicitly; it is usually fast and pulls p75 down.