How to Compute p75 Correctly from Raw RUM Beacons

This guide covers the aggregation side of RUM Beacons & Field Data Collection, part of Core Web Vitals & Measurement. Collecting beacons is the easy half. Turning them into a p75 that you can defend in a performance review — and that behaves like the assessed Core Web Vitals number — requires decisions most dashboards make implicitly and often wrongly.

The four that matter most: what one row represents (a page view, not a beacon), which value per page view counts (the final one), how the percentile is computed (and over which population), and which segments get their own p75 rather than being averaged together. Getting any one wrong produces numbers that drift from reality by hundreds of milliseconds, and teams then argue about whether a release helped.

From raw beacons to a trustworthy p75 Pipeline from raw beacons through deduplication, final-value selection and segmentation to a percentile per segment. From raw beacons to a trustworthy p75 Raw beacons many per page view Deduplicate one row per page view + metric Final value last reported value Segment device, template Percentile p75 per segment

Rapid Diagnosis

  • Count rows per page view. SELECT page_view_id, metric, COUNT(*) ... HAVING COUNT(*) > 1. Duplicates mean reportAllChanges or retries are inflating counts.
  • Check which value is used. If CLS or INP is reported incrementally, using the first or an average of reported values understates the final value.
  • Check the percentile function. Approximate quantile functions are fine at scale; averaging per-day percentiles is not.
  • Check the population. Are desktop and mobile mixed? Are bots excluded? Are prerendered or bfcache-restored page views treated consistently?

Root Cause Analysis

1. Beacons are not page views. With reportAllChanges: true, the web-vitals library reports CLS and INP every time they change. Each report is a beacon; only the last one per page view is the final value.

2. Averaging percentiles. A weekly p75 computed as the average of seven daily p75s is not the weekly p75. Percentiles do not average; compute them over the pooled data.

3. Mixing populations. Desktop and mobile distributions differ so much that a pooled p75 describes neither. Core Web Vitals are assessed per form factor for this reason.

4. Inconsistent handling of special navigations. Prerendered pages report near-zero LCP; bfcache restores report a fresh set of metrics. Including or excluding them inconsistently between periods creates false trends.

How modelling mistakes move a p75 INP Bar chart showing the reported p75 INP for the same data under correct modelling and three common mistakes. How modelling mistakes move a p75 INP Correct (final value per view) 236ms All beacons counted 188ms First reported value 151ms Mobile + desktop pooled 174ms 200ms INP Each mistake flatters the number — the page appears to pass while the correctly modelled p75 fails.

Step-by-Step Resolution

1. Give every beacon a page-view id and a metric id

The web-vitals library provides metric.id, unique per metric per page view; add your own page-view id for joins with other events.

javascript
import { onLCP, onINP, onCLS } from 'web-vitals';
const pageViewId = crypto.randomUUID();
const send = (m) => queue({ pv: pageViewId, id: m.id, name: m.name, value: m.value,
                           delta: m.delta, nav: m.navigationType, ts: Date.now() });
onLCP(send); onINP(send, { reportAllChanges: true }); onCLS(send, { reportAllChanges: true });
// trade-off: reportAllChanges improves delivery for long visits (you get a
// value even if the final beacon is lost) but multiplies beacons. Deduplicate
// on ingest, as below, or the counts are meaningless.

Expected outcome: every row can be attributed to exactly one page view and one metric instance.

2. Keep only the final value per metric instance

sql
CREATE OR REPLACE VIEW rum.final AS
SELECT * EXCEPT(rn) FROM (
  SELECT *, ROW_NUMBER() OVER (PARTITION BY id ORDER BY ts DESC) AS rn
  FROM rum.raw
) WHERE rn = 1;
-- trade-off: "latest timestamp wins" assumes client clocks are monotonic within
-- a page view, which holds for Date.now() in practice. If you send deltas
-- instead, SUM(delta) per id reconstructs the final value without ordering.

Expected outcome: one row per metric per page view, carrying its final value.

3. Compute percentiles over pooled, segmented data

sql
SELECT device, template, name AS metric,
       APPROX_QUANTILES(value, 100)[OFFSET(75)] AS p75,
       COUNT(*) AS page_views
FROM rum.final
WHERE ts >= UNIX_MILLIS(TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 28 DAY))
  AND nav IN ('navigate', 'reload')           -- see step 4 for the others
GROUP BY 1, 2, 3
HAVING page_views >= 500;
-- trade-off: the 500-row floor hides small segments whose p75 would be noise.
-- Report them in an "other" bucket rather than silently dropping traffic.

Expected outcome: p75 values per device and template over a 28-day window, comparable in spirit with CrUX's assessment.

4. Decide and document how special navigations count

The web-vitals library reports navigationType as navigate, reload, back-forward, back-forward-cache, prerender or restore. Report the main p75 over navigate and reload, and break out bfcache restores and prerenders separately — they are genuinely fast experiences, but mixing them in hides changes to normal loads.

Expected outcome: trends that move only when real load performance moves.

Handling each navigation type How to treat each web-vitals navigation type when computing the headline p75. Handling each navigation type navigationType In headline p75? Why navigate yes Standard page loads reload yes Real loads users wait for back-forward-cache separate series Near-instant; would mask regressions prerender separate series LCP from activation can be ~0 restore separate series Discarded tab restored

Verification

Cross-check your pipeline against an independent computation: export one day of raw beacons, compute the final value per metric id and the p75 in a notebook, and compare with the dashboard. They should match exactly (or within approximate-quantile error). Then compare your origin-level mobile p75 with CrUX's: they will not match exactly — CrUX is Chrome-only and opt-in — but a large and persistent gap usually reveals a modelling bug rather than a population difference.

Confidence Intervals: Is the Change Real?

A p75 computed from a sample has uncertainty, and week-over-week changes of a few percent are often noise. For a quick estimate, bootstrap: resample page views with replacement a few hundred times, compute p75 for each resample, and take the 2.5th and 97.5th percentiles of those values as a 95% interval. If the intervals for "before" and "after" overlap substantially, do not claim an improvement. For high-traffic templates intervals are narrow and small changes are detectable; for low-traffic ones, compare over longer windows. Displaying the interval on dashboards prevents a great deal of arguing over noise.

A Reconciliation Checklist

When your RUM p75 and another source disagree, work through these in order before assuming either is wrong:

  1. Population: same browsers, same form factor, bots excluded in both?
  2. Window: same date range, same 28-day versus calendar-week definition?
  3. Unit: one value per page view in both, final values only?
  4. Navigation types: prerender and bfcache treated the same way?
  5. Sampling: weights applied if rates differ by segment?
  6. Percentile method: exact or approximate, and pooled rather than averaged?

Most disagreements resolve at steps 1 or 3. Writing the answers down for your pipeline once makes every future comparison faster.

FAQ

Should I use the delta or the value field from web-vitals?

value is the current total for the metric; delta is the change since the last report. If you keep the latest report per metric id, use value. If your pipeline can only append, summing delta per id gives the same final value without needing ordering — useful for streaming aggregation.

Is the median ever the better statistic?

For tracking typical experience and detecting broad regressions, the median is more stable and responds to changes that affect everyone. For the Core Web Vitals pass/fail question, p75 is the defined statistic. Many teams chart both: the median for "did everything get slower?" and p75 for "are we passing?".

How should CLS zeros be handled?

Keep them. A large share of page views with CLS of exactly zero is common and legitimate; excluding them shifts the p75 upward and misrepresents the experience. Make sure your beacon sends CLS even when it is zero — some custom implementations only send when a shift occurs.

Should INP be computed from every interaction or from the per-page value?

From the per-page value. INP is defined per page visit as the worst interaction (ignoring one outlier per 50), so each page view contributes exactly one INP number, and the p75 is taken across page views. Computing a percentile over all individual interactions answers a different question — how fast a typical interaction is — which is useful for diagnosis but will be much lower than the assessed INP and should be labelled differently.