A platform team wants a pull request to fail when a route's field Interaction to Next Paint at the 75th percentile regresses. The route gets about 4000 page views a day and real-user monitoring samples 10% of sessions. Roughly how long before a shift from 25% to 30% of interactions above 200 ms is detectable, and what does that rule out?
Show the full answer Hide the answer
The assumptions, stated
- A p75 at or below 200 ms is the same statement as "at most 25% of measured interactions exceed 200 ms", so the quantity to detect is a change in that proportion.
- The regression to catch moves that proportion from 0.25 to 0.30. Five percentage points is about the smallest shift a team would call a regression rather than noise.
- One INP value is reported per page visit, and roughly 60% of visits to a list route produce at least one qualifying interaction. This assumption is the shakiest and it dominates the answer.
- 95% confidence, 80% power, two-sided.
The arithmetic
Sample size per arm for two proportions:
n ≈ (z(α/2) + z(β))² × [p1(1-p1) + p2(1-p2)] / (p2 - p1)²
= (1.96 + 0.84)² × [0.25×0.75 + 0.30×0.70] / 0.05²
= 7.84 × 0.3975 / 0.0025
≈ 1250 measured interactions per arm
Measured interactions available per day:
4000 page views × 10% sampled = 400 sampled visits
400 × 60% with a qualifying interaction ≈ 240 INP values per day
1250 / 240 ≈ 5.2 days per arm
Comparing a before window against an after window needs both, so about 10 to 11 days for one verdict — and that assumes nothing else changed in the traffic mix during those eleven days.
The number and its range
Call it 10 days, with a plausible range of 7 to 20 days. The lower end assumes almost every visit interacts; the upper end assumes a third do, or that device and network mix shifts enough to need stratification.
Which assumption dominates the error
The interaction rate per visit, which can plausibly be 30% or 90% and swings the result by three times. Effect size is next, and it is non-linear: halving the detectable regression to 2.5 points multiplies the required sample by four, pushing the verdict past 40 days. Chasing small field regressions is arithmetically hopeless at this traffic level, which is why CrUX itself reports on a 28-day rolling window.
What this rules in and out
- Ruled out: field INP as a blocking per-pull-request gate on anything except the two or three highest-traffic routes. A gate that needs ten days of data is not a gate; it is a weekly report wearing a gate's costume.
- Ruled in, per pull request: synthetic measurement on a named device-class baseline, decoded-byte budgets per route, and a long-task count for the primary interaction. These are deterministic and answer in minutes.
- Ruled in, per release: field INP as a trend with release annotations, reviewed weekly, with a rollback decision owned by a human.
- The cheap fix first: raise the sample rate on the routes you intend to gate. Going from 10% to 100% on one route cuts the wait to about one day per arm and costs almost nothing, because the volume is small — which is the same fact that caused the problem.
Common weak answers
- "Use a longer window." A longer window gives power and guarantees the signal arrives after the deploy it describes, which defeats the purpose of a gate.
- "Compare p75 values directly." A percentile estimate from 240 samples has a wide interval of its own; converting to the proportion above the threshold is what makes the test tractable.
- "Alert on any movement." At this sample size the day-to-day noise band is several percentage points, so the alert fires constantly and gets muted within a fortnight.