AnyLearn
All lessons
Businessintermediate

Peeking: How Watching Your Experiment Ruins It

The most natural behaviour in experimentation, checking results daily and stopping when they look significant, quietly destroys the statistical guarantee everyone thinks they have. This lesson shows the peeking mechanism with honest arithmetic, then the fixes: fixed-horizon discipline, group sequential designs, and always-valid inference.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 7

What the 5 percent actually promises

The standard experiment closes with a significance test: if the observed difference would arise by chance less than 5 percent of the time under no true effect, declare significance.

The fine print matters more than the ritual. That 5 percent false positive guarantee is a property of a procedure, not of a number, and the procedure it belongs to is: choose a sample size in advance, collect exactly that much data, test once. Every piece of the guarantee is conditional on the once.

Why once? Because the p-value's meaning is how often would chance produce this, under a defined way of looking. Look differently, look repeatedly, stop when you like what you see, and you are running a different procedure whose false positive rate is something else entirely, usually something much worse, while still stamping 5 percent on the report.

Key idea: statistical guarantees attach to the whole decision procedure, including when you look and why you stopped. Change the looking and you change the guarantee, even though every individual calculation along the way was performed correctly. That is what makes the trap in this lesson so quietly destructive: no step looks wrong.

Full lesson text

All 7 steps on one page, for reading, reference, and search.

Show

1. What the 5 percent actually promises

The standard experiment closes with a significance test: if the observed difference would arise by chance less than 5 percent of the time under no true effect, declare significance.

The fine print matters more than the ritual. That 5 percent false positive guarantee is a property of a procedure, not of a number, and the procedure it belongs to is: choose a sample size in advance, collect exactly that much data, test once. Every piece of the guarantee is conditional on the once.

Why once? Because the p-value's meaning is how often would chance produce this, under a defined way of looking. Look differently, look repeatedly, stop when you like what you see, and you are running a different procedure whose false positive rate is something else entirely, usually something much worse, while still stamping 5 percent on the report.

Key idea: statistical guarantees attach to the whole decision procedure, including when you look and why you stopped. Change the looking and you change the guarantee, even though every individual calculation along the way was performed correctly. That is what makes the trap in this lesson so quietly destructive: no step looks wrong.

2. The peek, anatomised

Here is the behaviour, performed daily in dashboards everywhere. The experiment launches. Each morning, someone checks the running results. The difference is not significant, so the experiment continues. One morning it crosses the threshold; the experiment is stopped and the win announced.

The corruption is in the asymmetry of the stopping rule. Under no true effect, the running p-value is a random walk that wanders as data accumulates: it dips and rises on noise. Checking daily and stopping at the first dip below 0.05 is fishing the walk for its minimum; the more mornings you look, the more chances noise gets to produce one qualifying dip. And the rule never stops early for the opposite reason, a day the treatment looks bad is just continue.

Stopping at the first favourable noise, never at unfavourable noise, and calling the result a 5-percent-grade finding: that is the peeking problem. Nothing was miscalculated; each day's p-value was arithmetically correct. What broke is that the correct answer to a different question, is it significant today, was substituted for the question the guarantee covers, is it significant at the pre-planned end.

The experimenters most at risk are the diligent ones: watching your experiment closely feels like rigour, and the stopping rule optimised for finding wins is indistinguishable, from inside, from responsiveness.

3. How bad does it get? The arithmetic

Put numbers on the inflation, with the assumptions stated honestly.

If k looks at the data were fully independent tests, the chance that at least one clears the 5 percent bar by luck alone would be one minus the chance that all k stay clean:

P(false positive)=1(10.05)kP(\text{false positive}) = 1 - (1 - 0.05)^k
False positive ceiling vs number of looks, alpha 0.05
%02040608015102030
Source: computed: 1-(1-0.05)^k, independent looks; accumulating-data peeks are correlated, so realised rates sit below this curve but well above 5%

The honesty caveat is in the caption: daily peeks at one accumulating dataset are correlated, today's data contains yesterday's, so the true inflation sits below this independent-looks ceiling. But it sits far above 5 percent: simulation studies in the experimentation literature consistently put continuous peeking at several times the nominal rate, and the qualitative conclusion needs no precision to be damning. A team that peeks daily and stops on significance is running experiments whose real false positive rate is a multiple of the one printed on the readout.

Combine that with the previous lesson's base rate, most tested ideas are truly null, and inflation lands on a population that is mostly nulls: a meaningful share of an undisciplined platform's celebrated wins are noise, promoted to roadmap.

4. Fix one: the fixed horizon, taken seriously

The oldest fix is to actually run the procedure the guarantee describes, and modern platforms enforce rather than request it.

Before launch, a power calculation sets the sample size: given the metric's variance, the minimum effect worth detecting, and the desired sensitivity, conventionally 80 percent power at 5 percent significance, the calculator outputs how many exposed users the experiment needs, and therefore roughly how long it must run. The experiment then runs to that horizon.

What about the dashboard in the meantime? The discipline is not blindness; it is separating monitoring from verdicts:

  • Guardrails may be watched, and acted on, continuously. Stopping because errors spiked or revenue cratered is safety, not peeking; what is forbidden is stopping early because the primary metric looks like a win.
  • Displays can enforce the separation: platforms hide the primary metric's significance until the horizon, show data-quality panels throughout, and label mid-flight numbers as unstable.
  • Whole weeks, not partial ones: traffic composition cycles weekly, weekday professionals, weekend browsers, so horizons are set in full weeks to sample the mix evenly.

The fixed horizon's real cost is inflexibility: a change that is catastrophically bad still runs its course unless a guardrail catches it, and a change that is obviously winning cannot ship early. Those costs are what the sequential designs in the next step were invented to price correctly.

5. Fix two: designs that budget the looking

If looking early has value, and it does, catching disasters, shipping clear wins sooner, the statistically honest move is to pay for it in the design.

Group sequential designs, imported from clinical trials, schedule a small number of interim analyses in advance, three looks, say, and spend the total 5 percent error budget across them along a chosen curve. Early looks face brutally strict thresholds, so only overwhelming effects clear them; the final look pays a slightly stricter bar than 0.05 as the price of the earlier chances. The whole procedure's false positive rate stays at 5 percent, by construction, and early stopping is legitimate when it happens.

Always-valid inference goes further, and is built for the dashboard age: methods in the confidence-sequence and sequential-testing family produce intervals and p-values designed to be checked at every moment, with the guarantee holding no matter when or why you stop. This is the machinery behind the continuous monitoring modes of modern experimentation platforms. The price is conservatism: at any fixed moment, an always-valid interval is wider than the classical one, so detecting the same effect needs more data. You purchase the right to peek freely, and the currency is sensitivity.

The choice between the three regimes, fixed horizon, scheduled interims, always-valid, is a real trade with no dominant answer: maximum sensitivity per user at the cost of rigidity, a structured middle, or full flexibility at a sensitivity discount. What is not on the menu is the free option the naive dashboard pretended to offer.

6. The same disease in other organs

Optional stopping is one member of a family, and recognising the family matters more than memorising the member. The shared pathology: choosing what to report based on what the data showed, then reporting it as if the choice had been fixed in advance.

Predict first

An experiment's primary metric shows no effect. The analyst slices by platform, country and user tenure, and finds a significant 6 percent lift among tablet users in one market. The team ships to that segment, citing the finding. What just happened?

The same structure hides in metric shopping, the OEC missed but engagement moved, in re-running flat experiments until one run clears, and in launching only the third variant that finally worked while forgetting the two nulls that calibrate it. The unified defence is pre-registration: metrics, segments, horizon and decision rule written before data, everything else labelled exploratory, and exploratory findings graduated only through fresh confirmation. It is bureaucracy exactly the way double-entry bookkeeping is bureaucracy: the boring structure that makes the numbers mean things.

7. What a healthy statistical culture looks like

Closing the lesson at the level where the problem actually lives, since peeking is a incentive problem wearing a statistics costume.

The platform enforces what policy cannot: horizons and sequential boundaries are configured at launch, primary-metric verdicts are unavailable until legitimate, guardrails alert automatically, and any early stop is recorded with its reason. The path of least resistance is the honest procedure.

The culture handles what the platform cannot:

  • Nulls are results. A team measured by shipped wins will find wins in the noise; a team credited for decisive answers, including this idea does nothing, has no reason to fish. The published base rates from the majors, most ideas fail, are the strongest culture tool available: they make a null unremarkable.
  • Effect sizes over verdicts. Reporting the confidence interval, we can rule out effects larger than half a percent, informs the next decision; a bare not significant merely disappoints.
  • Surprising wins get replicated, especially convenient ones, before the roadmap reorganises around them. One fresh confirmation filters most noise-wins at the cost of a week.

With the statistics honest, one hazard remains, and it is structural rather than procedural: everything so far assumed each user's outcome depends only on their own arm. Marketplaces, social products and shared infrastructure violate that assumption by design, and when they do, both arms contaminate each other and the clean difference-of-means quietly measures the wrong thing. That, plus the variance-reduction machinery that makes experiments affordably fast, is the final lesson.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. The 5 percent false positive guarantee of a standard significance test is conditional on what?
    • Using a sample of at least ten thousand users
    • The full procedure: sample size fixed in advance, data collected to that horizon, and the test performed once
    • The metric being normally distributed
    • Running a two-sided rather than one-sided test
  2. Why does stopping at the first significant daily reading inflate false positives?
    • Daily samples are too small for the test's assumptions
    • Significance thresholds tighten automatically over time
    • Under no true effect the running p-value wanders on noise, and stopping at its first favourable dip, never at unfavourable ones, fishes the walk for a qualifying moment
    • The dashboard recomputes variance incorrectly
  3. What does the computed curve 1-(0.95)^k represent, honestly stated?
    • The exact false positive rate of daily peeking
    • The power of the experiment after k days
    • The fraction of experiments that fail SRM after k looks
    • An upper-bound illustration assuming independent looks; correlated accumulating-data peeks land below it but still at several times the nominal 5 percent
  4. How do group sequential designs make early stopping legitimate?
    • They schedule interim analyses in advance and spend the total error budget across them, with strict early thresholds and a slightly stricter final bar
    • They double the sample size to compensate for looks
    • They only permit stopping for harm, never for benefit
    • They replace p-values with confidence intervals
  5. A flat experiment shows a significant lift only among tablet users in one market, found by searching twenty segments. What is the honest next step?
    • Ship to that segment, since the p-value cleared the threshold
    • Treat it as a hypothesis and run a fresh, pre-registered experiment on that segment, or correct for the multiplicity of slices examined
    • Extend the original experiment for another week
    • Merge the segment result into the overall average

Related lessons

Business
intermediate

Why Everything Gets Tested, and What a Test Actually Is

The companies famous for experimentation did not adopt it out of statistical enthusiasm: they adopted it because their own data showed most confident product ideas fail to improve the metrics they target. This lesson covers why observational product data misleads, what randomisation actually buys, the choice of randomisation unit, and the humbling base rates reported by the teams who measured.

7 steps·~11 min
Business
advanced

CUPED and Interference: Faster Experiments, and When Arms Contaminate Each Other

Two advanced problems decide how much an experimentation platform is actually worth. Variance: most product metrics are so noisy that detecting small effects takes painful sample sizes, and CUPED buys the reduction with data you already have. And interference: in marketplaces and social products the arms affect each other, so the measured difference misstates what full launch will do.

7 steps·~11 min
Business
intermediate

Assignment, Exposure, and the Smoke Detector Called SRM

Most wrong experiment results are not statistical subtleties; they are plumbing. This lesson covers how assignment actually works, hashing, not coin flips, why exposure must be logged at the moment of treatment, and the sample ratio mismatch check: the humble comparison of observed to expected group sizes that catches more broken experiments than any other single test.

7 steps·~11 min
Business
advanced

Selection Bias and the Deflated Sharpe Ratio

The statistical core of backtest overfitting: why the best of many trials is inflated even when nothing works, how much to discount it, and why finance needs a far higher significance bar than the usual one.

8 steps·~12 min