The tyranny of variance
The silent constraint on every experimentation programme is that product metrics are extraordinarily noisy. Revenue per user spans zero to enormous; sessions per week ranges from one to hundreds; and the effects worth detecting are small, a 1 percent lift on a core metric is often a large win at scale.
Detectable effect size shrinks only with the square root of the sample: to halve the minimum detectable effect, quadruple the users. For a metric whose standard deviation is large relative to its mean, powering an experiment to see 1 percent can demand millions of exposed users or months of runtime, and most companies have one of those at best.
This is not a pedantic concern; it decides what the platform can do. A team that needs six weeks per experiment runs eight experiments a quarter; a team that needs one week runs fifty. The entire compounding value of test everything rests on individual tests being cheap, which makes variance reduction, squeezing more sensitivity from the same users, among the highest-leverage machinery in the field.
Key idea: you cannot buy more users, but you can spend the noise smarter. The largest share of a metric's variance is usually not caused by the experiment at all: it is baseline differences between users that existed before the experiment started. Remove that share, and the same users answer harder questions.

