The Law of Large Numbers: Small Samples Are Wild
Flip a fair coin ten times and getting seven heads would not surprise anyone. It happens often.
Flip it a hundred thousand times and getting seventy per cent heads would be astonishing — so unlikely that you would reasonably suspect the coin.
The underlying probability never changed. What changed is how much room chance has to move the result. This is the law of large numbers: as the number of independent trials increases, the observed average converges toward the true underlying value.
It sounds obvious stated plainly. Its consequences are not obvious at all, and the ways it gets misapplied are among the most common errors in reasoning about data.
Small samples are volatile by nature
The most useful immediate implication is about variability, not averages.
Small samples produce extreme results routinely. Not because anything is wrong with them, but because with few observations, chance has not had the opportunity to cancel itself out. Three coin flips can easily give you all heads. Three thousand cannot.
This has a counterintuitive consequence that catches people constantly: the most extreme results in any dataset tend to come from the smallest samples.
If you rank regions by disease rate, the highest and lowest rates will both tend to be sparsely populated areas — not because small places are healthier or sicker, but because small numbers vary more. If you rank schools by test scores, small schools appear disproportionately at both the top and the bottom.
Interpreting that as evidence about small places, or small schools, is a real and expensive mistake that has driven actual policy in the past. The pattern is a property of sample size, not of the thing being measured.
This is also the mechanism underneath regression to the mean. An extreme result usually reflects a small sample, real effects, and favourable chance together — and the chance component does not repeat.
Why dramatic findings so often shrink
The same logic explains a pattern that recurs across research fields.
A study with a small number of participants reports a large, striking effect. It gets attention. Larger follow-up studies find a much smaller effect, or none.
The disappointing follow-up is usually the more accurate one. Small studies can only detect large effects, and when the true effect is modest, the only way a small study finds statistical significance is if chance happened to inflate it. So among published small studies with dramatic findings, an unusually high proportion are overstatements.
This is not fraud and usually not incompetence. It is the arithmetic of small samples interacting with publication practices that favour interesting results — which is why replication matters and why a single striking study is weak evidence regardless of how good the story is.
The practical filter: when you see a surprising finding, look for the sample size before anything else. A dramatic result from a few dozen participants and a dramatic result from tens of thousands are very different claims.
Two things the law does not say
It says nothing about the next outcome. This is the gambler's fallacy: after a run of heads, people feel tails is "due" because the average must even out. The coin has no memory, and each flip remains fifty-fifty. Convergence happens not because past deviations get corrected, but because they get diluted — a run of five extra heads is enormous in ten flips and irrelevant in a million.
Nothing pulls the average back. The early result simply stops mattering.
It requires independent trials. This condition is easy to overlook and it is where the law fails in practice. If outcomes influence each other, averages need not converge to anything stable.
This matters enormously in finance and risk. Loan defaults are not independent — a recession makes many borrowers default at once. Insurance claims are not independent when a single storm hits thousands of policyholders. Models that assume independence work well in ordinary conditions and fail precisely when correlations spike, which is exactly when the results matter.
That failure mode connects directly to black swan events and fat-tailed distributions: in domains where extremes dominate, averages converge slowly if at all, and a sample that has not yet included a rare large event will systematically understate both the mean and the risk.
So the honest summary is narrower than the slogan. The law of large numbers reliably describes long-run averages of independent trials with well-behaved distributions. It tells you nothing about the next event, and it quietly stops applying when things are correlated or when the tails are heavy — which happens to describe a good deal of the world people most want to predict.