Power Laws: When the Average Describes Nobody
Imagine a room holding a hundred randomly chosen people. Now imagine the world's tallest person walks in. The average height of the room barely moves — perhaps a centimetre. The tallest human being who has ever lived was roughly two and a half times the height of an average adult, and no taller person will ever exist, because bone and cardiovascular limits forbid it.
Now empty the room and refill it with a hundred random people, and let the world's wealthiest person walk in. The average wealth of the room does not rise by a fraction. It rises by a factor of millions. Every other person in the room becomes, statistically, a rounding error.
Same species. Same room. Completely different mathematics — and almost all statistical intuition, including the intuition taught in introductory courses, is built for the first case and quietly fails in the second.
Two distributions, two universes
A normal distribution — the bell curve — describes quantities that cluster around a typical value with rapidly thinning tails. Human height, measurement error, the sum of many independent dice. These have three comfortable properties: the average is a meaningful summary, extreme deviations are vanishingly rare, and no single observation can dominate the total.
A power law describes quantities where the relationship between size and frequency follows a constant ratio: each tenfold increase in magnitude is a fixed factor less common. Wealth, city populations, earthquake energy, book sales, casualties in conflicts, venture returns, word frequency. These have three deeply uncomfortable properties: the average may be nearly meaningless, extremes are far more common than a bell curve predicts, and a single observation can exceed the sum of everything else.
The reason this matters practically is that most people's statistical reflexes — formed on bell curves — actively mislead in power-law domains. "That's a once-in-a-century event" is a bell-curve sentence. In a power-law world, events of that magnitude may be once-a-decade, and the confident use of the word "unprecedented" usually reveals that someone fitted the wrong distribution rather than that something genuinely impossible occurred.
The practical diagnostic is a single question: could one observation dwarf all the others combined? If a single customer, a single book, a single investment, or a single outage could exceed everything else put together, you are in power-law territory and should reason accordingly.
Where power laws come from
Power laws are not arbitrary curiosities; they are generated by a specific and very common mechanism — preferential attachment, sometimes phrased as "the rich get richer."
The mechanism is a reinforcing feedback loop. Whoever has more of something finds it easier to acquire more. A website with many links is more discoverable and attracts more links. A wealthy person can invest and compound. A popular researcher's papers are read and therefore cited more, which makes them more read. A city with more jobs attracts more people, which creates more jobs.
Two consequences follow, and both are uncomfortable.
Small early advantages become enormous later ones. Because the loop compounds, an initially trivial lead — often attributable to timing, luck, or noise — can produce a final distribution of staggering inequality without anyone at the top being proportionally more capable. This is not a claim that skill is irrelevant. It is a claim that outcome ratios wildly exceed skill ratios wherever compounding operates.
The distribution is stable even as its members change. The specific identities at the top may churn constantly while the shape remains fixed. Analysing why a particular individual or firm reached the top often explains far less than recognising that the structure guarantees someone would.
Reasoning correctly in a power-law world
Averages become misleading, and medians are usually better. Average income in a power-law economy describes almost nobody; the median describes a typical person. Any time a distribution is skewed, reporting the mean is either a mistake or a rhetorical choice.
Optimise for exposure to the tail, not for the typical case. This is the entire logic of venture investing: a portfolio where ninety per cent of positions fail can be enormously profitable if the remainder includes one extreme outcome. Strategies that minimise variance also systematically eliminate access to the upside tail, which is fatal in a domain where the tail is where all the return lives.
Conversely, on the downside, survive the tail. If a single event can exceed everything else combined, then risk management means asking not "what is the expected loss?" but "what is the largest loss I can survive?" This is why margin of safety matters more than expected-value calculations wherever fat tails exist — an expected value is an average, and averages are the wrong tool here.
Beware sample-size intuitions. In normal-world statistics, more data reliably converges on the truth. In power-law domains, the largest and most consequential observations are rare, so a sample that has not yet contained one will systematically understate both the mean and the risk. A decade of calm is not strong evidence of safety; it may simply be a sample that has not yet included the event that defines the distribution.
Do not confuse a power law with a moral judgement. That outcomes are extremely unequal tells you about the generating mechanism, not about desert. Preferential attachment produces vast inequality from near-identical starting participants. Any argument that reads outcome distribution as a direct measure of merit has skipped the mathematics.
The tallest person who will ever live is already alive, or has been. There is no comparable ceiling on wealth, on a book's readership, or on the magnitude of the next financial crisis — and reasoning about the second category with tools designed for the first is one of the most consequential errors available.