Epidemiology: Everything Turns on One Number Crossing One
Epidemiology is the study of how disease moves through populations, and most of it turns on one number.
R is the average number of further cases each case produces. If R is above one, each generation of infections is larger than the last and the outbreak grows. If it is below one, each generation is smaller and it fades out.
The threshold at one is what makes this the central quantity. It is not a smooth relationship where higher is worse — it is a switch between two completely different behaviours, growth and decay.
And because the effect compounds across generations of infection, small differences in R produce enormous differences in outcome. An R of 1.1 and an R of 0.9 sound similar and describe opposite futures.
Why R is not a property of the disease
The most common misunderstanding is treating R as a fixed characteristic, like a chemical's boiling point. It is not. It is a property of a disease in a particular population under particular conditions, and it has roughly three components:
How easily transmission happens on contact.
How many contacts people have — which depends on density, behaviour, and season.
How long someone remains infectious — which depends on the disease and on whether they know they are infected.
Only the first is intrinsic. The others change with behaviour, which is why the same disease has very different R values in different places and at different times, and why quoted values should always be read as describing a context.
This is also why interventions work: each targets one component. Reducing contacts lowers the second. Isolating cases shortens the third. Vaccination reduces the pool of people transmission can reach.
Herd immunity falls directly out of the arithmetic. If enough of a population cannot transmit, the average case cannot produce a further case even if contacts continue as normal. The threshold depends on R — higher R requires a higher proportion — and it is a threshold for decline, not a guarantee nobody gets infected.
The average hides most of what matters
R is an average, and for many diseases the distribution behind it is extremely uneven.
A small proportion of cases and settings frequently account for most transmission, while a majority of infected people pass it to nobody. This is a power law pattern rather than a bell curve, and it changes what control measures make sense.
If transmission were evenly distributed, uniform measures would be the only option. Where it is concentrated in identifiable circumstances — particular settings, particular kinds of gathering — targeted measures can achieve much of the effect at a fraction of the cost.
The practical implication is that two diseases with the same R can require completely different responses, depending on the shape of the distribution behind the average and on one further variable: whether people transmit before they feel unwell.
That timing question is decisive. If infectiousness begins after symptoms, isolating sick people is highly effective and outbreaks are relatively controllable. If people transmit before feeling ill, isolation after symptoms arrives too late for a large share of transmission, and control requires measures applied to people who appear healthy — which is far more costly and far harder to sustain.
What the field can and cannot tell you
It is good at direction and mechanism. Whether an outbreak is growing, roughly which settings drive spread, what a given intervention does to which component of R.
It is much weaker at precise prediction. Models require assumptions about behaviour, and behaviour changes in response to the situation being modelled — including in response to the model's own publication. A forecast that changes what people do has invalidated itself, which is a genuine methodological problem rather than an excuse.
Early estimates are unreliable, and for a structural reason. During the early phase of an outbreak, cases are few and testing is incomplete, so parameter estimates come from small, biased samples. This is precisely the situation the law of large numbers warns about, and it is why early figures are revised so often.
Detection is a screening problem. Testing a large population for something rare produces mostly false positives even with a good test, exactly as Bayes' theorem describes. Interpreting a positive result requires knowing how common the condition currently is, which is why the same test means different things at different stages of an outbreak.
Attribution is genuinely hard. When several measures are introduced together and cases fall, separating their contributions is difficult and often not possible with confidence. Strong claims that a particular measure caused a particular decline usually rest on assumptions doing most of the work — which is worth remembering, because such claims are made frequently and confidently in both directions.
The honest summary is that the core arithmetic is simple and robust, the parameters are context-dependent and hard to measure well, and the gap between those two facts is where most public confusion about the subject lives.