← Archive

Bayes' Theorem: The 99% Accurate Test That Is Wrong 91% of the Time

by ·July 25, 2026·9 min read·Science & Math
इस निबंध का पूरा हिंदी अनुवाद अभी तैयार नहीं है — नीचे का लेख अंग्रेज़ी में है। चित्रों के लेबल और साइट का बाकी हिस्सा हिंदी में दिख रहा है।

A disease affects 1 person in 1,000. There is a test for it that is 99% accurate.

You take the test. It comes back positive.

What is the chance you have the disease?

Most people — including, in published studies, a majority of doctors asked this question — answer around 99%. The actual answer is roughly 9%.

This is not a trick. Both numbers in the question are exactly what they appear to be. The gap between 99% and 9% comes from a single piece of information that most people leave out of the calculation entirely: how rare the disease was before you tested.

Getting this right is what Bayes' theorem is for. The mathematics has a fearsome reputation and the underlying idea is genuinely simple — it is just careful bookkeeping for changing your mind.

Why a 99% accurate test can still bewrongHealthy peopletested9,990False positives(1%)~100Actually sick10
Figure 1.Test 10,000 people for a disease that 1 in 1,000 has. The test catches all 10 real cases — but also flags about 100 healthy people. A positive result means roughly a 9% chance of being ill, not 99%.

Working through it with people, not algebra

Forget the formula. Imagine testing 10,000 people.

How many actually have the disease? It affects 1 in 1,000, so about 10 people.

How many of those does the test catch? It is 99% accurate, so it correctly flags about 10 of them. Call it all 10.

How many healthy people are there? 10,000 − 10 = 9,990.

How many of those does the test wrongly flag? The test is 99% accurate, meaning it is wrong 1% of the time. One percent of 9,990 is about 100 people.

Now count everyone who received a positive result: 10 genuinely ill, plus 100 healthy but wrongly flagged. That is 110 positive results, of which only 10 are real.

10 out of 110 is about 9%.

The reason the intuition fails is that healthy people vastly outnumber sick ones. Even a small error rate applied to a very large healthy group produces more false positives than there are true cases. The test's accuracy was never the whole story — the base rate was doing most of the work, and it was invisible in the way the question was posed.

Bayes in three steps, no algebraStart with the base ratehow common is it anyway?Add the new evidencehow much does it shift things?Get an updated beliefnot a yes/no answer
Figure 2.The theorem is just bookkeeping for changing your mind. Start from how likely something was before, adjust by how strong the new evidence is, and end with a revised probability.

What the theorem actually says

Stripped of notation, Bayes' theorem says:

Your updated belief = how likely it was before × how much this evidence shifts it.

Three ingredients:

The prior. How likely was this before you saw the evidence? For the disease, 1 in 1,000. This is the piece people skip, and skipping it is the base rate fallacy.

The evidence. How much more likely is this evidence if the thing is true than if it is false? A test that flags 99% of sick people but also 1% of healthy people is informative but not decisive — it is about 99 times more likely to fire for a sick person, which sounds overwhelming until you remember there are 999 times more healthy people.

The posterior. The revised probability. And crucially it is a probability, not a verdict. Bayesian reasoning does not output "yes" or "no." It outputs "more likely than before, by this much."

That last point is the practical heart of it. The theorem models belief as something you adjust rather than something you switch on and off. A positive test moved you from 0.1% to 9% — a ninety-fold increase, genuinely significant, and still nowhere near certainty. Both facts are true at once, and everyday reasoning tends to collapse them into one or the other.

Two things that decide the answerHow strong is the evidence?Strong evidence,rare thing:still doubtStrong evidence,common thing:trust itWeak evidence,rare thing:ignoreWeak evidence,common thing:maybeHow rare is it to begin with?
Figure 3.Evidence never speaks alone. The same test result means very different things depending on how rare the condition was before you tested — which is the single most ignored fact in everyday reasoning.

Using it without the mathematics

You will rarely calculate this. But three habits fall out of it that improve reasoning immediately.

Always ask how common the thing was to begin with. Before assessing any piece of evidence, ask what fraction of cases look like this anyway. A striking symptom that matches a rare disease usually still means the common illness. An extraordinary claim genuinely does require stronger evidence — not because of a rhetorical rule, but because the prior is low and the arithmetic demands more to overcome it.

Ask what else would produce this evidence. Evidence is only informative to the extent it is more likely under your hypothesis than under the alternatives. If a result would occur just as often either way, it tells you nothing, however dramatic it seems. This is the question that punctures most conspiracy reasoning and most confident business post-mortems: the evidence fits the story, but it also fits three duller stories nobody considered.

Update in proportion, and keep updating. Strong evidence should move you a lot, weak evidence a little, and neither should move you to certainty. This is the direct antidote to confirmation bias, which treats supporting evidence as decisive and contrary evidence as dismissible — the exact opposite of proportional updating.

There is one honest limitation worth stating. Bayes requires a prior, and you often do not have a good one. Where the base rate is genuinely unknown, people substitute a guess and the output inherits that guess's error, dressed in mathematical clothing. The theorem is a tool for reasoning consistently; it cannot manufacture information you do not have.

Still, the disease example is worth carrying around. It is not an artificial puzzle — it is the everyday structure of screening, security alerts, fraud flags, and any search for something rare. Whenever a system looks for a needle in a haystack, most of what it finds will be hay, no matter how good the detector is.

Dr Nadeem Khudboddin Shaikh
Dr Nadeem Khudboddin Shaikh
Ex–Wells Fargo · Ex–Goldman Sachs · Columbia University alumnus