← Archive

Simpson's Paradox: True in Every Group, False Overall

by ·July 25, 2026·9 min read·Science & Math
इस निबंध का पूरा हिंदी अनुवाद अभी तैयार नहीं है — नीचे का लेख अंग्रेज़ी में है। चित्रों के लेबल और साइट का बाकी हिस्सा हिंदी में दिख रहा है।

Here is a genuinely uncomfortable statistical fact.

Two medical treatments are compared. Treatment A has a better success rate for patients with small kidney stones. Treatment A also has a better success rate for patients with large kidney stones.

Combine all the patients together, and Treatment B has the better overall success rate.

Nothing has been miscalculated. All three statements are true at once. This is Simpson's paradox, and it is the clearest demonstration available that a number can be entirely accurate and still point you the wrong way.

Better in every group, worse overallTreatment A,small stonesbetterTreatment A,large stonesbetterTreatment B,overalllooks better!
Figure 1.A treatment can win in every subgroup and still lose when the groups are combined. Nothing was miscalculated — the totals are simply weighted by unequal group sizes.

How the reversal happens

The mechanism is unequal group sizes combined with unequal difficulty.

Suppose large stones are simply harder to treat — both treatments do worse on them. Now suppose the two treatments were not given to similar patients. Treatment A, being the more invasive option, was used mostly on the difficult large-stone cases. Treatment B was used mostly on the easier small-stone cases.

So Treatment A's overall number is dominated by hard cases, and Treatment B's overall number is dominated by easy ones. Averaging across all patients does not compare the treatments — it mostly compares the patient populations they were given.

Within each group, where the comparison is fair, A wins both times. Across the total, where it is not, B appears to win.

The general recipe requires two things together:

A hidden variable that affects the outcome — here, stone size.

That variable distributed unevenly between the things being compared — here, A got the hard cases.

Remove either and the paradox disappears. If both treatments had been used equally on both stone sizes, the totals would agree with the subgroups.

The ingredient that creates the reversalA hidden variable splitsthe dataseverity, age, sourceGroups are very unequal insizethis does the damageThe total hides the pattern
Figure 2.The paradox needs a lurking variable that both affects the outcome and is distributed unevenly between the groups being compared. Without that imbalance, the totals and the subgroups agree.

Where it shows up outside medicine

Hiring and admissions. A famous historical example examined graduate admissions data showing women admitted at a lower overall rate than men. Broken down by department, most departments admitted women at a slightly higher rate. The explanation was that women applied in greater numbers to departments with low admission rates overall. The aggregate and the departments told different stories, and which one answered the question depended on what question you were asking.

Pay comparisons. An organisation can pay women and men equally within every role and still show an overall gap, if the roles themselves differ in pay and people are distributed unevenly across them. This does not make the overall gap uninformative — it means the two numbers describe different things, and conflating them produces bad arguments in both directions.

Product and business metrics. Overall conversion can fall while every individual segment improves, if traffic shifts toward a segment that converts poorly. Teams routinely investigate a "decline" that is entirely composition.

Sports and performance statistics. A player can have a better rate than another in every season and a worse career rate, if the number of attempts per season differs sharply.

The common thread: the total is a weighted average, and the weights matter as much as the values.

Which number should you trust?Does the grouping affect the outcome?Use thesubgroupsUse the combinednumberEither is fineThink carefullyWas the grouping caused by the treatment?
Figure 3.Statistics alone cannot say which view is correct — that depends on why the groups differ. If the grouping happened before and independently of the treatment, split the data. If the treatment caused it, do not.

Which number is actually right

This is the part most explanations skip, and it is the part that matters. Simpson's paradox is not a puzzle about statistics — it is a question about causation, and statistics alone cannot resolve it.

The correct view depends on why the groups differ.

If the grouping happened before, and independently of, the thing you are evaluating, you should usually split the data. In the kidney stone case, stone size was determined by the patient's condition, not by the treatment. Comparing treatments within stone size is the fair comparison; the aggregate is contaminated by which patients each treatment received.

If the thing you are evaluating caused the grouping, splitting the data can remove the very effect you are trying to measure. If a treatment works partly by moving patients into a better category, then controlling for that category hides the benefit.

This is why the same data can support opposite conclusions depending on which variables you adjust for — and why "we controlled for X" is not automatically a mark of rigour. Controlling for the wrong variable introduces bias rather than removing it, and there is no purely statistical test that tells you which is which. You need a view about what causes what, which comes from knowing the subject rather than from the data.

The practical habits:

When you see an aggregate comparison, ask how the groups were composed. Different populations being averaged together is the single most common source of misleading numbers.

When you see a subgroup analysis, ask why those subgroups. Splitting data enough ways will eventually produce a flattering result somewhere.

Be suspicious of any single number describing two different populations. It is not necessarily wrong, but it is answering a narrower question than it appears to.

The deeper lesson is that data does not interpret itself. Two people looking at identical, correctly calculated numbers can reach opposite conclusions, and resolving it requires understanding the situation rather than recomputing the arithmetic — which is a useful corrective to the idea that being data-driven is the same as being right.

Dr Nadeem Khudboddin Shaikh
Dr Nadeem Khudboddin Shaikh
Ex–Wells Fargo · Ex–Goldman Sachs · Columbia University alumnus