A Statistic Is a Decision, Not a Fact
On this page
Watch & explore
Start with a few high-quality watches, then dive into the notes below.
Try an idea before you read. Test your understanding of how statistics can mislead. Explore →
In 1973, the University of California, Berkeley looked guilty. That autumn it admitted about 44% of its male applicants to graduate study and only about 35% of its women. Roughly 8,400 men and 4,300 women had applied, so the sample was large and the gap was wide. The conclusion seemed to write itself: the university was turning women away.
Then the statistician Peter Bickel and two colleagues did something simple. Instead of treating Berkeley as one number, they looked at it department by department. Almost every department, taken on its own, admitted women at about the same rate as men, and if anything slightly more often. Published in the journal Science in 1975, their study found no sign that departments were rejecting women for being women. The overall figure and the department figures told opposite stories, and both were correct.
How can a true number lie? Because a statistic is never handed to you the way an apple is. Someone chose what to count, how to group it, and how to draw it. Change any one of those choices and the very same facts can point the other way. The useful skill here is not memorising formulas. It is learning to ask one quiet question of every figure you meet: what decision produced this number?
The average that hides the crowd
The Berkeley puzzle has a name: Simpson's paradox. It is what happens when a pattern that holds inside every group vanishes, or even reverses, once you pour the groups together. At Berkeley the secret was in the choice of subject. Women were applying in larger numbers to the most competitive departments, the ones that turned away most applicants of either sex. Men were applying more often to departments that admitted almost everyone. Each department was roughly even. The overall figure was dragged down not by bias in the rooms where decisions were made, but by where the two groups were standing in the queue.
This is not a rare laboratory curiosity. It surfaces whenever an average is asked to speak for a crowd made of very different parts. A hospital with worse survival rates may simply be the one that treats the sickest patients. A city where average wages appear to fall may have hired thousands of junior workers while everyone already employed got a raise. The honest first move, every time, is to split the number back into the groups it was built from and see whether the story survives.
When the picture does the lying
Numbers can mislead in words. Charts can mislead faster, because the eye reads a shape before it reads a label. In 2014 the news agency Reuters published a graph of gun deaths in Florida before and after the state's 2005 Stand Your Ground law. The line climbed gently for years, then appeared to plunge after 2005. To almost every reader it said one thing: deaths collapsed.
They had not. The designer had flipped the vertical axis upside down, so that zero sat at the top and the counts grew downward. The plunging line was in fact a sharp climb. Firearm deaths had gone up after the law, not down. Nothing in the chart was factually false. The data were real, and the labels were there for anyone who checked. It lied with its shape while telling the truth in its ink.
The same quiet trick hides in the most ordinary charts you will see this week. An axis that starts at 90 rather than zero can turn a wobble into a cliff. A timeline that begins in a conveniently chosen year can bury the part that would spoil the point. Before you believe a graph, look at what the axes are doing, not just where the line goes.
Exact is not the same as complete
In 1973 a statistician named Francis Anscombe built four small sets of numbers and published them to make a single stubborn point. All four had almost exactly the same average, the same spread, and the same trend line. On paper they were near twins. Then he drew them. One looked like a rough straight climb. One bent into a curve. One ran perfectly straight until a single outlier yanked the line off course. One was a vertical stack of points with one distant straggler pulling everything sideways. They looked nothing like each other.
Anscombe's four datasets, still taught today, carry a warning worth keeping: a summary is exact, but it is not complete. An average, a percentage, a lone correlation squeezes a mountain of detail into one figure, and some of what it throws away is the part that mattered most. This is also why subjects that feel abstract on a page come alive the moment you attach them to real data, which is the whole idea behind the way every subject is mapped to the real world in the Learnacy Hub. A number is a doorway, not the room.
The oldest trick, and the questions that beat it
None of this is new. In 1954 a journalist called Darrell Huff, who was not a statistician at all, wrote a slim, cheerful book called How to Lie with Statistics. It went on to sell more than a million and a half copies in English and became one of the best-selling statistics books ever written. The reason it never dates is bleak and useful: the tricks do not change. Cropped axes, cherry-picked years, averages that hide the crowd, pictures that flatter the person who drew them. Huff simply named them so that readers could see them coming.
You do not need a degree to do the same. You need a short list of questions, asked out of habit rather than suspicion. Who counted this, and how? Compared with what? What is the base, the total this percentage is a slice of? What has been left out? Does the axis start at zero? And, quietly, who benefits if I believe this? These are not exam techniques. They are a way of standing in the world with your eyes open, and they are worth practising on the numbers you actually meet: in the news, in an argument, in a chart a friend forwards you. We keep a growing shelf of pieces like this in the resources library, and it is exactly the habit our mentors try to build through Learnacy: not answers to memorise, but questions that keep you honest.
Numbers do not lie. People choose frames, and a frame is an argument wearing the costume of a fact. The most numerate person in any room is rarely the fastest with arithmetic. It is the one who, before nodding along, still asks what decision produced this number. Learn to ask it, and no chart will ever quite own you again.
Sources
- Simpson's paradox and the 1973 UC Berkeley admissions figures (44% of men, 35% of women; Bickel, Hammel and O'Connell, Science, 1975)
- Anscombe's quartet: four datasets, near-identical statistics, different shapes (Francis Anscombe, 1973)
- The 2014 Reuters Florida gun-deaths chart with the inverted vertical axis
- How to Lie with Statistics, Darrell Huff, 1954, and its sales
- TED-Ed lesson: How statistics can be misleading (Mark Liddell)
- TED-Ed video on YouTube: How statistics can be misleading
Key takeaways
- Statistics are not objective facts, but rather the result of human choices regarding what to measure, how to group information, and how to display it.
- Combining different groups of data can hide or even reverse trends that exist within the individual groups, a phenomenon known as Simpson's paradox.
- To uncover the true story behind a broad average, it is often necessary to break the numbers back down into their original, smaller categories.
- Visual representations like charts can easily mislead viewers through design choices, such as flipping an axis upside down or starting it at a number other than zero.
- Statistical summaries, such as averages or trend lines, can be mathematically precise while still hiding crucial differences and outliers in the underlying data.
Test yourself
What is the name of the statistical effect where a pattern in separate groups disappears when the groups are combined?
Simpson's paradox.
How did a 2014 Reuters graph make an increase in Florida gun deaths appear to be a decrease?
The designer flipped the vertical axis upside down so that the numbers grew downward.
What did Francis Anscombe create in 1973 to show that statistical summaries can hide important details?
He built four datasets that shared the same average and trend line but looked completely different when drawn on a graph.
Try it
A Statistic Is a Decision, Not a Fact
Test your understanding of how statistics can mislead.
1The University of California, Berkeley's overall admission rates showed a large gap between men (44%) and women (35%). But when examined department by department, almost every department admitted women at roughly the same rate as men. What explained this discrepancy?
The study found no sign that departments were rejecting women for being women. The text explicitly states this.
This is the key insight from the text. The gap was caused not by bias in decision-making, but by where the two groups were standing in the queue—different application patterns, not different treatment.
The text states the sample was large: roughly 8,400 men and 4,300 women had applied. The sample size is precisely why the gap looked so significant.
2In 2014, Reuters published a graph of gun deaths in Florida before and after the 2005 Stand Your Ground law. The line appeared to plunge after 2005, suggesting deaths had collapsed. What was the trick?
The text explicitly states the data were real and the labels were there for anyone who checked. Nothing in the chart was factually false.
This is exactly what the text describes. The plunging line was actually a sharp climb—deaths went up after the law, not down. The chart lied with its shape while telling the truth in its ink.
While the text does mention that timelines can be conveniently truncated, this is not what the Reuters example involved. The specific trick here was the inverted axis.
