Correlation is not causation. So what is?
Everyone learns the slogan. Almost no one is taught what to do about it. Causal Blocks is a small project to change that: draw what you believe causes what, and see what those beliefs do to the answer.
The question behind most questions
Most questions people bring to data are really about cause. Will this treatment help? Would a smaller class raise grades? Does the new pricing bring in customers, or were they coming anyway? Each one asks what would happen if we changed something.
A correlation answers a different question: when one thing is higher, is the other usually higher too? That can be true without either causing the other. Treating one question as the other is the most common mistake in data analysis, and it is easy to make without noticing.
Three ways a correlation misleads
- A common cause. People who carry lighters get lung cancer more often. Lighters don't cause cancer; smoking causes both. The shared cause is called a confounder.
- The arrow points the other way. Areas with more police have more crime. That is mostly because police are sent where crime is, not because police create it.
- Only looking at some cases. Among hospital patients, two unrelated illnesses can look linked, because having either one is a reason to be in hospital. Selecting on a shared result creates a pattern. That shared result is a collider.
"Just control for it" is not the fix
The usual response is to add more variables to the regression. Sometimes that helps. Sometimes it quietly makes things worse. Adding a variable has three possible effects, and the data alone cannot tell you which one you are getting:
| If the variable is a… | Adjusting for it… |
|---|---|
| confounder | removes bias. You need it. |
| mediator | changes the question you answered, from the total effect to the direct one. |
| collider | creates bias that was not there before. |
Which role a variable plays depends on how everything is connected. So the connections have to be written down first.
Draw your assumptions
The tool for writing them down is a causal graph: a set of boxes and arrows, where each arrow says "this directly affects that". Once the graph is drawn, a rule called the backdoor criterion tells you exactly which variables to adjust for and which to leave alone. Libraries such as DoWhy apply that rule and estimate the effect.
The graph is an assumption, not a fact. That is its value. It puts your reasoning where someone else can see it and argue with it, instead of leaving it hidden inside the choice of control variables.
What Causal Blocks is for
Causal inference is decades old, standard in epidemiology and economics, and still missing from many statistics, analytics and data science courses. Causal Blocks exists to make it easier to pick up, especially for students:
- An interactive demo where you draw a graph over real data and watch DoWhy's estimate change with every arrow. Each variable is coloured by the role it plays, so the graph explains itself.
- Worked examples that show the mistakes happening, on real data and on simulated data where the true answer is known.
- Open-source code, so you can check every number.
The first example uses data on education in English towns, because it shows all of this in one dataset: the published correlation, the result you get by controlling for everything, and a different story once the causes are drawn. Try the demo. The subject is only the example. The point is the habit of asking what causes what before fitting anything.
Read more
- The Book of Why Judea Pearl and Dana Mackenzie, 2018 The best place to start. A non-technical history and explanation of causal thinking from the person who formalised much of it.
- Thinking Clearly About Correlations and Causation Julia Rohrer, 2018, Advances in Methods and Practices in Psychological Science A short, readable paper on causal graphs for observational data. Good for a single sitting.
- The Effect Nick Huntington-Klein, 2021 Free online A friendly textbook on research design and causal inference, with causal graphs throughout.
- Causal Inference: The Mixtape Scott Cunningham, 2021 Free online An economist's introduction, with code in R, Stata and Python.
- Causal Inference: What If Miguel Hernán and James Robins, 2020 Free online The standard text from epidemiology. More formal, and very clear.
- Statistical Rethinking Richard McElreath, 2nd edition, 2020 A statistics course that puts causal graphs at its centre. The recorded lectures are free.
- Introduction to Causal Inference Brady Neal Free course Lecture videos and notes, from the machine-learning side of the field.
- DAGitty Johannes Textor and colleagues Free tool Draw a causal graph in the browser and get the adjustment sets. The academic standard.
- Spurious Correlations Tyler Vigen Hundreds of real correlations with no causal link at all. A good reminder, and funny.