Causal Blocks

Correlation is not causation. So what is?

Everyone learns the slogan. Almost no one is taught what to do about it. Causal Blocks is a small project to change that: draw what you believe causes what, and see what those beliefs do to the answer.

The question behind most questions

Most questions people bring to data are really about cause. Will this treatment help? Would a smaller class raise grades? Does the new pricing bring in customers, or were they coming anyway? Each one asks what would happen if we changed something.

A correlation answers a different question: when one thing is higher, is the other usually higher too? That can be true without either causing the other. Treating one question as the other is the most common mistake in data analysis, and it is easy to make without noticing.

Three ways a correlation misleads

  1. A common cause. People who carry lighters get lung cancer more often. Lighters don't cause cancer; smoking causes both. The shared cause is called a confounder.
  2. The arrow points the other way. Areas with more police have more crime. That is mostly because police are sent where crime is, not because police create it.
  3. Only looking at some cases. Among hospital patients, two unrelated illnesses can look linked, because having either one is a reason to be in hospital. Selecting on a shared result creates a pattern. That shared result is a collider.

"Just control for it" is not the fix

The usual response is to add more variables to the regression. Sometimes that helps. Sometimes it quietly makes things worse. Adding a variable has three possible effects, and the data alone cannot tell you which one you are getting:

If the variable is a…Adjusting for it…
confounderremoves bias. You need it.
mediatorchanges the question you answered, from the total effect to the direct one.
collidercreates bias that was not there before.

Which role a variable plays depends on how everything is connected. So the connections have to be written down first.

Draw your assumptions

The tool for writing them down is a causal graph: a set of boxes and arrows, where each arrow says "this directly affects that". Once the graph is drawn, a rule called the backdoor criterion tells you exactly which variables to adjust for and which to leave alone. Libraries such as DoWhy apply that rule and estimate the effect.

The graph is an assumption, not a fact. That is its value. It puts your reasoning where someone else can see it and argue with it, instead of leaving it hidden inside the choice of control variables.

What Causal Blocks is for

Causal inference is decades old, standard in epidemiology and economics, and still missing from many statistics, analytics and data science courses. Causal Blocks exists to make it easier to pick up, especially for students:

The first example uses data on education in English towns, because it shows all of this in one dataset: the published correlation, the result you get by controlling for everything, and a different story once the causes are drawn. Try the demo. The subject is only the example. The point is the habit of asking what causes what before fitting anything.

Read more