Getting started with causal inference and experimentation
My own method tutorials, and four free resources I recommend to people learning this from scratch. All of it is free and I earn nothing from any of it, mine included.
This comes up often enough in coaching that it is worth writing down. People ask what to read to get credible on experimentation and causal inference, usually because an interview exposed a gap, and the honest answer is that the free material is now good enough that paying for an introduction makes little sense.
My own tutorials come first below, then the outside material I point people to. Treat that second list as a set of references rather than a syllabus, because almost nobody works through four resources end to end. Find the one that covers the method you are stuck on and read that part.
Selected causal inference and data science tutorials from my archive
I film these as the questions come up in coaching. They are free on YouTube, there is nothing to buy, and they are here so the whole answer sits in one place rather than scattered across my feed.
Read the transcript
When you are a data scientist, stakeholders will ask you all the time what the impact was of some action the company took, on an outcome you care about such as conversion rate. Most of the time they do not come to you first and set up an experiment. They simply expect you to work it out after the fact.
In a stakeholder’s mind the outcome usually looks simple. It sits flat beforehand, the treatment happens, and it jumps to a higher level. That version is easy to measure, and you could safely say the treatment caused the jump from 60 to 65 percent.
But suppose the conversion rate leading up to the treatment had some variability and was already trending up. After the treatment it is still trending up, with some volatility along the way. How do you know whether the increase is caused by the treatment or by the trend that was already there?
It helps enormously if you can identify a comparison group that was not treated. Say you launched a new checkout page on one site, and you see an upward trend afterwards. You also have a comparison product where no new checkout page was launched, and over the same period it had a similar upward trend. In that case you would probably argue the trend caused the increase, because the treated group is higher than before but so is the untreated group that had been moving with it.
Now suppose it looked different. In the pre-treatment period the treated group is trending up, the untreated group is trending up too, and the two track each other’s movements over time. That tracking is the important part, and it is something you have to test formally. After the treatment window, the treated group keeps increasing and the untreated group does not.
What you can estimate is the change in the gap. Before the new checkout page the gap was about two percentage points. Afterwards it is about five. The difference-in-differences estimate is the post-period gap of five percentage points netted against the pre-period gap of two, so we would say the treatment increased the conversion rate by three percentage points.
Read the transcript
We use synthetic controls to estimate the impact of something that happened when only one unit, or a small number of units, was affected and there was no randomized experiment. For example, your company starts a brand marketing push in one region by putting billboards up across a city.
To demonstrate it I simulated data: 50 markets, two years of monthly observations, and a six-month post period. All the markets are drawn from the same distribution with seasonality, plus a random walk and some noise, so every market differs. The important part is that we know the ground truth, because I did not generate any treatment effect at all. The true effect is zero, and the test is how close to zero the method gets us.
There are many techniques for getting synthetic control weights, but the idea is the same throughout. You form a weighted average of the untreated regions that, taken together, looks like the treated region in the period before the treatment started.
Here I fit a lasso model on the donor markets and use its predictions as the synthetic control. The treatment effects are then just the post-period actuals minus that weighted average, and from those you can read off the average treatment effect over the period, and the relative effect by dividing through by the baseline.
This is where the important part shows up. Everything before the treatment date is in-sample fit, and everything after it is out-of-sample fit. So the pre period looks like an amazing match, while the post period is a combination of out-of-sample fit and whatever treatment effect genuinely exists.
We know from the simulation that there is no real treatment effect, so every gap you see after the treatment date is the difference between in-sample and out-of-sample model fit. In the real world you do not have that luxury. All you have is the gap, and you are trying to work out how much of it is a treatment effect and how much is model noise.
One way to measure that is to drop the treated region and run the identical model on each comparison unit in turn, then look at the distribution of those placebo effects. With a limited number of donors it can be hard to eyeball that distribution and say whether your estimate falls inside or outside a 90 or 95 percent interval. So take the mean and standard deviation of the placebo effects, plot them as a normal distribution, mark the boundaries at plus and minus 1.96, and see where your original estimate lands.
That gap looked tremendous. A perfect fit in the pre period, and then the treated group plummets afterwards, which looks like it has to be a large negative effect. It turns out that a negative effect of that size is fairly typical of the in-sample versus out-of-sample fit across all the untreated units.
The full write-up with the simulation code →
Read the transcript
These techniques have come up in recent interviews across multiple FAANG companies in just the last few weeks, for senior and staff level data science roles. So imagine we want to understand the effect of frame rate on user watch time at Netflix.
The problem is that frame rate will not have a linear effect. There is a world where better quality plateaus at some point and people stop noticing a difference. But as the frame rate gets lower it starts to degrade, until eventually the video is unwatchable.
So I simulated a ground truth effect where above 30 frames per second there is effectively no difference, it is perfect. It then starts to degrade until it hits 24, and below that it gets rapidly worse. Because I simulated the data, we know the ground truth.
I also added a confounder: device quality. People with better devices might get better video quality through higher frame rates, but those people planned on watching more anyway. They already had a higher intent to watch, and they may also be the ones who buy better devices.
The data has users and shows. Shows have different watch times, users have different propensities to watch a show to the end, and shows have a quality level where higher quality means people watch more of them. Each session then gets some random variation in frame rate.
I gather the features that would realistically be available, such as historical watch time for that show and for that user, and estimate a double machine learning model. A random forest predicts watch time, a separate random forest predicts the frame rate, and double machine learning is then a regression of the residuals of the watch rate against the residuals of the frame rate.
Double machine learning does a good job of picking up the general curve, including the knee. The green line is the ground truth and the red dotted line is what the DML model recovers. Linear regression, by comparison, just gives a steady line.
The next question is uplift: what would happen to watch time if we increased frames per second? The important part is that the linear model predicts an increase even for users who already have a high frame rate, while the double ML and the ground truth both show the uplift only happening at the low end.
So if the question is opportunity sizing, say a 10 percent increase, naive OLS dramatically overstates it. An OLS that includes user and show fixed effects still overstates it. The double ML estimate comes out close to the ground truth.
These three are the longer walkthroughs. The shorter ones each come with a notebook that runs in the browser.
All video tutorialsFree resources to learn experimentation and causal inference
-
Course · free
A/B Testing, by Google on Udacity
The gentlest way in, and the right starting point if your only exposure to experiments has been a classroom.
-
Course · free · ~20 hours
Confidence Bootcamp, by Spotify
Where to go once the basics are in place, built by the team who run Spotify’s experimentation platform.
-
Book · free online · R and Stata
Causal Inference: The Mixtape, by Scott Cunningham
Deeper on identification and research design, in the language economists use, and better on why a design fails.
-
Book · free online · Python
Causal Inference for the Brave and True
The same methods in Python, written for people arriving at causal inference from predictive machine learning.
Knowing the methods is not the same as interviewing well
Everything above teaches the methods, but none of it teaches the interview, and those are genuinely different skills. The candidates I see struggle in case study rounds are usually not short on technique. They know how to design the experiment and reach for it too early, before saying what they would measure or who the change might hurt.
If that sounds familiar, the case study series is about that specific gap, and the Playbook covers the full interview loop. Both are paid, while the four resources above are not, and those are the right place to start if the methods themselves are what you are missing.