Experimentation
The False Positive Trap: The Dark Side of Bayesian A/B Testing
If you use a history of winning features to boost power, what happens when you test a bad idea? With ten historical winners and no discount, the false positive rate hits 75%.
Read the write-up →
Surviving the “Haircut”: Why Bayesian A/B Testing Beats Conservatism
Stakeholders rarely accept historical lifts at face value. I ran 1,000 simulations to find out how much of a conservative discount a Bayesian prior can absorb before it loses its power advantage.
Read the write-up →
Don’t Trust the Closed-Form: Sizing Switchback Experiments with Permutation Tests
Standard power calculators assume your observations are independent. Time-series data is not, and on a two-week switchback the closed-form standard error understated the true variance by 22%.
Read the write-up →
Slashing Amazon Ad Spend by 69%: Automating a Switchback Experiment to Find the Optimal Bid
Standard A/B testing is impossible in walled gardens. Here is how I built an automated switchback pipeline on my own seller account to bypass platform attribution, cut ad spend by 69%, and increase net profit.
Read the write-up →
Speed vs. Power: The Hidden Tradeoffs of Sequential Testing
Sequential testing fixes the peeking problem, but the safety net costs you power. Here is exactly how much, across 10,000 simulations of O'Brien-Fleming against Pocock.
Read the write-up →
Beyond the Peek: Using Sequential Boundaries to Protect Experiment Integrity
Peeking daily over a 30-day experiment inflates your false positive rate from 5% to 26%. Group sequential boundaries fix it, and choosing between O'Brien-Fleming and Pocock depends on what you are optimizing for.
Read the write-up →