Find me on

The False Positive Trap: The Dark Side of Bayesian A/B Testing

If you use a history of winning features to boost power, what happens when you test a bad idea? With ten historical winners and no discount, the false positive rate hits 75%.
Read the write-up →

Surviving the “Haircut”: Why Bayesian A/B Testing Beats Conservatism

Stakeholders rarely accept historical lifts at face value. I ran 1,000 simulations to find out how much of a conservative discount a Bayesian prior can absorb before it loses its power advantage.
Read the write-up →

Don’t Trust the Closed-Form: Sizing Switchback Experiments with Permutation Tests

Standard power calculators assume your observations are independent. Time-series data is not, and on a two-week switchback the closed-form standard error understated the true variance by 22%.
Read the write-up →

Slashing Amazon Ad Spend by 69%: Automating a Switchback Experiment to Find the Optimal Bid

Standard A/B testing is impossible in walled gardens. Here is how I built an automated switchback pipeline on my own seller account to bypass platform attribution, cut ad spend by 69%, and increase net profit.
Read the write-up →

Speed vs. Power: The Hidden Tradeoffs of Sequential Testing

Sequential testing fixes the peeking problem, but the safety net costs you power. Here is exactly how much, across 10,000 simulations of O'Brien-Fleming against Pocock.
Read the write-up →

Beyond the Peek: Using Sequential Boundaries to Protect Experiment Integrity

Peeking daily over a 30-day experiment inflates your false positive rate from 5% to 26%. Group sequential boundaries fix it, and choosing between O'Brien-Fleming and Pocock depends on what you are optimizing for.
Read the write-up →