In previous posts we established that peeking at data without a plan destroys statistical validity. My simulations showed that checking a dashboard daily over 30 days inflates the false positive rate to 26%. Sequential testing solves this: methods like O’Brien-Fleming and Pocock keep your error rate at about 5% by adjusting significance thresholds at each look.
But my latest work reveals the cost of that safety net: you sacrifice statistical power. The question is how much, and whether the ability to stop early makes up for it.
The chart below summarizes 10,000 simulations comparing O’Brien-Fleming, the conservative standard, against Pocock, the aggressive alternative. Both use 5 planned looks over 30 days.

Scenario 1: the marginal win
At the minimum detectable effect of 0.039, a 3.9% relative lift, the power tradeoff is real.
O’Brien-Fleming achieves 78.5% power, close to the 80% nominal target, and takes 3.74 looks on average, stopping about 25% earlier than the planned 5 looks. Pocock achieves 70.3% power, eight percentage points lower, but is faster at 2.93 looks, stopping about 40% earlier.
Pocock buys you speed but sacrifices power. At marginal effect sizes, you’re accepting a higher risk of missing real wins.
Scenario 2: the home run
At an effect roughly double the MDE, the power tradeoff disappears entirely. Both methods hit 100% power, because the signal is so strong that neither misses it. Pocock maintains its speed advantage at 1.58 looks against 2.37 for OBF, a 33% reduction. For home runs you pay nothing in power for Pocock’s faster boundaries.
What should you do?
Use O’Brien-Fleming to catch marginal wins, when the cost of running the experiment is low but the cost of missing a marginal win is high. It preserves near-nominal power at MDE while still letting you stop 25–50% earlier if results are clearly positive. This is the right choice for most product experiments.
Use Pocock for high-burn or exploratory tests, when you expect large effects are likely, or when speed dominates and you can accept reduced power at smaller effects. If running to full sample costs $20k a day, Pocock’s faster stopping times justify the power loss.
One important caveat: Pocock’s advantage only materializes when there’s a signal to detect. For flat results, all sequential methods typically run to the full 5 looks anyway.
Takeaway
Sequential testing gives you a safety net against peeking, but it isn’t free. O’Brien-Fleming preserves more power while enabling early stopping when effects are large. Pocock sacrifices power at moderate effects in exchange for maximum speed when effects are detectable. Choose OBF when marginal wins matter, and Pocock when speed is the priority and you’re hunting for home runs.
New write-ups when I finish them
I simulate a method, run it against known ground truth, and publish what happened. No schedule, and no filler between posts.
Unsubscribe anytime.