✎ Blog
Notes from building an experiment engine.
What we have learned about assignment, statistics, growth teams, and the stubborn work between an experiment idea and a clean result. Product claims are tied back to the system we are building; comparisons link to the vendors' own documentation.
P-values for product teams: a working interpretation
A p-value measures how surprising your A/B test result would be if the change did nothing. A working interpretation, common misreadings, and a better lens.
Minimum detectable effect: the number that decides your test before it starts
Minimum detectable effect is the smallest lift your A/B test can reliably see. How to choose an MDE from business value, and why 2% hunts on low traffic fail.
How long should an A/B test run?
Most A/B tests should run one to four weeks: full business cycles, and enough sample for the effect size. How to work out your number, and when to stop early.
Guardrail metrics: how to win without breaking something else
Guardrail metrics protect revenue, retention, and performance while your A/B test chases a win. How to pick them per funnel stage and act on a breach.
11 A/B testing mistakes that quietly ruin your results
The most common A/B testing mistakes — peeking, ignored sample ratio mismatch, testing trivia, skipped cleanup — and the concrete fix for each one.
A/B test sample size: how to size a test without a statistics degree
A/B test sample size comes down to three inputs: baseline rate, minimum detectable effect, and power. The intuition for each, and what to do when n is huge.