What is the winner's curse in A/B testing?
hardAnswer
- Because we only 'ship' experiments that cross the significance / effect threshold, the observed effect on the winners is a biased over-estimate of the true effect (selection on the outcome).
- Empirical rule at Microsoft / Bing: shipped effects shrink 10-40% on re-measurement.
- Fixes: (1) empirical Bayes shrinkage of shipped estimates; (2) look at long-run replicable effects; (3) hold-out samples to re-estimate on unbiased data.
- Foundation of realistic ROI accounting.
Check yourself — multiple choice
- Random
- Selection on significance/effect biases shipped-effect estimate up (10-40% shrinkage on re-measurement); fix via Empirical Bayes shrinkage / holdout re-measurement
- Same as p-hacking
- Not real
Winner's curse: shipped-effect estimates biased up; shrink via EB / holdout.
#ab-testing
Practise Statistics Fundamentals
215 interview questions in this topic.