Skip to content

What a delta actually costs

What a delta actually costs

A two percent lift is not a number until you know how wide the interval around it is. Here is how we report both, and why the second one changes decisions more often.

Most teams ship an experiment, read a single number off a dashboard, and move on. The number is usually a lift: conversion went from 4.1% to 4.2%, so the change worked. That reading is not wrong so much as unfinished. It answers "which direction" and stays silent on "how sure", and the second question is the one that decides whether you keep the change.

The interval is the finding

A delta of two percent with a confidence interval spanning minus one to plus five is a different result from the same two percent spanning one and a half to two and a half. The first says you learned almost nothing and spent two weeks doing it. The second says you learned something worth rolling out. Both render as "+2%" on a tile.

We report the interval next to the point estimate everywhere, in the same weight, because a number whose uncertainty is a click away is a number people will quote without its uncertainty.

Sample size is a budget, not a gate

The usual framing treats sample size as a threshold you wait to cross. It is closer to a budget you allocate. Every experiment you run is traffic you are not spending on another one, and a test powered to detect a five percent effect will not see a one percent effect no matter how long you leave it.

The practical version of this is unglamorous. Decide the smallest effect worth acting on before you start. If you would not ship a one percent lift, do not power the test to find one.

What we do differently

Nothing exotic. We compute the interval by default, we show it at the same size as the estimate, and we refuse to render a result as a single number in any surface where a decision gets made. That last rule is the one that took the most argument internally and has been worth the most since.

Written by

Julian Vandermolen

Last updated