A report lands on your desk: the campaign generated +$2.10 per customer. Do you scale it?
You cannot answer that question. Not because the number is wrong, but because it is alone. +$2.10 with a plausible range of +$1.80 to +$2.40 is a green light. +$2.10 with a plausible range of -$3.00 to +$7.20 is a coin flip wearing a suit. Same point estimate, opposite decisions, and the report that omits the range has quietly made the decision for you.
This note explains, in operator terms, what a confidence interval is, what it looked like in a controlled trial we ran, and why we treat a number without one as unfinished work.
The rerun-the-world question
Every marketing result you have ever seen is one draw from a noisy process. Run the identical campaign on the identical audience in a parallel week and the number lands somewhere else: different customers happen to be in-market, different orders happen to be large. The question that matters for a budget decision is never "what did the number do?" It is "if we reran this, where would the number land?"
A 95% confidence interval is the honest answer to that question: the range where the true effect plausibly lives, given the noise in your data. Wide interval, you know little. Narrow interval, you know a lot. Interval clear of zero, the effect is real enough to act on. Interval spanning zero, you cannot yet distinguish your campaign from nothing, no matter how good the point estimate looks.
What it looked like in practice
In an eight-week randomized controlled trial with a DTC apparel brand, the primary endpoint was net 12-month forward value built per customer versus control, read weekly on dates registered before launch. Watch what the offer arm's reads did, and notice how differently the point and the interval behave:
| Read | Point estimate | 95% CI | What it licensed |
|---|---|---|---|
| Week 1 | +$3.40 | -$0.60 to +$7.40 | Nothing yet. Spans zero. |
| Week 4 | +$2.40 | -$1.30 to +$6.10 | Still nothing. Hold. |
| Week 8 | +$5.20 | +$1.30 to +$9.10 | Clear of zero. Scale. |
The point estimate wandered between +$2.40 and +$5.20 across the eight weeks. A team reading unaccompanied points would have celebrated week 2, panicked at week 4, and made at least one wrong call. The interval told a steadier story: not yet, not yet, now. And when the trial's other arm, a no-offer reminder email, was read the same way, its interval spanned zero in all eight reads. Point estimates alone would have credited it; the interval retired it.
There is one more discipline hiding in that table: the read dates were registered before launch. An interval you compute on the flattering week is theater. Committing to the calendar in advance is what keeps the measurement from chasing its best look.
Where honest intervals come from
Ours are bootstrapped: the trial population is resampled at the customer level thousands of times, and the effect is recomputed on every resample. The spread of those recomputations is the interval. No distributional hand-waving, no formula chosen for convenience; the uncertainty you see is the uncertainty that was actually in the customers. This matters in DTC specifically because order values are wildly skewed, and a handful of large baskets can impersonate a trend. Customer-level resampling prices that in.
The same machinery is what makes a null result trustworthy. "The interval spanned zero for eight consecutive weeks" is a strong, useful, money-saving finding. A measurement system that cannot produce a confident zero cannot produce a confident anything.
The strongest objection
"Leadership wants one number. Ranges read as hedging."
We would put it exactly backwards: the range is the confidence. Anyone can print a point; the interval is the part that had to be earned. And if a single number is genuinely required for planning, the right one is the lower bound. "This program builds at least +$1.30 per customer at 95% confidence" is a stronger sentence than any unaccompanied average, because it still holds on the bad weeks. Plans built on lower bounds get beaten pleasantly. Plans built on points get revised.
The verdict
A marketing number without an interval is not a measurement; it is an opinion with units. The interval is what converts a report into a decision: clear of zero, scale it; spanning zero, stop paying for it; no interval at all, the analysis is not done.
- Take your biggest recent "win" and ask for its range. If nobody can produce one, you have a point, not a result.
- Plan against lower bounds, not averages. The lower bound is the only number that holds on bad weeks.
- Fix read dates before launch, and treat a confident zero as a deliverable, not a failure.
Every result we publish from an experiment carries its interval, computed on your own customers. If a measured number can't carry its interval, we don't ship it.