Skip to content
Insights

June 24, 2026 · Trial notes · 5 min read

The metric you optimize decides which strategies survive

Two emails, nearly identical on the engagement dashboard. One built +$5.20 of forward value per customer; the other built nothing. If conversion rate had made the call, the wrong one might have lived.

Every metric is a filter. Strategies that score well on it survive and get budget; strategies that score badly get killed. This is so obvious it barely seems worth writing down, until you watch two strategies score nearly the same on the metric everyone uses, while an honest measurement shows one working and one inert.

We watched exactly that in a completed eight-week randomized controlled trial with a DTC apparel brand. Client details and figures are altered to preserve confidentiality; the design and the direction of every result are exactly as run.

The two treatments

Both arms received one extra email on top of the brand's normal marketing, sent to high-value one-time buyers selected in advance by forward CLV. One arm's email carried a modest second-purchase offer. The other arm's email was the same touch without the offer.

On the conversion dashboard, the difference between them was small: a couple of points, the kind of gap that gets explained away by seasonality or creative in an ordinary reporting meeting. Neither looked like a failure. Neither looked like a breakout. If you have ever stared at two rows of a campaign report and shrugged, you know the feeling.

Then we read the same two arms on the trial's actual endpoint:

SECOND-PURCHASE CONVERSION NET FORWARD VALUE / CUSTOMER 8.1% 7.6% OFFER REMINDER +$5.20 $0.00 OFFER REMINDER NEARLY TIED NOT CLOSE
The same two emails under two metrics. Conversion could not tell them apart. Forward value told them apart in every read. Figures altered for confidentiality.

The forward-value read

We measure differently. Every customer carries a 12-month forward CLV score: expected purchases, expected spend, expected returns, expected margin. The trial's primary endpoint was the change in that score, per customer, versus control, with a 95% confidence interval, read weekly on dates registered before launch.

On that read, the two arms were not close. The offer arm built +$5.20 of net forward value per customer, and its interval, +$1.30 to +$9.10, pulled clear of zero and stayed clear through the final read. The no-offer arm was statistically zero in all eight reads.

The customers the offer converted were not just converting; they were converting into buyers worth meaningfully more over the following year. The gap invisible on the conversion report was the entire result.

The uncomfortable general case

Here is the uncomfortable version for any operator: your reporting stack is a survival filter for strategies, and if the filter measures the wrong thing, it selects for the wrong survivors. Conversion-optimized reporting selects for strategies that generate cheap conversions, including conversions that would have happened anyway and conversions of customers who never pay back their acquisition cost.

Run that filter for three years and the compounding is brutal. Every quarter, a few strategies get scaled and a few get killed, each call made on a metric that cannot see forward value. Some of the scaled ones are inert; some of the killed ones were building value the report could not show. No single decision looks wrong. The calendar just fills, season by season, with activity that photographs well and compounds nothing.

The dashboard is not lying to you. It is answering a different question than the one your P&L is asking.

The strongest objection

"Conversion is at least real. It counted actual orders. Forward value is a model's opinion."

We take this one seriously, because it is half right. Forward value is a model output, which is why ours is validated on a held-out slice of the client's own history before it directs anything, and why the trial that tested it reported its interval instead of a confident-sounding point.

But the other half is the part operators miss: conversion embeds a model too. It assumes every counted order was caused by the campaign, which the trial directly disproved; the reminder's conversions were overwhelmingly orders that would have happened anyway. That assumption is a model of customer behavior, an extremely naive one, and it never publishes its error bars. The choice is not "measurement versus model." It is a model that states its assumptions and scores itself against reality, versus one that hides inside a report and never gets graded.

The verdict

If two strategies are indistinguishable on your current metric and differ by the whole result on forward value, your current metric is deciding part of your budget by coin flip. The fix is not a smarter dashboard. It is a metric that carries the future in it, and a measurement discipline honest enough to say zero.

  • Add a forward-value endpoint to your next A/B test alongside conversion, and see whether the two agree on the winner.
  • Require an interval on any number that moves budget. A point without a range is not a decision input.
  • Audit last quarter's kill list. Anything killed on conversion alone may deserve a retrial on value.

One brand's result, presented as a range, not a law of nature. The way to know what your own filter is selecting for is to run the measurement on your own customers, which is what the free diagnostic does.

Get the next verdict when it ships.

New insights, when there’s one worth sending. Email only, unsubscribe anytime.


Written by the team that builds and runs the model. Nothing here ships without a method behind it.

Run your numbers.

Six numbers you already know, on your own Shopify and Klaviyo data, no obligation. We score the forward value of every customer and walk you through where it concentrates, and where it leaks.

Read-only access · No deck required · You keep the analysis either way