Skip to content

Revenue Uplift

The revenue uplift model answers one question: how much extra revenue did a specific change earn so far, and how much will it keep earning over the next year?

This page explains how that number is built, what makes it deliberately cautious, and where it stops counting. The figures it produces are the ones on your account performance page and on each deployed experiment.

Uplift is measured per deployed winner and then summed across your account. Each winner contributes two parts:

PartPeriodHow it is produced
Earned during the testFrom experiment start to deployMeasured. The winning variant’s actual revenue minus the control’s, adjusted for the traffic each side received.
Earned since deploy, and projected forwardFrom deploy onwardEstimated. The per-visitor revenue difference the test proved, applied to your traffic and faded out over 18 months.

The first part is money that already landed and can be counted directly. The second part is a projection, so it is built on the conservative end of the measurement rather than the best case.

While an experiment is live, only part of your traffic sees the winning variant, so the extra revenue is real but limited to that share.

Shuttlebase takes the winning variant’s total revenue and subtracts the control’s revenue, scaled to the same number of visitors. If the winner served 40,000 visitors and the control 38,000, the control’s revenue is scaled up by 40,000 / 38,000 before the subtraction, so an uneven split does not distort the comparison.

No statistical discount is applied here. This is revenue that was actually recorded, not a forecast.

Once the winner is live for everyone, there is no control group left to compare against, so this part has to be modeled. It starts from the difference in revenue per visitor between the winning variant and the control.

CI-90: projecting the low end, not the average

Section titled “CI-90: projecting the low end, not the average”

The measured average difference is the single most likely value, but the true difference could land above or below it. Projecting the average means being wrong high about half the time.

Instead Shuttlebase uses a method called CI-90: rather than projecting the average difference, it projects the low end of the range it is 90% confident the true difference sits above.

The size of that discount depends on how clean the result is. A large, consistent win on plenty of traffic sits close to its average. A small win measured on thin traffic, or on revenue data with a few very large orders stretching the spread, gets discounted much harder — sometimes all the way to zero.

This is why the uplift figure is usually smaller than the headline lift on the experiment’s results page. The results page reports what was measured. The uplift model reports what your site can be relied on to keep delivering.

Shuttlebase does not book a permanent gain. Traffic changes, customers change, and the site around the change keeps moving, so a win from two years ago cannot still claim its original per-visitor effect.

So the per-visitor uplift declines in a straight line to zero over 18 months from the deploy date. On the day of deploy it counts in full; nine months later it counts at half; after 18 months it contributes nothing further. Uplift for any period is the accumulated area under that declining line, which is why a recent win adds to the total faster than an older one.

This applies both to the revenue accrued since deploy and to the next 12 months of projection, so a winner deployed 15 months ago has only three months of fading value left to contribute.

The per-visitor figure is multiplied by your expected visitors for the period. That expectation comes from the experiment’s own observed traffic, annualized using your account’s seasonality so a test that ran through a peak season is not projected as if every month looks like that one.

The visitor estimate is frozen at the moment of deploy. Later shifts in your seasonality data do not retroactively rewrite what a past win is credited with.

An experiment on a product page runs to a clear win:

  1. Revenue per visitor: $4.20 for the winner, $4.00 for the control. The measured difference is $0.20.
  2. Accounting for how much that difference could vary, the CI-90 lower bound comes out at $0.12 per visitor. That is the figure the projection uses, not $0.20.
  3. Expected annual visitors for the experiment: 2,000,000. Using the CI-90 value we calculated, the expected uplift is at least $0.12 × 2,000,000 = $240,000 per year.
  4. The fade applies. Over the 12 months following deploy, the per-visitor value declines from full strength toward zero on the 18-month schedule, so the first year contributes roughly $160,000 rather than the full $240,000.
  5. Revenue measured during the test itself is added on top, and the result is the experiment’s contribution to your account total.

The same model is applied to purchases too, not just revenue

Section titled “The same model is applied to purchases too, not just revenue”

The same model runs a second time on purchase conversion rate to produce the extra purchase count attributable to each win. It uses the same CI-90 lower bound, the same 18-month fade, and the same 100-purchase minimum, so the purchase and revenue figures always tell a consistent story.

  • It does not include experiments that were ended without deploying a winner, however much was learned from them.
  • It is not the number your analytics tool will show. Analytics reports total revenue; uplift is the difference against a control that no longer exists after deploy. See Google Analytics 4 for comparing the two.
  • Locked-in revenue on the performance page is a projection from wins already deployed, not a guarantee. Traffic that falls far below expectation reduces what actually arrives.