Skip to content

Reading Results

Every experiment has its own results page in the dashboard. It shows what each variant changed, how each variant performed, and how confident you can be in the difference. This page walks through each section and explains how to act on what you see.

The per-variant results table for an experiment: visitors, click-through, reached PDP, purchase conversion, average order value, and the target metric, with the change and confidence under each figure and the deployed winner highlighted.

Each variant card shows the change that variant makes on your site, along with screenshots of the modified page. Use the screenshots to confirm what visitors in each variant actually saw. This matters when you review results weeks later, or when a teammate who did not approve the experiment reads the report.

The baseline variant represents your site without changes. All other variants are compared against it.

You can also export a winning variant’s code from this view. See Experiment terminal states for how exports fit into rollout.

For each variant, the results page shows how it performed on the experiment’s primary metric and on any guardrail metrics. Conversion-rate metrics show the share of visitors who completed the action. Revenue metrics show purchase conversion rate, average order value, and revenue per visitor. See Metrics and guardrails for what each metric means.

Next to each variant’s result, the dashboard shows a significance level of high, medium, or low. This combines two things: how confident the statistics are that the difference is real, and whether enough visitors have taken part. The label always reflects the more cautious of the two, so a large measured lift on a small sample still shows as low.

Here is how to read each level:

LevelWhat it meansWhat to do
HighThe difference is very likely real and the sample is large enough.Safe to act on. Deploy or end with confidence.
MediumThe result points in a clear direction but is not yet settled.You can act if you need to, but letting it run reduces risk.
LowThe difference could easily be noise.Keep the experiment running. Do not act on the numbers yet.

The dashboard pairs these labels with automated recommendations. Statistics explains how significance is calculated and how recommendations are made.

The results-over-time chart shows how each variant’s performance developed across the experiment’s run. You can bin the data by hour, day, week, or month.

Use this chart to spot patterns:

  • Results that converge and hold steady suggest the measured difference is stable.
  • A gap that keeps flipping direction means the experiment needs more time.
  • A sudden shift on one day may line up with a promotion, a site change, or a traffic spike. Check the activity log for that date.

The segment breakdown splits results by visitor group, so you can see whether a variant works better for some visitors than others. Available segments:

  • New vs returning visitors
  • UTM source, medium, and campaign
  • Landing page

Segments help you understand a result, not overturn it. A variant that wins overall but loses in one small segment is usually still worth deploying. If a segment difference is large and the segment matters to your business, raise it with your Shuttlebase team. They can turn the finding into a follow-up experiment targeted at that group.

The activity log lists everything that happened to the experiment: launch, pauses and resumes, variant changes, approvals, and status changes, each with a timestamp. When a chart shows an unexpected shift, check the log first. A pause, a disabled variant, or a modification often explains it.