Galley A/B · Significance report
EXP · 14 D · N=– · –
01
Success metric
Test duration (days) *
Start date optional
Confidence
Winner selection
Multiple-comparison correction
Guardrail metrics
Secondary metrics that must not regress (e.g. return rate, refunds, page load).
Segment breakdown
Check the result holds within sub-groups (e.g. Mobile vs Desktop). Enter control & variation per segment.
CUPED variance reduction
Removes pre-experiment noise to tighten the confidence interval, and can reveal a real effect a raw test misses. Runs on Control vs the best variation.
Metric to adjust
Pre/post correlation (ρ) · 0.50
Simpson's paradox detector
Checks whether the win reverses when two segments are combined, a classic trap from imbalanced traffic.
Segment 1
Segment 2