Do the maths first
Take your baseline conversion rate, the minimum effect size that would actually be worth acting on, and your weekly traffic per variant, and run them through a standard sample-size calculation before a single line of test code is written. If the required runtime to reach statistical significance exceeds roughly six weeks for most mid-traffic Shopify stores, the test will be contaminated by seasonality, promotional calendar changes, or a Google algorithm update before it has a chance to conclude honestly.
This calculation is not exotic statistics — it is a five-minute exercise with a free calculator — and yet it is routinely skipped because teams are eager to start testing and treat the arithmetic as a formality rather than a gate. The result is a graveyard of tests marked ‘inconclusive, extended for another two weeks’ that eventually get quietly closed with whichever variant happened to be ahead when someone lost patience.
The uncomfortable implication is that a large share of stores running a CRO programme simply do not have the traffic to test the changes they most want to test — a new checkout flow, a pricing page redesign — at least not in isolation. That is not a reason to abandon experimentation, but it is a reason to be honest about which changes deserve a test and which deserve a different kind of evidence.
The alternative to a bad test
Where traffic genuinely cannot support a conclusive test, ship a considered change based on qualitative evidence — session replay footage, support ticket themes, a structured usability review with real customers — and measure the resulting trend over time rather than pretending to a statistical significance the data cannot actually support.
This is not a lesser methodology, it is a different and often more appropriate one for lower-traffic surfaces. A checkout flow issue identified independently by session replay, support tickets and a usability session is far better evidenced than a single underpowered A/B test that happened to show a positive result by chance.
- Reserve formal A/B testing for high-traffic surfaces and reasonably large expected effects
- Use qualitative evidence — replay, support data, usability sessions — for everything else
- Report inconclusive results as inconclusive, in writing, rather than letting them quietly disappear
The cost of pretending
The real damage from underpowered testing is not the wasted development time on the losing variant, it is the false confidence that accumulates when a team repeatedly acts on results that were never statistically sound. Six months of ‘winning’ tests that were each individually underpowered can leave a site measurably worse off, with nobody able to point to which change caused the decline because each one was ‘validated’ at the time.
This is compounded when tooling makes it easy to declare a winner the moment a dashboard shows green, without checking whether the sample size the tool used to reach that conclusion was ever adequate for the effect size claimed. Most testing platforms will happily report a 95% confidence interval on a dataset that is nowhere near large enough to support one.
Keep the record
Every test, including the ones that failed to reach significance and the ones that were deliberately not run because the maths did not support it, belongs in a written log with the hypothesis, the mechanism believed to be at work, and the actual outcome. That log, more than any single winning variant, is the actual asset a mature CRO programme produces over time.
Teams that maintain this discipline stop re-litigating the same hypothesis every eighteen months when a new hire proposes testing something the team already tried and learned from. That institutional memory is worth more than any individual test result, and it is the part of a CRO engagement that compounds.