How Long to Run an A/B Test on Shopify: Our 1,000-Order Rule
[ SUMMARIZE WITH AI ]
[ FREE CRO TEARDOWN ]
Find the 3 biggest revenue leaks on your store.
Every day a conversion leak goes unfixed, you're paying for traffic that doesn't buy. Get a 5-minute Loom through your PDP, cart, and checkout, with mockups of the fixes. No pitch.
Get My TeardownHow long to run an A/B test on Shopify: until every variation has at least 1,000 orders and at least two full weeks have passed, whichever comes last. Days are the wrong unit. Orders are the right one. A store doing 300 orders a day hits the bar in a week of traffic and still waits for the second weekend. A store doing 30 orders a day needs over two months for a simple control vs variation test, and should probably test something bigger instead.
We didn't always work this way. We used to call tests at 50 to 150 orders per variation when the numbers looked clear. Below are three of our own tests where the early read and the final read told different stories, the math behind the 1,000-order bar, and the rules we run every test by now.
Who this is for: Shopify brands running A/B tests in Intelligems, Shoplift, or a similar tool.
What you'll learn: how long to run an A/B test, why early results lie, and what to do if your store is too small for the bar.
The data: three real tests from our client work, final numbers pulled from Intelligems. Clients anonymized.
Why orders, not days, decide how long to run an A/B test
Most advice on how long to run an A/B test gives you a number of days: "two weeks", "30 days". That ignores the only thing that actually matters, which is how many people bought in each variation.
Revenue per visitor, the metric we judge tests on, is driven by orders. With 50 orders in a variation, two customers buying a big bundle can swing the result by 10% or more. With 1,000 orders, they barely move it.
Days still matter for one reason: shopping behavior changes across the week. Weekend shoppers are not weekday shoppers, and payday is not mid-month. So we use both rules. Hit the order count, and cover at least two full weeks.
Three of our tests where the early read was wrong
These are real tests from the last few months. The numbers are the probability to beat control that Intelligems reports, and the change in revenue per visitor (RPV) against control.
| Test | Early read | Final read |
|---|---|---|
| Desktop navigation (control + 1 variation) | -30% RPV, 9.7% chance to beat control, 43 vs 54 orders. We called it a loss. | Never ran longer. Under today's rule, this was too thin to call either way. |
| Collection page navigation (control + 3 variations) | Day 13: +36.6% RPV, 97% chance to beat control, 138 vs 123 orders. We recommended a deploy. | 153 vs 151 orders: +15.2% RPV, 82% chance to beat control. |
| Product page buy box (control + 4 variations) | Day 10: two variations ahead by +3.3% and +3.1% RPV at ~70%, ~1,200 orders each. One segment showed +17.9%. | Final read, 2,100+ orders each: all four variations behind control, -1.1% to -3.3% RPV. |
Test 1: the loss we called at 54 orders
A desktop navigation test with product cards in the dropdown. After 12 days the variation was down 30% on revenue per visitor with a 9.7% chance of beating control. It looked like a clear loss, so we stopped it.
The problem: 43 orders against 54. At that size, only a change of roughly 50% or more shows up as real. A 30% drop is well inside the noise. The honest answer was "we don't know yet". Desktop traffic on that store was too thin to test navigation at all.
Test 2: the +36.6% winner that shrank to +15.2%
A collection page test for a home decor brand. On day 13 the best variation was up 36.6% on revenue per visitor at 97% probability to beat control. The trend had been stable from day one. We recommended a deploy, and the client shipped it.
About 30 more orders per variation later, the final read was +15.2% at 82%. Still ahead, but less than half the lift, and no longer a confident result. The variation is live and we're watching it post-deploy. Our mistake was the promise of +36.6%, not the deploy itself.
Test 3: 1,200 orders each, and still wrong
A product page buy box test for a supplement brand, control + 4 variations. By day 10, every variation already had about 1,200 orders. Two variations were ahead by just over 3%, and among visitors who clicked a variant option one variation was up 17.9%.
By the end, with over 2,100 orders per variation, all four were behind control by 1.1% to 3.3%. A wash. This test already passed the order bar at day 10. What it hadn't done was finish its second week. That's why the rule has two parts.
It also shows the segment trap. A +17.9% lift inside one slice of visitors is a clue for the next test, not a result. We decide on the whole-test number.
Why 1,000 orders per variation
There's simple math behind the bar. At 95% confidence and 80% power, the smallest change in conversion rate you can reliably detect is roughly 2.8 × √(2 ÷ orders per variation).
| Orders per variation | Smallest change you can trust |
|---|---|
| 54 | ~54% |
| 150 | ~32% |
| 500 | ~18% |
| 1,000 | ~12.5% |
| 2,000 | ~9% |
A 10% lift in revenue per visitor is a strong result on a store that has already been optimized. Below 1,000 orders, wins that size are invisible, and the "big" results you see are mostly noise. That's exactly what happened in tests 1 and 2.
And this is for conversion rate. Revenue per visitor is noisier, because order values vary, so it needs even more. 1,000 orders is a floor, not a finish line.
The rules we run every test by now
- 1,000 orders per variation, minimum. Control counts too. Control + 3 variations means 4,000 orders through the tested page.
- At least two full weeks. Start and stop on the same weekday so every day of the week is covered equally.
- Set the stop date before launch. Estimate it from your daily orders on the tested page. Checking daily is fine. Stopping because the number looks good today is not.
- Judge on revenue per visitor for the whole test. Segments (device, new vs returning, engaged visitors) explain a result. They don't make one.
- "Inconclusive" is an answer. It means the change doesn't matter much either way. Ship the cheaper version and test something bolder.
For how we pick what to test in the first place, see our findings from our winning Shopify tests and the 7 reasons A/B tests fail.
What if your store can't reach 1,000 orders?
Do the math first. Take the daily orders that pass through the page you want to test, and multiply the order bar by the number of versions (control included).
- Control + 1 variation needs 2,000 orders. In four weeks that's ~70 orders a day through the tested page.
- Control + 3 variations needs 4,000. In four weeks that's ~140 a day.
If you're below that, you have three options:
- Test fewer variations. Every extra variation adds another 1,000 orders to the bill. See A/B vs multivariate testing for the traffic cut-offs.
- Test bigger changes. A pricing, offer, or bundle change can move RPV by 20%+, which shows up with fewer orders. A button color never will.
- Test where the orders are. Sitewide elements like the cart drawer or header see every order. A single landing page sees a fraction.
If none of that works, don't test. Make the change based on research and watch the before and after. A test that can't reach a real answer costs you development time and gives you false confidence.
Common questions about A/B test length
How long should you run an A/B test on Shopify?
Until every variation, including control, has at least 1,000 orders and at least two full weeks have passed. Whichever takes longer decides. For most stores that's two to six weeks.
How many orders do you need per variation?
We use 1,000 as the minimum. At that size you can reliably spot a change of about 12.5% in conversion rate. Smaller lifts need more orders, and revenue per visitor needs more than conversion rate.
Can I stop a test early if it hits 95% probability?
Not before it reaches the order bar and two full weeks. Early probabilities swing a lot. One of our tests read 97% at 138 orders and fell to 82% about 30 orders later.
What if my store is too small for 1,000 orders per variation?
Test fewer variations, test bigger changes like pricing or offers, or test sitewide elements that see every order. If you still can't reach the bar in about six weeks, skip the test and make research-backed changes instead.
Want tests you can trust?
Our team runs testing programs for 8 and 9-figure Shopify brands, from research and design to development, QA, and the final read. Every test is sized before it launches, so the result means something when it ends. Book a free strategy call to talk through your testing roadmap.