BLKDG is a Shopify agency that runs testing and optimization programs for ecommerce brands. The first decision in any of those programs is what’s worth testing, because traffic limits how many trustworthy tests a store can run. Ron Kohavi, Alex Deng and Lukas Vermeer, who write that they’ve been involved in tens of thousands of A/B tests at Airbnb, Booking, Amazon and Microsoft, put the floor this way in a 2022 paper on common A/B testing misunderstandings (opens in new tab): “A/B tests are useful to detect effects of reasonable magnitudes when you have, at least, thousands of active users, preferably tens of thousands”.
For ecommerce A/B testing on Shopify, Shopify Rollouts runs Experiments natively on your theme, your checkout and accounts configuration, your catalogs and your discounts. It reaches further than a theme split, and it reports a short, fixed list of metrics.
Ecommerce A/B Testing Starts With What's Worth Testing
Optimizely’s own 2023 benchmark report (opens in new tab) counted more than 127,000 true experiments run on its platforms between 2018 and 2023 and found that around 12% of them won on the primary metric. A win in that report meant the metric moved in the winning direction at 90% statistical significance or higher, which is a lower bar than the 0.05 threshold Kohavi, Deng and Vermeer call the industry standard. That’s a testing vendor describing its own customers, and it still leaves 88% of experiments without a win. Most ideas in a backlog won’t produce a detectable lift, so each test should answer the highest-value question your traffic can support.
Size of change is the second filter. Microsoft’s experimentation team wrote in 2013 that teams whose products have thousands to tens of thousands of users are typically looking for larger effects (opens in new tab), which are easier to detect than the small effects large sites worry about. For a store, that points at the offer, the price, the page structure and the checkout layout, and away from button colors and microcopy.
A backlog built that way comes from evidence about where shoppers drop out. A conversion rate optimization audit is where we find those candidates before any ecommerce A/B testing begins, and our list of five product page tests worth running shows what a well-formed hypothesis looks like on a PDP. Whether your ad spend is producing sales that wouldn’t have happened anyway is a separate question with its own method, covered in our guide to incrementality testing for paid media.
Shopify A/B Testing Is Native: What Rollouts Experiments Cover
Rollouts sits under Markets > Rollouts in the Shopify admin. You can create three types of rollout (opens in new tab): a Launch, an Event and an Experiment. A Launch is a change you keep, an Event runs for a set period, and an Experiment is the A/B test. An Experiment tests two versions of a change set, a treatment and a control, against each other without affecting all of your visitors.
A rollout can carry four kinds of change: one change to your main theme, one change to your checkout and accounts configuration, one or more catalog changes, and one discount change. Each rollout takes only one theme change and one checkout and accounts change (opens in new tab).
| Change type | What the treatment does | Limit |
|---|---|---|
| Theme | Edits the main theme in the theme editor, or replaces it with another theme | Vintage themes are excluded, and Liquid templates can’t be changed as part of a rollout |
| Checkout and accounts configuration | Changes the checkout and accounts configuration | Online store only. Headless and custom storefront checkouts aren’t supported |
| Catalog | Sets a market catalog to Active or Draft, which publishes or withdraws its prices and product availability | Changes the catalog’s status only. Catalogs created for a B2B company location can’t be added |
| Discount | Activates or deactivates existing discounts for visitors in the rollout | Up to 250 discounts per rollout, and a discount can be in only one unended rollout at a time |
The theme limits decide whether a test fits Rollouts at all. You can’t apply a rollout to a vintage theme (opens in new tab), and you can’t change Liquid templates as part of one. Experiments, like any rollout that publishes to less than 100% of visitors, apply only to your online store. If your storefront is headless, theme and checkout Experiments aren’t available to it.
Plan Gates And Traffic Settings
Rollouts has three separate plan lines. Rollouts are available on the Basic plan or higher (opens in new tab), and Experiments need the Grow plan or higher. When Shopify announced Rollouts on March 31, 2026, it made market-specific targeting (opens in new tab) available on the Advanced and Plus plans. A Basic store can schedule a Launch but can’t run the A/B test, and a Grow store can run the test but can’t aim it at one market.
Two dials control exposure. Traffic allocation is the share of eligible visitors who enter the rollout, and traffic split is how those visitors divide between control and treatment. A new Experiment defaults to a traffic split of 50% control and 50% treatment (opens in new tab), with an end date 90 days after the Experiment starts. Set allocation to 50% with a 50/50 split and a quarter of your eligible visitors see the treatment, which doubles the time the test needs.
Concurrency costs traffic the same way. You can run multiple Experiments on the same resource (opens in new tab) at once, and each one gets its own mutually exclusive segment of visitor traffic. Set four overlapping theme rollouts to 100% each and each rollout’s effective allocation is 25% of your traffic.
Four simultaneous theme Experiments each take four times as long to reach the same sample, and an active Launch on the same resource gets its traffic first.
Three housekeeping rules protect a result. Editing the treatment or control copy while an Experiment is active (opens in new tab) can affect its results, so freeze both versions once the test starts. After you apply a rollout, you can’t revert the changes. Archiving a rollout permanently deletes its theme changes, so keep a losing treatment unarchived until you’ve recorded what it changed.
What Each Experiment Type Reports
You can’t customize the metrics (opens in new tab) a rollout reports. A theme Experiment reports four: conversion rate over time, bounce rate, reached checkout rate and add to cart rate. A checkout and accounts Experiment reports one, checkout conversion rate, which is online store sessions with completed purchases relative to the visits that reached checkout.
Gross sales, average order value and orders are Launch metrics, so a theme Experiment reports no revenue metric. Conversion rate over time is the percentage of online store sessions that lead to an order over time, which means a treatment that converts more sessions into smaller orders reads as a win. If your hypothesis is about order value, such as a bundle, an upsell or a free shipping threshold, a theme Experiment’s four metrics can’t confirm it.
Two more reporting details affect how you read the numbers. Launches and Experiments calculate metrics with slightly different algorithms, so the figure you see after applying a winner as a Launch won’t necessarily match the figure from the Experiment. And the analytics measure visitors on your online store, cart and checkout, with headless storefronts and Shopify POS left out of the results.
Set the sample size, the threshold and the stopping date before the Experiment starts. Replace the 90-day default end date with one set by the sample your test needs.
Price And Discount Tests In Shopify Rollouts
A catalog change is a native price test. Setting a catalog to Active (opens in new tab) publishes its prices and product availability, and setting it to Draft withdraws it.
The Experiment controls the catalog’s status and nothing else. Edits to the catalog’s prices, price adjustments and products (opens in new tab) are saved right away, and they apply whenever the catalog is active. Finish editing the catalog before the test starts, because a price you change mid-test changes the treatment.
Catalog Experiments are limited to catalogs you create for markets, and a catalog can be part of only one unended rollout at a time (opens in new tab). When a catalog change is in an Experiment, or in a rollout that publishes to less than 100% of visitors, your online store can display different prices or products than Google or other sales channels. Product ads, listings, feeds and third-party apps might not display the same offer.
Check your Shopping ads and feeds against both prices before you launch, and note that the View as preview shows base prices during an Experiment.
Discount changes test a promotion against your store’s standard discount behavior. A rollout takes up to 250 discounts (opens in new tab), and in an Experiment no discount is included in the control.
Two caveats weaken the read. A discount code that’s part of an Experiment can reach visitors outside the group you’re testing (opens in new tab), and for a rollout with a discount change, exposures match assignments, which might skew results. Discount-specific metrics, such as the revenue or the orders attributed to a discount, aren’t available.
How Much Traffic Ecommerce A/B Testing Needs
The standard sizing rule, as Kohavi and his co-authors print it (opens in new tab), sets the users needed in each of two equal variants at 16 times the variance, divided by the square of the change you want to detect, assuming 80% power and a p-value threshold of 0.05. The variance of a conversion rate is the rate multiplied by one minus the rate, and the change is the smallest absolute difference you want to detect. Their worked example uses a 3.7% baseline and a 10% relative lift, and it comes to 41,642 users per variant.
The same formula appears in Kohavi’s 2009 survey of web experiments (opens in new tab). Applied to a 2% baseline, it works out to 78,400 users per variant for a 10% relative lift, 19,600 for a 20% lift and 8,711 for a 30% lift. Halving the lift you want to detect quadruples the sample. That arithmetic is why ecommerce A/B testing on a mid-sized store works for large changes and stalls on small ones.
Three details change the calendar. The formula counts users, while Shopify defines its Experiment metrics on sessions, so a visitor who returns three times isn’t three units of sample. The survey’s authors recommend running an experiment for at least a week or two, then extending it by multiples of a week so day-of-week effects can be analyzed. They also show that uneven splits are slow: a test run at 99%/1% has to run about 25 times longer than one run at 50%/50%, so leave the split at Shopify’s 50/50 default.
Consent settings can shrink the measured sample too. In cookie banner regions, Shopify collects non-essential data only after obtaining consent (opens in new tab), which can be observed in decreased session counts and in other metrics that rely on session data, including conversion rates. Shoplift, a testing app, says that when analytics consent is declined, no visitor data is collected, logged or stored (opens in new tab) in any way. If you sell into consent-gated markets, size the test on measured visitors, not total traffic.
Our guide to ecommerce conversion rate optimization works a full store model through this formula and covers peeking, sample ratio mismatch and the winner’s curse in detail.
Ecommerce A/B Testing When You Can't Reach The Sample
Should a low-traffic store run A/B tests at all? Sometimes it shouldn’t. Evan Miller puts the subjects needed per branch (opens in new tab) for a 10% relative change at 30,244 on a 5% baseline and 157,697 on a 1% baseline, and warns that tests on single-digit conversion rates have a good chance of being a waste of time. If the sample for your test is more than a few months of traffic away, ship the change on evidence from customer research and session data, and save testing for a question you can answer.
Testing a bigger change is the first alternative. The same Microsoft team wrote in 2013 that startups have used controlled experiments when they’ve had thousands of active users and are typically looking for large effects, and the sample-size rule explains why: a tenfold gain in sensitivity costs a hundred times the users. A redesigned product page, a different offer or a restructured collection can produce an effect a store’s traffic can detect. A reworded button usually can’t.
Running the test underpowered and believing the result is the option to rule out. Kohavi, Deng and Vermeer (opens in new tab) calculate the share of statistically significant wins that are false positives when 10% of ideas truly work, at the standard 0.05 two-tailed threshold: 22.0% at 80% power and 52.9% at 20% power. They recommend replicating any surprising result, and they point out that in software, replication is much cheaper and easier.
A sequential design lets you stop early. Miller’s sequential procedure (opens in new tab) fixes a total number of conversions in advance, splits traffic evenly, and stops when the treatment’s lead over control reaches a preset boundary or the total is hit with no winner. Tools with sequential statistics do the same job at a price: Eppo says its sequential method produces wider confidence intervals (opens in new tab) for the same data than its fixed-sample method.
The last option is a metric with a higher base rate. On Miller’s figures, a 20% baseline needs 6,347 subjects per branch for the same 10% relative change, about a fifth of what a 5% baseline needs. A theme Experiment in Rollouts reports add to cart rate and reached checkout rate, and both reach a sample sooner than purchases do because more sessions add to cart than buy. The trade is that you’ve measured add-to-cart behavior, and a lift there isn’t proof of a lift in orders.
Reading An Ecommerce A/B Testing Result When Tools Use Different Thresholds
Kohavi, Deng and Vermeer call a p-value threshold of 0.05 the industry standard. Testing tools don’t all apply it, and several report a different quantity. The table below gives each vendor’s own account of its statistics.
| Tool | What it reports | What the vendor says about calling a result |
|---|---|---|
| Optimizely sample size calculator (opens in new tab) | Statistical significance | Treats 95% as an accepted standard for statistical significance, adjustable from 80% to 99% |
| Optimizely 2023 benchmark report (opens in new tab) | Statistical significance | Counts a winning experiment at 90% or higher |
| Eppo (opens in new tab) | Confidence intervals | The confidence level is 95% by default |
| AB Tasty (opens in new tab) | Bayesian chance to win | The progress bar turns green when the chance to win is higher than 95% |
| Intelligems (opens in new tab) | Bayesian probability, from Monte Carlo simulations | Recommends at least 300 orders per test group and 7 days, and says meeting those minimums “does not guarantee statistical significance” |
| Shoplift (opens in new tab) | Bayesian Probability to Win | Probability to Win has to hold at or above a level for multiple days before Shoplift assigns a stage |
A Bayesian 95% probability to win and a frequentist 95% significance level are different quantities, so a result that clears one tool’s bar hasn’t necessarily cleared another’s.
Shoplift says of the traditional bar of 95% confidence, “Most real ecommerce tests never reach it in a reasonable window”. Its reports use stages, and it says of the one it calls Leaning, “This is enough to act on when the stakes are low or you want to move quickly.” That’s a defensible rule for a reversible change like a headline. For a price or a checkout layout, wait for a Test Complete result.
Shoplift also declares no consistent difference when a test has run for two full weeks with no meaningful signal, and calls that “a real, conclusive result, not a failure.” Two weeks of a small store’s traffic can sit far below the sample the formula above asks for. A test that short of its sample can’t separate no difference from a difference it didn’t have the users to detect.
Vendor minimums aren’t power calculations either. At 80% power and a 0.05 threshold, Intelligems’s 300 orders per group on a store converting 1% to 2% of sessions is enough to detect a relative lift of about 24%, and smaller lifts need more. Intelligems says as much about its own reports: a test can reach its threshold while the confidence interval around uplift and value is still very wide.
So which threshold should you use? Pick it before the test and write it down with the rest of the plan. Kohavi and his co-authors give the reason in their 2009 survey: decide the success metric in advance, as a planned comparison, or you raise the risk of finding results that only appear significant by chance. OEC is their term for that one agreed success metric. A plan for an ecommerce A/B testing result you can trust has five entries:
- The single primary metric the test will be judged on
- The baseline rate and the smallest lift worth detecting
- The sample size per variant that follows from those two numbers
- The threshold, and whether the tool reports significance or a Bayesian probability
- The end date, rounded up to whole weeks
Checking the dashboard during the run is fine. Miller’s rule (opens in new tab) is that peeking at the data is OK as long as you can restrain yourself from stopping an experiment before it has run its course.
Ecommerce A/B Testing Without Harming Search
Google sets four rules for website testing (opens in new tab). Don’t cloak, which means don’t show Googlebot one set of URLs and humans another. Put a canonical link on every alternate URL pointing at the original, which Google recommends over a noindex tag. Redirect with a 302 (temporary) redirect, not a 301 (permanent) one, and note that JavaScript redirects are acceptable.
The fourth rule is about time. Google’s policy reads, “If we discover a site running an experiment for an unnecessarily long time, we may interpret this as an attempt to deceive search engines and take action accordingly,” and that applies most when you’re serving one content variant to a large percentage of your users. Once the test ends, remove all of its elements, such as alternate URLs or testing scripts and markup, as soon as possible. An Experiment left on Rollouts’ 90-day default after it has reached its sample runs longer than it needs to, so end it on the planned date and clean up.
Googlebot generally doesn’t support cookies, so on a cookie-controlled test it sees the version shown to browsers that don’t accept them. Small changes carry little risk: they often have little or no impact on a page’s search result snippet or ranking.
Bots matter to the statistics, since Kohavi and his co-authors warn in their 2009 survey that robots can introduce significant skew into estimates, and both Intelligems and Shoplift say they filter them. Intelligems says of its bot handling (opens in new tab), “We detect all major bots and will block execution when one is detected – this ensures that the bot will see the page without intelligems running, which avoids SEO impact.” Shoplift says of its own (opens in new tab), “Shoplift detects and blocks all major search engine bots (including Googlebot) from seeing test variants. Search engines always see your original pages”.
Google’s rule is “Cloaking counts whether you do it by server logic or by robots.txt, or any other method.” If organic search carries a large share of your revenue, ask the vendor how its bot handling fits that sentence before you install the app.
The Performance Cost Of Client-Side A/B Testing For Ecommerce
A client-side testing tool changes the page in the browser, after the original HTML arrives. To keep visitors from seeing the original first, many tools hide content until the variant is ready, which delays when the page is first displayed (opens in new tab) until the changes in any test have been applied. Google’s web.dev guidance is to weigh that cost to page performance against any benefit the test brings.
The cost lands on Largest Contentful Paint, and it isn’t confined to visitors in the test. An element that has already loaded can’t render while other code is hiding it (opens in new tab), such as an A/B testing library that’s still determining which experiment the user should be in. That delay can worsen the experience even for users who aren’t part of the experiment (opens in new tab).
The vendors describe the same trade. Optimizely says it uses a synchronous snippet to prevent flickering (opens in new tab), and that with synchronous loading the browser shows nothing but a white screen until all the external resources have fully loaded. Intelligems offers a render-blocking option (opens in new tab) that it says reduces flicker at the cost of slightly higher FCP and LCP times. Shoplift, which says it works through Shopify’s theme and template structure, puts its cost (opens in new tab) at “a minimal 1-2 point reduction in the overall Lighthouse score.”
web.dev’s first best practice is to limit A/B testing tools to the pages that are part of the test, not to delay every page. It goes further for implementation: use server-side testing (opens in new tab) to avoid any rendering cost. Performance can also corrupt the test itself. Researchers from Microsoft, Booking.com and Outreach.io note in a 2019 paper on sample ratio mismatch (opens in new tab) that a slower treatment can cause the two groups to come out unevenly sized, which they say “in most cases completely invalidates experiment results.”
Measure LCP on the tested templates before you install a testing script and again after. Our guide to Shopify speed optimization covers how to take that measurement by template and how to budget third-party scripts.
When A Third-Party Ecommerce A/B Testing Tool Is The Right Call
Rollouts is the place to start if you’re on the Grow plan or higher and your test fits its four change types and its metrics. A third-party tool earns its cost when the test needs something Rollouts doesn’t report or can’t change.
Revenue is the first case. Intelligems measures each test (opens in new tab) on conversion rate, AOV, revenue per visitor and gross profit per visitor, and Shoplift’s goal metrics (opens in new tab) include AOV and revenue per visitor. If the hypothesis is about order value or margin, you need one of those numbers as the primary metric.
Test type is the second. Shoplift runs template tests (opens in new tab) that put your current live template against a different version, which suits a single product or collection template. Intelligems tests shipping rates (opens in new tab), a change type outside Rollouts’ four. It says the carrier-calculated shipping those tests depend on needs the Advanced or Shopify Plus plan, or a paid add-on on the plan below.
AB Tasty’s Shopify app (opens in new tab) runs A/B tests, multivariate experiments and personalized experiences, and Convert offers full-stack, server-side experiments (opens in new tab), the category a headless storefront needs.
Price tests split by catalog structure. Rollouts tests market catalogs. Shoplift’s price testing (opens in new tab) requires a single-currency store and doesn’t support multi-currency setups such as Shopify Markets and Global-E.
Intelligems says it applies test prices (opens in new tab) through Shopify Cart Transform Functions on any Shopify plan, and that any channel connected to Shopify (opens in new tab) will display the highest price from your test. Whichever route you take, your feeds and ads can show one price while some visitors see another, so plan the test’s length and your ad copy around that.
Checkout is the case where the native tool often wins. Shopify restricts checkout UI extensions (opens in new tab) for the information, shipping and payment steps to Shopify Plus. Optimizely’s snippet doesn’t load on Shopify’s checkout pages (opens in new tab), Shoplift doesn’t currently support testing them directly (opens in new tab), and Intelligems’s checkout modifications (opens in new tab) require a Shopify Plus account. A Rollouts checkout and accounts Experiment is available from the Grow plan.
Redirect-based tests need one extra check. Shoplift’s theme tests (opens in new tab) send variant visitors to a preview of an unpublished theme, which takes a brief page redirect on their first page load. The sample ratio mismatch researchers cited above warn that an experiment that doesn’t redirect every variant of a page may end up with a mismatch, because some redirects may fail. Compare the visitor counts in each group against your intended split before you read any conversion numbers.
Shopify’s App Store keeps a collection of A/B testing apps (opens in new tab) if you’re comparing options.
Who Runs Your Ecommerce A/B Testing Program
An ecommerce A/B testing program needs an owner for four jobs: choosing what’s worth testing, writing the plan before launch, building the treatment without slowing the page, and refusing to call a result early. On most teams those jobs fall between the marketer who wants the answer and the developer who builds the variant. The plan in particular needs someone who’ll hold the end date when a dashboard shows an early lead.
That’s the work BLKDG does through its A/B testing and optimization services. We decide with you which questions your traffic can answer, set the sample and threshold up front, run the test, and report the result against the plan.
If your store has the traffic for a few good tests a year, spend them on the changes that could move revenue, and size each one before it starts. If you’d like a team to run that program with you, take a look at BLKDG’s A/B testing and optimization services.
Not sure where the gap is? That's exactly what the Digital Marketing Growth Audit is for.
A free, no-obligation look at where your site can win more traffic and conversions, with a clear digital marketing roadmap to get there. Just a straight read on where your digital presence stands and where it's headed.

