Skip to content
MarketingSep 23, 202622 min

Incrementality Testing: What Your Paid Media Actually Added

Your ad platform reports the conversions it was assigned credit for, not the conversions it caused. Google draws that line itself, in two different help articles. Its attribution models documentation (opens in new tab) says “Attribution models let you choose how much credit each ad interaction gets for your conversions,” while its Conversion Lift documentation (opens in new tab) describes incrementality experiments as “the way to measure the causal impact of ads.” Incrementality testing is how you measure the second thing.

Every number below links to its source. Google and Meta both sell the media and grade the media, so their figures are labeled as vendor material where they appear, and the sample-size arithmetic is ours, computed from published formulas and labeled as ours everywhere it shows up.

We run paid media in-house, and whether a client’s account can support an incrementality test is a question that comes up every time a budget needs defending. This is the analysis we run before anyone spends money on a lift study.

What Incrementality Testing Measures That Attribution Doesn't

Attribution divides a conversion that already happened. Google’s attribution documentation (opens in new tab) explains that customers may interact with multiple ads from the same advertiser, and that the models set how much credit each interaction receives. Last click “Gives all credit for the conversion to the last-clicked ad and corresponding keyword,” and the data-driven model is the default for most conversion actions. Switching between models moves credit between campaigns and keywords.

What it doesn’t move is the number of sales your advertising produced. Someone who searched your brand name, clicked the ad at the top of the page, and bought something they’d already decided to buy generates a conversion in every attribution model available. None of those models ask whether they’d have scrolled to the organic listing and bought anyway.

That’s the question incrementality testing asks instead. Google’s Conversion Lift page (opens in new tab) describes the tool as measuring “the number of purchases, site visits, and any other conversions directly driven by people seeing your ads.” The same page states that “Conversion Lift isn’t available for all Google Ads accounts. To use Conversion Lift, contact your Google account representative.”

Incrementality Testing Requires A Group That Sees No Ads

The mechanism that separates an incrementality test from every other kind of measurement is a holdout: a randomly chosen group of users or geographies that the ads are withheld from. You compare what that group did against what the exposed group did, and the difference is the lift. Randomization is what makes the comparison causal, because the two groups differ only in whether they saw the ad.

Without a holdout, you’re comparing people who saw ads against people who didn’t see ads for reasons you never controlled. Those reasons usually correlate with buying. Someone who visits your site three times a week sees your retargeting ads because they’re already interested, and crediting the ad with their purchase is exactly where observational estimates break.

What Incrementality Testing Found When Someone Ran It At Scale

eBay Turned Off Brand Search And 99.5% Of The Clicks Came Back

eBay ran two separate suspensions of its paid search and measured what happened to sales: one that turned brand keywords off on Yahoo! and MSN while leaving them running on Google as a control, and one that turned non-brand keywords off across a large share of US markets. Both ran in 2012, brand keywords in March and non-brand across April to July.

The study is published as NBER Working Paper 20171 (opens in new tab) by Thomas Blake, Chris Nosko and Steven Tadelis, and appeared in Econometrica in January 2015. The cover page discloses that Nosko and Tadelis were employed by eBay Research Labs at the time, so read it as an insider experiment with published methods rather than as independent work.

On brand keywords, the substitution was close to total. The paper reports (opens in new tab) that “The experiment revealed that almost all (99.5 percent) of the forgone click traffic from turning off brand keyword paid search was immediately captured by natural search traffic from the platform, in this case Bing.” Nearly every click eBay had been paying for was a click it would have received for free.

Non-brand keywords did produce incremental traffic. “Advertising clicks dropped 41 percent, and total clicks fell 2 percent as a result of the non-brand experiment,” which the authors read as roughly 42 percent of paid search clicks being newly acquired. The design (opens in new tab) suspended ads in roughly 30 percent of DMAs, using 68 test DMAs against 142 control DMAs, 65 of them matched to the test markets by an algorithm that matched historical serial correlation in sales and the remaining 77 simply never drawn as candidates, analyzed as difference-in-differences.

The return-on-investment table is where the two measurement approaches separate. Measured observationally, the same spend (opens in new tab) returns 4,173 percent without time and geographic controls and 1,632 percent with them. Measured off the experimental variation, it returns negative 63 percent, with a 95 percent confidence interval of negative 124 to negative 3 percent. The spend under examination was about $51 million in annual US paid search against $2,880.64 million in gross revenue.

our best estimate of average ROI using the experimental variation is negative 63 percent as shown in Columns (4) and (5). This ROI is statistically different from zero at the 95 percent confidence level.

That result doesn’t transfer to a small brand, and the paper explains why. Its abstract reports that “new and infrequent users are positively influenced by ads but that more frequent users whose purchasing behavior is not influenced by ads account for most of the advertising expenses,” and page 16 states that “Figure 4 implies that search advertising works only on a firm’s least active customers.” eBay’s problem was that nearly everyone searching for it already knew where to find it, which is the opposite profile from a DTC brand nobody searches by name. It’s also a short-run return on search spend, not a verdict on paid media generally.

Facebook's Own Experiments Read 416% Naive And 77% Measured

The second set of experiments comes from inside Facebook. Brett Gordon and Florian Zettelmeyer of Kellogg, with Neha Bhargava and Dan Chapsky of Facebook, compared randomized experiments against the observational methods advertisers normally use, in a white paper dated July 14, 2016 (opens in new tab). The cover discloses that Gordon and Zettelmeyer hold no financial interest in Facebook and weren’t compensated. That white paper covers 12 US lift studies comprising 435 million user-study observations and 1.4 billion total impressions.

One study makes the gap concrete. The naive comparison (opens in new tab) of exposed users against unexposed users estimates a lift of 416 percent. Balancing on age and gender brings it to 221 percent, and the best propensity-score model the authors fit brings it to 102 percent. The randomized experiment on the same users reports 77 percent.

Each step in that sequence is a more sophisticated correction, and the best observational estimate still landed 25 percentage points above the experimental one. In study 9 (opens in new tab) the experimental lift is 2.5 percent while the observational methods range from 1,413 to 3,288 percent.

Across checkout outcomes, exact matching disagreed (opens in new tab) with the experiment in 10 of 11 studies, or 91 percent of the time, with an average absolute lift deviation of 661 percentage points against an average experimental lift of 57 percent. The best-performing method in the paper, inverse-probability-weighted regression adjustment with the richest variable set, still disagreed in 3 of 10 studies with an average deviation of 173 percentage points. The deviations run in both directions, so there’s no correction factor to apply after the fact.

Facebook didn’t show its control users nothing. It served each control user the ad they’d have seen if the advertiser’s campaign had never run, which is the second-place ad in the auction. So the measured lift is lift over the next advertiser’s ad rather than lift over an ad-free world, which is the right counterfactual for a spend decision and a different quantity from what most merchants picture when they hear “incremental.”

The paper also notes that because users log in across devices, control users were never inadvertently exposed, which sidesteps the cookie-identity problem that leaks into smaller holdouts. All 12 studies were randomized controlled trials held in the US across retail, financial services, e-commerce, telecom and tech, each running January 2015 or later on at least 1 million users with conversion tracking already in place. A version of the paper was later published in Marketing Science in 2019; the counts above are from the 2016 white paper.

The Four Incrementality Testing Designs You Can Run

Google’s guide to implementing campaigns for geo experiments (opens in new tab) names three study types, and splits one of them in two at implementation time. Google frames the first as “Holdback (net new / prove new value): Validate the incremental returns of entirely new ad strategies and unproven channels.”

  • Holdback. You withhold the ads from a share of users or geographies and leave everything else running as it was.
  • Go-dark, uncapped. You switch the campaigns off entirely in the test geographies, which is Google's targeting-only path for campaigns that aren't budget constrained.
  • Go-dark, capped. Google's separate path for budget-constrained campaigns, where you switch off in the test geographies and cap the budget so the freed spend doesn't pour into the control geographies.
  • Heavy-up. You increase spend in the test geographies and measure what the increase bought.

The capped version exists because turning campaigns off in half your markets doesn’t stop the budget from spending. Google’s instruction (opens in new tab) is to work out what share of spend historically landed in the control geographies, then multiply the existing daily budget by that share to get the capped campaign’s new one. Skip it and the control geographies absorb the test budget, which lifts the control arm and shrinks the measured difference.

The same page names single-cell, multi-cell and cross-publisher multi-cell structures, which is how you test more than one change or more than one platform inside a single geographic split.

A/B Testing Isn't Incrementality Testing

Does a campaign experiment measure incrementality? No. Google’s custom experiments (opens in new tab) split traffic between an original campaign and an experiment campaign, and Google recommends “using 50% to provide the best comparison between the original and experiment campaigns.” Both arms see ads, so the output is which configuration performed better, never whether either one added a sale.

The split method changes what the result means. Cookie-based assignment “randomly assigns users to either your experiment or original campaign and ensures that a given user only views either the original or the experiment.” Search-based assignment reassigns on every search, so per the same page (opens in new tab), “If a user runs multiple searches, the same user could view both the experiment and your original campaign,” and Google notes that option may reach statistically significant results faster.

Meta’s A/B testing works on the same principle. Per Meta’s split testing documentation (opens in new tab), it splits the audience with no overlap between groups, supports treatment percentages per cell, and covers variables including audience, delivery optimization, placement, creative and budget, with Meta advising one variable per test. Its documented limits are up to 100 concurrent studies per advertiser, up to 150 cells per study and up to 100 ad entities per cell. Meta says tests with larger reach, longer schedules or higher budgets tend to deliver more statistically significant results, without quantifying that.

Neither Google’s custom experiments page nor the Ads API experiments overview (opens in new tab) states a significance threshold or a minimum duration. The duration and the read are yours to decide.

What Google's Incrementality Testing Products Require

Conversion Lift Based On Users: A $5,000 And 1,000-Conversion Floor

Google publishes its entry requirement for the user-based study directly. Per the setup page for Conversion Lift based on users (opens in new tab): “For advertisers looking for directional lift (lower than 90% confidence), any budgets above $5,000 USD and 1,000 conversions will allow access to these results. You won’t be able to save your study if the budget is below $5,000 USD.” That floor belongs to the user-based study, which is the only Conversion Lift variant with a published numeric threshold.

The same page (opens in new tab) allows “studies as short as 7 days” while typically recommending more than 14 days, requires that “All conversions must be set up to fire unconditionally,” and states that if “your study power is below 90%, budget guidance will be provided to help you achieve 90% certainty of lift.” Certainty is displayed as a percentage range from 50% to 95%, in 5% increments.

The holdback is adjustable across a 1% to 50% range (opens in new tab), and studies running alongside Brand Lift and Search Lift have to leave the 30% holdback in place. Eligibility requires the account to track at least one compatible conversion action, directly or through a manager account, and Google recommends enhanced conversions for web to raise what it calls data strength.

Google also claims on that page that studies with a long conversion lag that run for less than 14 days can see up to a 17% drop in Absolute Lift (opens in new tab). That figure is published with no methodology, no sample size and no date, so it’s a vendor rule of thumb rather than a measured effect you can plan against.

How long should an incrementality test run? Long enough that the window covers your conversion lag, on top of Google’s own recommendation of more than 14 days. A store where people browse for three weeks before buying will under-report lift in a 14-day study, because purchases the ads caused during the window land after it closes.

Conversion Lift Based On Geography Publishes No Minimum At All

Google publishes a hard dollar floor for the user-based study and publishes no numeric minimum spend, minimum conversion count or minimum number of regions for the geography-based study (opens in new tab) as of September 15, 2026. In place of a number it shows a feasibility grade with three levels, High, Medium and Low, computed inside the tool against your own campaigns. Google says a “High” status “will give you the best chance at generating statistically significant results,” says it doesn’t recommend proceeding at low feasibility, and directs you to update your budget against its recommendation. Google does say elsewhere that the geography-based study’s budget requirement tends to run higher than the user-based study’s, without attaching a number to it.

An absent number isn’t the same as no minimum. It means there’s no published figure to check your account against before you build the study, and the feasibility grade is the only signal Google exposes. The page does define the effect-size concept it grades you on: “Minimum detectable iROAS is the effect size we need during the test in order to have a high chance to detect significant lift.”

Eligibility has one structural constraint worth checking before anything else. Per the same page (opens in new tab), “In addition to conversions, campaigns must target a single country to be eligible for the geo-split methodology.” The experimental units are Google Marketing Areas, the product is in beta and Google restricts it to holdback and go-dark studies only, so heavy-up isn’t on the menu here, and the supported campaign types are App, Demand Gen, Discovery, Display, Video, Search, Shopping and Performance Max.

Compatible conversion sources are the Google Ads Conversion Tag, the Firebase Conversion Tag, Display and Video 360 Floodlight, and bring-your-own-device support for any offline conversion type aggregated to ZIP or city level. If your conversions arrive through something outside that list, the geo study can’t read them.

The implementation guide (opens in new tab) is where the design gets specific. It tells you to expand the Location options section and “select Presence to prevent location leakage,” to use “City IDs or ZIP codes for international targets (non-USA) and DMA regions for USA-based targeting,” and that “Pre-test data must be at least 3 times the duration of the test period.” A four-week test needs twelve weeks of clean history behind it, which rules out testing on a campaign you restructured last month.

Learning periods are the other constraint that quietly decides whether the read is valid. Google states that “Any budget change greater than 20% necessitates a new learning period, whereas bid modifiers under 20% can be applied without a severe learning reset,” and for heavy-up designs says to “Factor in the typical 4-5 day campaign-dependent learning period, which must be fully included inside the test period.” Duplicating a campaign triggers a learning period as well, because the duplicate has no baseline history while the original keeps its own, and Google says to remove the learning period from the final analysis.

Geographies you excluded from the experiment altogether don’t receive business-as-usual budget when campaigns are duplicated, per the same guide (opens in new tab). The test still reads correctly, but those markets quietly stop being advertised to, so Google’s advice is to run a separate business-as-usual campaign for them.

Brand Lift And Search Lift Answer A Different Question

Brand Lift is a YouTube and Demand Gen product. Per Google’s lift comparison page (opens in new tab), it grades YouTube campaigns on perception measures: ad recall, association, awareness, consideration, favorability and purchase intent. None of those are sales.

Google’s Brand Lift setup page (opens in new tab) publishes a minimum budget requirement across 10 days that scales with the number of survey questions and the country tier, from $5,000 for one question in a Country A market up to $60,000 for three questions in Country B or Country C markets. The page states the requirements are in USD, adjusted by exchange rate and refreshed four times a year at the beginning of January, April, July and October, and it lists which countries sit in each tier. The United States is a Country B market, so a US advertiser’s floor is $10,000 for one question, $20,000 for two and $60,000 for three. That table belongs to Brand Lift, not to Conversion Lift.

Google’s Search Lift page (opens in new tab) says the product “measures the increase in searches for your product or brand after users have viewed your ad, rather than traditional metrics such as clicks, impressions, or views,” and calls it “a free tool for measuring the effectiveness of your ads.” Free refers to the measurement, not the media: Google’s lift comparison page (opens in new tab) says Search Lift carries budget minimums similar to Brand Lift’s. For a merchant asking whether ads added revenue, an increase in branded searches is an intermediate signal rather than the answer.

Google’s published entry requirements for each product, as of September 15, 2026:

Product What it measures Published entry requirement Access
Conversion Lift based on users Conversions driven by seeing your ads Above $5,000 USD and 1,000 conversions for directional results; study won’t save below $5,000 Contact your Google account representative
Conversion Lift based on geography The same, split across Google Marketing Areas No numeric minimum published; High / Medium / Low feasibility grade instead Contact your Google account representative
Brand Lift Ad recall, awareness, consideration, favorability, purchase intent on YouTube $5,000 to $60,000 across 10 days, by question count and country tier; the US is a Country B market, so $10,000 / $20,000 / $60,000 Contact your Google account representative
Search Lift Increase in searches for your product or brand after ad views No figure of its own; Google says budget minimums are similar to Brand Lift’s, and calls the tool itself free Contact your Google account representative
Custom experiments Which of two campaign configurations performs better No significance threshold or minimum duration stated Self-serve

Google Reports A Credible Interval, Not A Confidence Interval

Google’s Bayesian methodology page (opens in new tab) states: “Bayesian methodology uses ‘Credible Intervals’ which are similar to ‘Confidence Intervals,’ but have a different interpretation. The true value of your lift has an 80% probability of being within the Credible Interval, which is determined by combining your experiment’s data with the historical campaign data.”

That last clause changes what the number is. The interval blends your experiment with your account’s history rather than reading the experiment alone, so if your historical campaign data leans positive, the reported interval inherits some of that lean. A Google credible interval and a confidence interval from a test you ran yourself are two different objects, and putting them side by side in a deck implies a comparison neither one supports.

If you’re reporting lift to a board or a client, call it a credible interval and say it’s blended with historical data. Calling it 80% confidence misstates what Google computed.

Meta's Incrementality Testing Runs Through A Rep

Meta’s Lift Studies developer guide (opens in new tab) opens by stating that access to Conversion Lift measurement is currently limited, and telling advertisers to contact their Meta representative for access. Meta’s incrementality product isn’t self-serve, the same as Google’s Conversion Lift on the campaign types an ecommerce advertiser actually runs, since Google gates Display, Search, Shopping and Performance Max behind a rep while leaving Video, Discovery and Demand Gen open. If you’re running Meta spend without a rep relationship, the holdout study isn’t on the menu, and the A/B testing that is available compares ads against ads.

Per the same guide, creating a study creates a randomized test group of Accounts Center accounts that see the ads and a control group that doesn’t. A holdout percentage defines the accounts that won’t see ads, treatment plus control percentages have to total 100, and the control percentage sets the holdout for each test group relative to the total population.

Meta’s Ad Study API reference (opens in new tab) documents a per-cell treatment percentage minimum of 10, with cell treatment percentages summing to no more than 100, and each cell needing at least one associated ad account, campaign or ad set. Google’s 1% to 50% holdback range belongs to Google’s user-based study, and Meta publishes no equivalent holdout range, so don’t carry Google’s numbers across.

Operationally, a study’s start time has to be in the future, it needs at least one objective, and objectives can’t be modified once it’s running. Start time and per-cell treatment percentage lock once a study runs, while end time can be extended while the study is still going. Cell-level results require the cell_id breakdown, and the cooldown_start_time parameter is marked deprecated.

Studies started after July 13, 2021 no longer report buyers metrics or breakdowns by gender, age and country. If you’re planning to slice a lift result by demographic, check what the API will actually return before you design around it.

Meta documents GEO_LIFT as a study type alongside LIFT, SPLIT_TEST, CONTINUOUS_LIFT_CONFIG, BACKEND_AB_TESTING and others, so geo designs exist on Meta as well as user-level holdouts.

We checked the lift studies guide, the Ad Study reference and the top level of Meta’s Marketing API changelog (opens in new tab) for a dated retirement or rename of a Meta lift product and found none as of September 15, 2026. The dated changes we did find are the cooldown_start_time deprecation and the metrics cutoff above, and neither retires a product. That’s an absence in those pages rather than a statement that nothing has changed.

The Arithmetic Behind Incrementality Testing

Whether a test can work is a sample-size question, and it’s answerable before you spend anything. The arithmetic below is ours, computed from the standard two-proportion formula rather than taken from a vendor. Assumptions, stated once: a two-sided test at alpha 0.05, 80% power, two equal-sized arms, conversion measured per user with one outcome per user, no peeking at interim results, no clustering and no covariate adjustment.

Real tests come in worse than this, not better. Peeking, seasonality, holdout contamination and multi-device identity all cost power, and an uneven split costs more. Treat every figure below as a floor.

The useful version expresses the requirement in conversions rather than users, because conversions are the number a merchant already knows. Detecting a 10% relative lift takes roughly 1,600 conversions per arm. Detecting 5% takes roughly 6,300 per arm, detecting 20% takes about 420, and detecting 50% takes about 78.

Those figures barely move across conversion rates. At a 0.5% conversion rate a 10% lift needs 1,640 conversions per arm, and at 3% it needs 1,597. The binding input is the size of the lift you’re trying to detect, not your conversion rate.

BLKDG computation. Baseline conversions needed per arm, two-sided alpha 0.05, 80% power, equal arms:

Baseline conversion rate +5% lift +10% +20% +50%
0.5% 6,404 1,640 430 78
1.0% 6,371 1,631 427 78
2.0% 6,305 1,614 423 77
3.0% 6,239 1,597 418 77

Run it the other way and you get the smallest lift your existing conversion volume can detect. At 250 conversions per arm, that’s about 26.6%. At 1,000 it’s about 12.9%, at 2,000 about 9.0%, at 5,000 about 5.7%, and at 10,000 about 4.0%.

BLKDG computation. Minimum detectable relative lift, same assumptions:

Conversions per arm CVR 0.5% 1.0% 2.0% 3.0%
250 26.6% 26.5% 26.4% 26.2%
500 18.5% 18.4% 18.3% 18.2%
1,000 12.9% 12.9% 12.8% 12.7%
2,000 9.0% 9.0% 9.0% 8.9%
5,000 5.7% 5.7% 5.6% 5.6%
10,000 4.0% 4.0% 4.0% 3.9%

What "No Significant Lift" Means At Google's Incrementality Testing Floor

Google will let you save a user-based Conversion Lift study at $5,000 in budget and 1,000 conversions, and labels results at that level “directional lift (lower than 90% confidence) (opens in new tab).” At 1,000 conversions per arm, the smallest relative lift detectable at 95% and 80% power is about 12.9%, by our arithmetic above.

So a “no significant lift” readout from a study sitting at Google’s published floor means the lift probably isn’t enormous. It doesn’t mean the lift is zero. A true incremental lift of 9% sits below what that study can reliably detect, and 9% incremental revenue on a paid budget is a different decision from 0%.

Google’s page states the threshold as 1,000 conversions without specifying whether that’s per arm or across the whole study. Read as per arm, our table gives about 12.9%. Read as the study total with an even split, it’s 500 per arm and about 18.5%, and Google’s holdback range runs from 1% to 50% (opens in new tab), so any split more uneven than 50/50 pushes the detectable lift higher again.

Decide the smallest lift that would change what you do before you run anything. If your conversion volume can’t detect that lift, the study will report no significant lift whatever your ads are actually doing, and you’ll have bought a result the arithmetic could have given you for free.

Geo Experiments Aren't A Shortcut Around Incrementality Testing Math

A geo experiment sidesteps the user-level identity problem and introduces a variance problem instead. The Gordon et al. authors simulated (opens in new tab) 10,000 random allocations of 40 matched markets in a setting where the true lift is 33%, and report that “95% of the time the researcher will estimate a lift between -2% and 80%.”

A true lift of 33% can read as a small negative or as more than double the truth, depending on which markets happened to land in which arm. Increasing to 80 markets narrows the range without closing it. Reporting a single matched-market test’s point estimate as “the” incremental lift claims more precision than the design produces.

Our own market-level arithmetic reaches the same place from the design side. The number of markets you need per arm depends on how much your markets vary relative to their mean, and at a coefficient of variation of 0.4, detecting a 10% lift takes about 252 markets per arm. There are only 210 Nielsen DMAs in the United States.

BLKDG computation. Markets needed per arm, where the standard deviation across markets equals the coefficient of variation times the mean, unpaired, two-sided alpha 0.05, 80% power:

Coefficient of variation +2% lift +5% +10% +20% +30%
0.2 1,570 252 63 16 7
0.4 6,280 1,005 252 63 28
0.6 14,128 2,261 566 142 63
0.8 25,117 4,019 1,005 252 112
1.0 39,245 6,280 1,570 393 175

That model is deliberately naive, and most of the table is unreachable by construction given 210 DMAs. It excludes pairing, pre-period covariates and synthetic control, all of which cut variance substantially, and all of which are why serious geo work uses them. eBay’s design sat at the serious end: 68 test DMAs against 142 controls, with 65 of those controls matched to the test markets on historical serial correlation in sales, analyzed as difference-in-differences.

Read the table as a statement about effect size rather than as a sample-size recommendation. Geo tests work when the effect is large, the markets are matched carefully, or the analysis is smarter than a two-sample comparison. When none of those hold, the test still returns a number, and the number carries the spread shown above.

There's No Average Incremental Lift To Benchmark Against

We went looking for a third-party, methodologically documented average incremental lift for ecommerce paid media and didn’t find one. The figures that circulate as benchmarks are vendor case studies, published without sample sizes, test durations or intervals, and they dead-end at the vendor’s own marketing page.

The best available evidence on the spread is Table 6 of the Gordon et al. white paper (opens in new tab). Across the eleven of its twelve studies that measured checkout, experimental lifts range from negative 3.6% to 418%, with several not statistically different from zero. Eleven advertisers, one platform, one methodology, and the answers span more than 400 percentage points.

An average drawn from that distribution wouldn’t describe any of the eleven advertisers in it. If someone quotes you an average lift figure, ask for the sample size, the test duration and the interval before it goes anywhere near a forecast.

Incrementality Testing When You're Too Small To Power A Test

Plenty of stores can’t produce 1,600 conversions per arm inside a window they’d be willing to wait out. Measurement doesn’t stop there, it just changes shape.

Split brand from non-brand before anything else. eBay’s experiment measured the two separately and got opposite answers, with brand clicks almost entirely substituted by organic and non-brand clicks roughly 42% newly acquired. Brand search is where substitution risk concentrates, which makes it the first line to examine even when your volume can’t support a formal study.

A go-dark test on one channel across a set of markets is still a real experiment, and it’s the design Google documents for unproven channels. Run it long enough to cover your conversion lag, hold the pre-period at three times the test length as Google’s implementation guide (opens in new tab) requires, and read the output as directional. A before-and-after comparison with no control geography is one of the observational approaches the Facebook paper found unreliable.

Fix conversion tracking before you test anything at all. Google’s user-based study requires that all conversions fire unconditionally, and a study built on a purchase event that stopped firing will report a confident number that’s wrong in both arms. If your tracking lived in Shopify’s Additional Scripts field, check it against what broke in Shopify conversion tracking on August 26, and confirm your GA4 setup and reporting is measuring purchases rather than key events you don’t care about.

Attribution still has a job once you know what a channel adds. Our piece on multi-touch attribution and ETL tools covers how to consolidate the data those models run on. Incrementality sets the channel budget, and attribution distributes it inside the channel.

How We Approach Incrementality Testing For Clients

We run paid media in-house, and the sequence is the same for every account. Confirm that conversion tracking fires correctly and unconditionally. Compute the smallest lift the account’s conversion volume can detect in a window the client will actually wait out. Then pick the design that fits, which is a user-level holdout when the platform grants one and the volume supports it, and a geo holdback or go-dark when it doesn’t.

When the arithmetic says a formal study can’t return a readable answer, we say so instead of running it. A study that was always going to report no significant lift spends budget and calendar time without changing a decision.

If you’re carrying paid spend you can’t defend, start with a free Growth Audit from our ROAS optimization team. Not a sales call. A no-obligation look at what your account is producing, what’s measurable at your current volume, and what it would take to prove it. We also handle paid media management across Google, Meta and Shopping for brands that want the buying and the measurement under one roof.

About the author

BLKDG Team

Not sure where the gap is? That's exactly what the Digital Marketing Growth Audit is for.

A free, no-obligation look at where your site can win more traffic and conversions, with a clear digital marketing roadmap to get there. Just a straight read on where your digital presence stands and where it's headed.