Incrementality Testing: A Practitioner's Guide
Geo lift design, holdout sizing, the lift and iROAS formulas, and how a test result calibrates a marketing mix model.
Definition
Incrementality testing measures the causal lift a marketing activity actually caused, by comparing a treated group against a controlled holdout that did not receive the activity. It answers what would have happened anyway, which correlational attribution cannot, and it is the strongest way to validate a marketing mix model.
Incrementality testing is a controlled experiment that measures how many conversions your advertising actually caused. You expose one randomly assigned group to a campaign, withhold it from a comparable control group, and treat the difference between them as the incremental effect. Whatever the control group converts anyway is baseline demand, not marketing performance.
Marketing attribution describes the paths that converting customers took. Incrementality testing answers a harder question: what would have happened if the campaign had never run. The two methods produce different numbers for the same media, and only one of them is a causal estimate.
What is incrementality testing?
Incrementality testing is the applied form of a randomised controlled trial. You split a population into two comparable groups, run the campaign for one and suppress it for the other, then compare outcomes over a fixed window. Because assignment is random, the only systematic difference between the groups is exposure to the media, so the difference in conversions can be read as the effect of the media.
The vocabulary is worth pinning down before anything else.
- Treatment group: the users, audiences or geographies that receive the advertising.
- Control group or holdout: the comparable population that is deliberately suppressed.
- Counterfactual or baseline: the conversions the treatment group would have produced without advertising, estimated from the control.
- Lift: the difference between actual conversions and the counterfactual.
- Incremental ROAS, or iROAS: incremental revenue divided by the incremental media spend that produced it.
Every conversion recorded in the holdout happened without the ad. If a channel reports a strong ROAS and the holdout converts at close to the same rate, that channel is largely harvesting demand that already existed.
What incrementality testing means in marketing
Search terms like "incrementality testing meaning" and "what is incrementality testing in marketing" usually point at the same confusion: teams expect a new reporting metric and get a research method instead. A platform-reported ROAS is a description of observed behaviour. An incrementality test result is a causal claim with an error bar attached. Those are different objects, and treating the first as if it were the second is how budgets stay misallocated for years.
Why incrementality testing matters now
Signal loss, consent frameworks and walled gardens have made user-level tracking progressively less complete, which weakens every measurement method that depends on following individuals across sites and devices. Experiments do not depend on that chain. They depend on assignment, and assignment still works when tracking does not.
Adoption has followed. In a July 2025 survey of 196 US marketing professionals, fielded by EMARKETER with TransUnion, 52.0% of brand and agency marketers said they already use incrementality testing or experiments to measure campaigns, and 36.2% planned to invest in it over the following twelve months. The barriers they named were concerns about accuracy or reliability (44%), difficulty applying tests across ad types and retailers (43%) and limited tooling (41%).
The alternative methods are also less dependable than they look. Using data from twelve US advertising lift studies at Facebook covering 435 million user-study observations and 1.4 billion impressions, Gordon, Zettelmeyer, Bhargava and Chapsky (Marketing Science, 2019) found that common observational methods, including exposed versus unexposed comparisons, matching, model-based adjustment, synthetic matched-market tests and before-after comparisons, often fail to reproduce the results of a true randomised experiment. That held even after conditioning on thousands of behavioural variables and using non-linear models. Adding more covariates does not rescue a design that lacks randomisation.
How incrementality testing works
Every valid test needs three things, and most failed tests are missing one of them.
- A unit of randomisation you can actually control: a user, a device cluster, an audience segment, or a geographic market.
- A credible counterfactual: either a randomly assigned holdout, or a statistically constructed control such as a synthetic control built from untreated markets.
- Enough signal for the effect you care about to be distinguishable from ordinary week-to-week noise.
The mechanics are then simple. Freeze everything else, run the campaign in treatment only, and measure the difference across the test window. The reported result should always be a range, not a single number.
The main types of incrementality test
Platform conversion lift tests
The ad platform randomises its own users into exposed and held-out groups, then reports the difference. These are user-level tests, they are cheap to run, and they measure the platform's own media only.
Geo lift tests
Whole geographic markets become the unit of randomisation. Some markets get the campaign, others do not, and the difference is measured at market level. Geo lift testing needs no user-level identifiers, which is why it has become the default design for privacy-constrained and offline-heavy businesses.
Audience holdouts
A fixed share of a CRM list, app audience or retargeting pool is suppressed for a period. This suits owned channels, email, push and retargeting, where you control suppression directly.
Time-series on and off tests
A channel is switched off for a defined period and back on again, and the swing is measured against a modelled baseline. It is the weakest of the four designs because time confounds everything, but it is sometimes the only option available to a single-market advertiser.
| Design | Randomisation unit | Best for | Main weakness |
|---|---|---|---|
| Platform conversion lift | Individual user | Paid social, YouTube, display | Covers that platform only, and the platform runs it |
| Geo lift test | Geographic market | Cross-channel, offline, CTV, retail media | Needs many comparable markets and larger spend |
| Audience holdout | Segment or list | CRM, retargeting, app push | Contamination from other channels |
| On and off time series | Time period | Single-market advertisers | Seasonality and trend confound the result |
Geo lift testing: a method, not a label
Most guides treat "geo lift test" as a category name. It is really a family of estimation methods, and the one you choose determines what the result means.
- Difference-in-differences compares the before and after change in test markets against the same change in control markets. It is transparent, and it assumes the two groups would have moved in parallel without the campaign.
- Matched-market testing pairs each test market with a similar control market on size, seasonality and category behaviour, then compares the pairs.
- Synthetic control builds a weighted composite of untreated markets that reproduces the test markets' pre-period behaviour, and uses that composite as the counterfactual. It is the most flexible of the three and the most demanding on pre-period data.
Four design details decide whether a geo lift test is worth reading.
- Pre-period fit quality. If the counterfactual does not reproduce the test markets' history closely, nothing after the start date is interpretable. This is the most useful single diagnostic to demand in a read-out.
- Market selection. Markets should be chosen by an explicit procedure, not by convenience. Meta's open-source GeoLift R package, an MIT-licensed implementation of synthetic control methods, covers power analysis and data-driven market selection as part of the workflow.
- Spillover between markets. People travel and media leaks across boundaries. Google's geo-based Conversion Lift uses Google Marketing Areas, sub-country regions built with a spectral clustering algorithm specifically to serve as experimental units, and applies a contamination model that estimates how often people move between them, because a user exposed in one region who converts in another shrinks the measured difference.
- Feasibility before launch. Google reports geo test feasibility at three levels, High, Medium and Low, and recommends increasing budget to reach High. A Low-feasibility test is a budget commitment with a predictable answer of "we cannot tell".
Google incrementality testing and Meta incrementality testing
Both large platforms offer first-party experiments, and both have real limits worth understanding before you rely on them.
Google Ads Conversion Lift comes in two designs: user-based, where test groups are built from aggregated user attributes such as age and gender, and geography-based, which can support offline data. User-based studies report incremental conversions, relative conversion lift, incremental conversion value, incremental cost per action and incremental ROAS. Access is not universal. Google states that Conversion Lift is not available for all Google Ads accounts and that advertisers must contact their Google account representative to use it.
Meta runs user-level conversion lift studies inside its own tooling and also publishes GeoLift for market-level work, which means a Meta incrementality test can be run either on the platform's terms or on yours.
The platform marks its own homework
The obvious objection to platform incrementality testing is that the seller designs the experiment, chooses the holdout and reports the result on its own inventory. The evidence here is more nuanced than the cynicism suggests. An analysis of 3,204 Meta Lift tests and 181,890 Meta A/B tests by Burtch, Moakler, Gordon, Zhang and Hill (2025) found no meaningful audience imbalance between treatment and control in the Lift tests, which supports their validity for causal inference. The A/B tests were a different story. They showed clear imbalance caused by divergent delivery, where the delivery algorithm sends different ad variants to different audience segments. Campaign configuration choices can reduce that imbalance but not eliminate it.
The practical reading is straightforward. Platform lift tests are a legitimate instrument. Platform A/B tests are not a substitute for one, and creative or bid A/B tests should never be reported as incrementality evidence.
How to calculate incrementality
Three formulas cover almost every read-out.
- Incremental conversions = actual conversions in test, minus counterfactual conversions
- Relative lift = incremental conversions, divided by counterfactual conversions
- iROAS = incremental revenue, divided by incremental media spend
Watch the denominator. Some tools express lift over the baseline (incremental divided by counterfactual) and others over the treated total (incremental divided by actual). The same experiment then produces two different percentages, so state which convention you are using.
A worked illustration, using round numbers rather than benchmark figures:
| Line | Value |
|---|---|
| Conversions in test markets, four weeks | 12,400 |
| Counterfactual conversions from the synthetic control | 11,000 |
| Incremental conversions | 1,400 |
| Relative lift over baseline | 12.7% |
| Average order value | $95 |
| Incremental revenue | $133,000 |
| Incremental media spend in test markets | $110,000 |
| iROAS | 1.21 |
If the platform had reported a 4.2 ROAS on that same spend, the distance between 4.2 and 1.21 is the part of the reported return that was always going to happen. Closing that distance is the entire commercial case for incrementality testing, and it is why the finance team usually becomes the method's strongest internal advocate.
How to run an incrementality test, step by step
- Write the decision first. State what you will do differently at each possible outcome. If no result changes a decision, do not run the test.
- Define one KPI and one variable. Multiple changes inside one cell make the result uninterpretable.
- Set the minimum detectable effect. Decide the smallest lift that would change the decision from step one, then check whether the design can detect it.
- Run the power analysis before committing budget, not after the result disappoints.
- Select markets or audiences by procedure, and record the pre-period fit.
- Pre-register the analysis: test window, primary metric, estimation method, confidence level and stopping rule, all written down before launch.
- Execute with a change freeze. No creative refreshes, no bid changes, no promotions running in test markets only.
- Read out with an interval, then feed the result into the measurement system rather than filing it as a slide.
Related product
MMM Singularity
An interpretation layer connecting attribution, incrementality, saturation and prediction into one explainable view.
Statistical power: the question most guides skip
Every guide says "right-size your groups" and stops there. The honest version is less comfortable.
Required sample size scales with the square of the outcome's variability and inversely with the square of the effect you want to detect. Halving the effect you need to detect roughly quadruples the sample you need. Advertising is a hard case for this arithmetic because individual purchase behaviour is far more variable than advertising's effect on it.
Lewis and Rao (Quarterly Journal of Economics, 2015) quantified the problem across twenty-five large field experiments with US retailers and brokerages, collectively representing $2.8 million of digital advertising expenditure. The median confidence interval on advertising ROI was over 100 percentage points wide. They note that a coefficient of variation of 10 is common for individual-level sales relative to per-capita ad cost, and that an informative advertising experiment can easily require more than 10 million person-weeks. For many advertisers, a properly powered test of a single channel is not merely expensive, it is infeasible.
That result should change how you plan rather than discourage you from testing. Three consequences follow.
- Test at the level where the effect is large enough to see: a whole channel rather than one campaign, a full-funnel change rather than one creative variant.
- Prefer fewer, larger, longer tests over a calendar full of underpowered ones.
- Treat a wide interval as a real finding about your measurement capacity, and plan around it instead of reporting the midpoint as though it were precise.
When not to run an incrementality test
No page in this category tells you when to walk away, so here is the gate we use.
Do not run the test when any of these hold.
- The minimum detectable effect your volume supports is larger than the effect that would change your decision.
- The sales cycle is longer than the test window you can afford, so most of the response lands after read-out.
- The channel cannot be cleanly suppressed, so the holdout will be contaminated by the same message arriving another way.
- Spend on the channel is small enough that the cost of the holdout exceeds the value of the answer.
- A major promotion, seasonal peak or brand campaign overlaps the window.
When the gate closes, the alternatives are a marketing mix model to estimate channel contribution from aggregate data, a larger test across a bundle of channels rather than one, or a longer observation period with a pre-registered analysis. Reaching for an underpowered test because it feels rigorous is worse than not testing, because it produces a number that looks like evidence.
What a holdout actually costs
Suppression has a price, and it belongs in the business case as a line item rather than as a footnote. The arithmetic is:
Holdout cost = (share of audience or markets suppressed) x (baseline revenue in that population) x (test duration) x (the incremental rate you expect to find)
Run that calculation before the test, not after. If the expected cost of suppression is larger than the budget you would reallocate on the strength of the answer, the test has failed its own business case. If the expected cost is trivial, you can probably afford a bigger holdout and a tighter interval.
How to read the result without fooling yourself
- Report the interval, not the point. An iROAS of 1.21 with an interval spanning 0.4 to 2.0 supports a very different decision from the same 1.21 with an interval of 1.1 to 1.3.
- Check the confidence level. Platform lift tests commonly default to 90% rather than 95%, which makes a result look more decisive than a stricter threshold would allow. Confirm the setting before comparing results across sources.
- A null result is not proof of zero effect. It usually means the test could not distinguish the effect from noise. Report the minimum detectable effect alongside the null so the reader knows what was actually ruled out.
- Watch multiple comparisons. Run enough cells and some will clear significance by chance. Fix the primary metric and the number of cells in advance.
Incrementality testing vs attribution vs MMM
| Incrementality testing | Multi-touch attribution | Marketing mix modelling | |
|---|---|---|---|
| Question answered | What did this media cause? | Which touchpoints preceded conversion? | How does spend across all channels relate to outcomes? |
| Evidence type | Experimental | Observational, user-level | Observational, aggregate |
| Granularity | One channel or tactic per test | Campaign and creative level | Channel level, full portfolio |
| Coverage | Only what you test | Only trackable digital paths | Everything, including offline and brand |
| Cadence | Episodic | Continuous | Weekly or monthly refresh |
| Main limitation | Cost, power, contamination | No counterfactual | Needs external validation |
Everyone ends this comparison on the word triangulation. Very few explain what triangulation actually involves, which is the next section.
Closing the loop: using tests to calibrate your MMM
An incrementality test gives a precise answer about one channel over one window. A marketing mix model gives a broad answer about every channel continuously. Neither is complete on its own. The connection between them is not a slide that says "we use both", it is a calibration step.
In practice the test result enters the model in one of three ways.
- As an informative prior. In a Bayesian MMM, the experiment's effect estimate and its interval become the prior distribution on that channel's coefficient, so the model is pulled towards the experimental evidence in proportion to how precise that evidence is.
- As a constraint. The channel's implied ROI is bounded to the experimentally supported range, which stops the optimiser proposing spend levels the experiment has already ruled out.
- As a validation check. The model is fitted independently, then its ROI estimate for the tested channel is compared against the test result. Agreement is evidence for the whole model, not just that one coefficient. Disagreement is a diagnostic worth chasing before anyone acts on the allocation.
This is the layer most measurement stacks are missing. The test lives in one deck, the model lives in another, and the two never meet. MMM Singularity exists to close that gap. It sits on top of MMM outputs and connects attribution, incrementality, saturation and prediction into one explainable view, with the EMMMY agent handling the interpretation work that normally consumes an analyst's week. Plans start at $99 per month. Once that reconciled view exists, Aryma Nebula turns it into an actual spend plan rather than a set of recommendations nobody actions.
Institutional memory matters as much as the maths here. Test results tend to be forgotten within two planning cycles, which is why the same channel gets re-tested every eighteen months. MMM Synapse holds the history of past models, decks and read-outs so the next question starts from what you already learned.
Incrementality testing when campaigns are AI-driven
Most guidance on this topic was written when you could switch one channel off cleanly. Performance Max, Advantage+ and Demand Gen changed that. Budget moves between placements automatically, audiences are assembled by the system rather than by you, and suppressing "one channel" often means suppressing something the platform is free to redefine mid-test.
Three adjustments follow.
- Prefer geo designs. Geography is one of the few boundaries an automated campaign still respects, which makes geo lift testing the most reliable option for AI-driven campaign types.
- Test the bundle, not the placement. If the system reallocates internally, the honest unit of analysis is the whole campaign type, not a placement inside it.
- Lengthen the pre-period. Automated bidding drifts. A longer pre-period gives the counterfactual something stable to fit against.
This is also why we argue that AI belongs around the statistical core rather than inside it, a position we set out in Peripheral Agentic MMM. Agents are excellent at preparing, summarising and interrogating measurement work. Deciding what a confidence interval justifies is still a human call.
A pre-registration checklist you can reuse
Write these down before launch and the read-out argues itself.
- Hypothesis, stated as a directional claim about one channel or tactic.
- The decision that each possible outcome triggers.
- Primary KPI, and the single secondary metric you will also report.
- Design, unit of randomisation, and estimation method.
- Test and control definition, with the selection procedure recorded.
- Minimum detectable effect, and the power calculation behind it.
- Test window, including a stated no-peeking rule.
- Confidence level, and the lift convention you will report.
- Expected holdout cost.
- Where the result will be applied: which model, which coefficient, which review.
Frequently asked questions
What is a geo lift test?
A geo lift test is an incrementality test that uses geographic markets as the unit of randomisation. Some markets receive the campaign, comparable markets do not, and the difference in outcomes estimates the causal effect. It needs no user-level tracking, which makes it suitable for offline conversions, CTV, retail media and any market where identifiers are limited. Its accuracy depends on how well the control reproduces the test markets' pre-period behaviour.
How long should an incrementality test run?
Long enough to cover your conversion lag, plus enough periods for the estimator to separate the effect from normal variation. Four to six weeks is a common window for fast-cycle ecommerce, and considerably longer for considered purchases. Set the duration from the power analysis and the sales cycle, not from the reporting calendar, and fix it in advance so the test cannot be stopped early when the numbers happen to look favourable.
How big does the control or holdout group need to be?
There is no universal percentage. The right size falls out of the power analysis: the smaller the effect you need to detect and the noisier the outcome, the larger the holdout has to be. Work backwards instead. Decide the minimum detectable effect that would change a decision, calculate the holdout that supports it, then price the suppression. If that cost exceeds the value of the answer, redesign the test.
What is incremental ROAS and how is it different from ROAS?
ROAS divides attributed revenue by spend, and attributed revenue includes purchases that would have happened without the ad. Incremental ROAS divides only the experimentally measured incremental revenue by the incremental spend that produced it. iROAS is therefore almost always lower than platform-reported ROAS, and it is the version that survives a finance review, because it reflects revenue the business would not otherwise have earned.
How is incrementality testing different from attribution?
Attribution observes journeys that already ended in conversion and distributes credit across the touchpoints it can see. It has no counterfactual, so it cannot tell you what would have happened without the media. Incrementality testing creates that counterfactual deliberately through a control group. Attribution is useful for operational optimisation within a channel. Incrementality testing is what you use to decide whether the channel deserves the budget at all.
Can incrementality testing be automated?
Parts of it, safely. Market selection, power analysis, data preparation, the read-out and the write-up are all mechanical enough to automate, and both Google and Meta ship tooling for pieces of that workflow. The judgement calls are not: choosing what to test, deciding whether the pre-period fit is good enough, and deciding what a wide interval justifies commercially. Automate the work around the experiment and keep the design decisions human.
How we approach it at Aryma
Causal marketing experiments are the core of what Aryma Labs does, and incrementality testing is the instrument we trust most when a budget decision has to hold up under scrutiny. The discipline is not in running more tests. It is in running the few that are properly powered, reading them honestly, and making sure each result changes the model that drives the next allocation instead of sitting in an archive.
That is the job MMM Singularity was built for: an interpretation layer that keeps attribution, incrementality, saturation and prediction in one explainable view, so a test result becomes a permanent input to your measurement rather than a one-off finding. If your measurement problem needs a design rather than a product, our custom solutions team builds the experiment and the calibration around it.
For help with MMM, causal marketing experiments and incrementality testing, get in touch with us.
Aryma Labs is a marketing mix modeling consultancy founded in 2019. Aryma AI is its Gen AI division, applying agents to the periphery of MMM while keeping the statistical core human-led.
Explore the Aryma AI suite
Gen AI products for marketing mix modeling, built on a human-led statistical core. Explore the suite, or talk to the team.
More from Aryma
Keep reading
Peripheral Agentic MMM
A deep dive into how AI is reshaping Marketing Mix Modeling (MMM), yet not replacing the foundation of MMM
Why LLM Observability Matters: How We Measure Every Query Inside MMMGPT
How Aryma Labs uses Langfuse to trace every MMMGPT query end to end, turning AI debugging from guesswork into evidence.
The Art of Subtraction: Training AI Agents in Marketing Mix Modeling with 'Via Negativa'
How exclusionary prompts and restraints are reshaping the future of smarter, more nuanced MMM AI - lessons from Aryma Labs