Knowledge
Incrementality Testing: Even Without a Data Science Team
How much revenue would have come without the ads? Holdout tests, geo splits and clean test designs – measure incrementality without a data science team.
By Denys Lichtenstein Prefer us on Google

Platform attribution answers many questions: which creative gets clicks, which campaign delivers, what ROAS shows in the account. One question it does not answer – and it is the most important of all: how much revenue would have come even without the advertising? That is exactly what incrementality measures. The term sounds like a data science team, statistics infrastructure and big-tech methodology – yet the core idea can be tested with simple means. This guide shows why the platform figures systematically miss this question, which test designs work for mid-sized businesses, what makes a clean test and which mistakes devalue even good tests.
The question behind every budget decision
Every euro of ad spend is only well invested if it generates revenue that would not have existed otherwise. That sounds obvious, but it is violated constantly in everyday practice. An example: a loyal customer wants to reorder anyway, sees a retargeting ad shortly before the purchase and clicks on it. In the ad account, the full order value appears as a campaign success – yet nothing about it is incremental. The purchase would have happened either way; the ad merely collected the last click.
Multiply this case by thousands of orders and you get an ad account that shines while your business gains less than the numbers suggest. Whoever wants to steer budgets must therefore keep two quantities apart: attributed revenue and caused revenue. Attribution delivers the first. Incrementality asks about the second – and only that ultimately justifies the spend.
Why platform attribution does not answer the question
Platform attribution answers a different question: which conversions had contact with an ad within the configured attribution window? That is an assignment rule, not proof of impact. The systems attribute conversions to themselves generously – every interaction within the window counts, regardless of whether it caused the purchase or merely happened to precede it.
On top of that comes a structural problem: every platform evaluates its own performance. If you add up the self-reported revenues of all channels, the sum often exceeds your actual total revenue, because several systems claim the same order for themselves. Why an overarching metric like MER is therefore the more honest compass, we explain in the article MER instead of ROAS. And what attribution can still do today – and what it no longer can –, the sister article Attribution today puts into perspective.
Important: none of this makes attribution useless. For comparisons within a channel it remains a workable tool. But the question about additional revenue it cannot answer as a matter of principle. That requires experiments – and they are simpler than the term incrementality suggests.
Four test designs that work without a data science team
Big-tech companies employ entire teams with their own statistics infrastructure for incrementality measurement. You do not need that apparatus. For robust directional answers, four designs suffice, sorted from radically simple to methodologically more demanding:
1. Holdout test: pause and observe. You deliberately pause a channel – in a defined region or for a defined period – and observe what happens to total revenue in the shop backend. If it drops noticeably, the channel made a genuine contribution. If it stays largely stable, part of the attributed revenue was not incremental. The price of this test is potential lost revenue during the pause – which is why it pays off especially where you already distrust the reported ROAS, for instance with pure retargeting campaigns.
2. Geo split: compare regions. Instead of pausing completely, you serve comparable regions differently: in one, the channel keeps running, in the other it does not – or with a clearly reduced budget. If revenue in the advertised regions develops measurably differently from the unadvertised ones, you have a strong signal of genuine effect. What matters is that the regions performed comparably before the test – otherwise you compare apples with oranges and measure regional quirks instead of advertising impact.
3. Meta’s conversion lift tests: the platform option. From sufficient conversion volume upwards, you can use Meta’s own lift tools. They split users into groups: one part sees your ads, a randomly selected part does not – the difference in purchasing behaviour is the measured lift. Methodologically, this is the cleanest design on this list. Two caveats: you need enough volume so the results do not drown in noise, and the test only answers the question for Meta itself. As a reality check for your biggest channel, it is still valuable.
4. On/off comparison with an honest baseline. The most pragmatic variant: switch a channel off for a period, then back on – and hold the revenue development against a baseline defined beforehand. The word beforehand is decisive. Determine what revenue level you expect without the channel – based on the preceding weeks, the trend and the typical seasonal pattern – before you switch off. Without this discipline, you end up comparing against wishful thinking and read into the result exactly what you want to see.
What makes a clean test
All four designs stand or fall by the same ground rules. Four principles separate an experiment from an anecdote:
One channel at a time. If you simultaneously pause Meta, restructure the Google budget and run a newsletter push, you cannot assign a cause to any effect afterwards. Change one variable and keep the rest as stable as possible.
A sufficiently long period. Advertising has a lagging effect: after switching off, your revenue lives off earlier contacts for a while, and depending on the product, days to weeks lie between first contact and purchase. A test must run clearly longer than your typical purchase cycle – otherwise you measure the echo of old campaigns instead of the real effect.
Factor out seasonality. Never compare blindly against the previous week. Use the pattern of comparable earlier periods, control regions running in parallel, or at least the trend of the weeks before the test to separate seasonal patterns from the test effect. A revenue decline that would have come even without the pause proves nothing.
Hypothesis and decision rule in advance. Before the start, write down what you expect and what you will do for each outcome – for example: if revenue during the pause falls clearly less than the paused spend would suggest, budget gets reallocated. A test without a decision rule is not measurement, it is busywork.
Typical mistakes that devalue tests
Most failed incrementality tests do not fail because of statistics, but because of four avoidable patterns:
Tested too short. A few days of pause almost always show: nothing happens. That is not because the advertising is ineffective, but because purchase cycles and lagging ad effects delay the impact. Whoever stops too early draws the wrong conclusion from an unfinished experiment.
Tested too small. A test in a tiny region or with a minimal share of budget disappears in the normal ups and downs of the business. The effect must be large enough to stand out from the daily noise – better one bold test with a clear answer than three homeopathic ones without.
Tested at the wrong time. Sale phases, product launches and peak season distort every result. Discounts pull purchases forward, launches create one-off effects, and in high season demand behaves differently from the rest of the year. Test in calm, representative phases.
Result does not fit – so it gets ignored. The most expensive mistake is not methodological but human: the test shows that a beloved channel contributes less than assumed, and the result gets argued away. Whoever only accepts results that confirm gut feeling does not need tests – they need more honest decision-making processes.
Rough and honest beats precise and wrong
None of these tests would hold up to the standards of a scientific experiment – and they do not have to. The alternative is not the perfect experiment; the alternative is blind trust in numbers from platforms that evaluate their own performance. A rough, honest answer to the question of a channel’s real contribution is worth more for budget decisions than a precise-looking illusion with two decimal places.
Besides, you do not test incrementality permanently but selectively: when a channel is about to be scaled significantly, when a ROAS looks too good to be true, or when a retargeting budget must justify why it exists at this level. One or two clean tests a year often create more clarity than twelve months of reporting.
Conclusion
The question of how much revenue would have come even without advertising is uncomfortable – and precisely for that reason so valuable. Platform attribution cannot answer it, simple experiments can: holdout tests, geo splits, Meta’s lift tools from sufficient volume upwards, and on/off comparisons with an honest baseline. What you need for this is not a data science team but discipline – one channel at a time, enough runtime, seasonality in view and a decision rule fixed before the test. If you want to know where your setup stands and which budgets deserve an incrementality check first, take a look at our free account check.
