FILED UNDER Consumer Commerce · Marketing · Measurement

Incrementality testing,
and what the answer
costs you.

Your attribution tools disagree because none of them measures causation. A holdout test does. Here is how to run one, what the learning costs in forgone revenue, and the volume you need before the answer means anything.

Author
Taylor Sicard
Published
September 2026
Read
7 min · ~1,650 words
Ring
I · Consumer Commerce
About the author
Taylor Sicard

Early Shopify employee who helped build and scale the Partner Program, co-founder of WIN Brands Group, and has built portfolios of consumer brands to mid nine figures in annual revenue, and founder of a Shopify-ecosystem SaaS company sold to Tiny. He advises DTC brands, Shopify app founders, and Fortune 500 commerce teams.

Full background →
Key takeaways

An incrementality test is the only method that answers whether your ads caused the sale. Attribution tools observe correlation and model the rest, which is why three of them can point at the same store and disagree. A holdout costs real money in forgone revenue, and most brands never price that before deciding not to run one.

  • Three designs are practical for a DTC brand: a geo holdout, a PSA or ghost-ad test, and a budget-split test. Only the geo holdout is fully self-serve.
  • The cost of the test is the baseline revenue in the holdout regions across the window, to the extent the channel was genuinely driving it. Calculate it before you argue about it.
  • Statistical power, not duration, is the binding constraint. A brand doing a handful of orders a day in the holdout cannot detect a realistic lift no matter how long it runs.
  • Run it when the spend at stake exceeds the cost of learning, and when you would actually change the budget based on the result.
  • Do not run one during a promotion, a launch, or a seasonal peak. You will measure the event, not the channel.
Source: Taylor Sicard, Taylor Sicard Consulting · Updated September 2026

Point three attribution tools at the same store and they will give you three different answers about the same Meta campaign. That is not a bug in one of them. It is what happens when every tool observes what people did and then models the part it cannot see. None of them ran an experiment, so none of them knows what would have happened if the ads had not run.

An incrementality test answers that question directly by turning the ads off for some people and leaving them on for others. It is the only method in the stack that measures causation rather than correlation. It is also the only one with a price tag most brands never calculate, which is the real reason it does not get run.

01/What it measures
PLATE 01 · CAUSATION

Attribution observes.
Incrementality
intervenes.

Attribution watches a customer touch an ad and later buy something, then assigns credit according to a rule: last click, first click, a 28-day window, a proprietary model. Every one of those rules is a guess about a counterfactual it never observed.

Incrementality removes the guess. You withhold the ads from a randomly chosen group, run the campaign to everyone else, and compare. The difference is the lift the advertising actually caused. Everything else is arithmetic on top of an assumption.

This matters most where attribution is most flattering. Branded search, retargeting, and anything reaching people already deep in your funnel will look excellent in a click-based model and can turn out to be close to worthless incrementally, because those customers were going to buy anyway. If a channel's reported ROAS seems too good, that is usually the channel to test first. What the tools actually do, and how far observational methods drift from experimental results, is covered in the attribution tools teardown.

02/The three designs
PLATE 02 · WHAT YOU CAN RUN

Three designs. Only
one is genuinely
self-serve.

Geo holdout. Split regions into test and control, turn the channel off in control, compare revenue. You control it entirely, it works across any channel including ones with no platform test tooling, and it does not need the ad platform's cooperation. The tradeoff is that regions are never perfectly matched, so you need enough of them on each side for the comparison to mean anything.

PSA or ghost-ad test. The platform shows a placebo, or logs who would have seen your ad, and compares those people against the exposed group. Cleaner randomisation than geo, because it happens at the user level. The catch is that you depend on the platform to run and report the test that grades the platform.

Budget-split test. Cut spend hard in one set of geos and hold it in another, then look at the difference. Easier to stomach politically than a full blackout because nothing goes dark, but a smaller signal, which means it needs more volume and more time to reach significance.

For most brands the geo holdout is the one to start with. It is the only design where you own the whole apparatus and nobody grades their own homework.

03/The price of the answer
PLATE 03 · FORGONE REVENUE

The test costs real
money. Price it before
you argue about it.

Every incrementality test buys information with revenue you deliberately do not earn. That is the whole trade, and it is why the test feels expensive even when it is cheap relative to what it protects.

The arithmetic fits on a napkin. Take your baseline revenue in the holdout regions for the length of the test, multiply by the share of that revenue the channel is genuinely driving, and that is roughly your exposure. If the channel turns out to be highly incremental, you lose most of that. If it turns out to be worthless, you lose almost nothing, which is itself the finding.

Set that number against the spend the decision governs. A brand arguing about whether to keep a meaningful monthly budget in a channel is deciding an annual number many times larger than a two-week holdout in a slice of the country. When the test costs a fraction of the spend it settles, refusing to run it is not caution. It is paying every month to avoid finding out.

The test is not expensive. Not knowing is expensive, and it renews monthly.

04/Power
PLATE 04 · THE BINDING CONSTRAINT

Volume, not duration,
decides whether the
answer means anything.

The most common way a holdout fails is not that it was run badly. It is that the brand never had enough conversions in the holdout to detect the effect it was looking for, so the test returns a range wide enough to include both "this channel is excellent" and "this channel does nothing." That is not a null result. It is no result, bought at full price.

Two things drive it. The smaller the true lift you are trying to detect, the more conversions you need, and the effects worth finding are often modest. And the noisier your baseline, the more ordinary week-to-week variance swamps the signal, which is why brands with lumpy, promotion-driven revenue need more volume than steady ones.

The check before you start: how many orders will land in the holdout during the window, and is that enough that a realistic lift would move the number visibly above normal weekly noise? If you cannot answer that, run the calculation first or do not run the test. A brand doing a handful of orders a day in the control region is buying an expensive shrug.

05/When not to
PLATE 05 · THE HONEST NO

When not to run
one at all.

During a promotion, launch, or seasonal peak. You will measure the event rather than the channel, and the result will not generalise to a normal month. Test in ordinary trading conditions or the number is decorative.

When you would not act on the result. If the channel stays funded whatever the test says, because of an internal commitment or a partner relationship, the test is theatre. Save the forgone revenue.

When the volume is not there. Covered above, and it is the most common case. Below a certain scale the right move is to accept that attribution is directionally useful and spend the effort on contribution margin instead, which you can actually measure. The mechanics are in contribution margin for DTC.

When you have not agreed the decision rule in advance. Write down what you will do at each outcome before you see the data. Teams that skip this reliably discover a reason the unfavourable result does not count, which wastes the test and the revenue that paid for it.

06/Questions
PLATE 06 · FAQ

Questions brands ask
about incrementality
testing.

Q: What is incrementality testing?

It is a controlled experiment that measures whether your advertising caused sales, rather than inferring it from what people clicked. You withhold a channel from a randomly selected group, run it to everyone else, and compare outcomes. The difference is incremental lift. Attribution tools cannot do this because they only observe behaviour that already happened and model the counterfactual they never saw, which is why three tools looking at the same store can report three different answers.

Q: How long should an incrementality test run?

Long enough to accumulate the conversions you need, which depends on your order volume and how large an effect you are trying to detect rather than on a fixed calendar length. Duration is the wrong variable to fix first. Work out how many orders will land in the holdout and whether a realistic lift would clear normal weekly noise, then set the window to reach that. Cover a whole number of weekly cycles so you are not comparing weekends against weekdays, and avoid promotions and seasonal peaks entirely.

Q: How much does an incrementality test cost?

There is no invoice. The cost is forgone revenue: the baseline revenue in your holdout regions across the test window, to the extent the channel was genuinely driving it. Calculate that figure before the debate rather than after, and set it against the annual spend the decision governs. In most cases where a brand is arguing about a channel at all, the test costs a small fraction of what continuing to guess costs over a year.

Q: Can I run incrementality testing without an expensive tool?

Yes. A geo holdout needs regional revenue data, the ability to turn a channel off by geography, and the discipline to leave it off for the full window. Every part of that is available in the ad platforms and your own order data. Paid platforms make the analysis faster and the region matching more rigorous, which matters at scale, but the constraint for most brands is order volume and organisational patience rather than software.

Q: What if the test comes back inconclusive?

Treat it as a power problem before you treat it as a finding. An inconclusive result usually means the holdout never contained enough conversions to detect the effect, which is a design failure rather than evidence the channel does nothing. Check whether the range of plausible results spans both a meaningful lift and zero. If it does, you learned nothing about the channel and something important about your testing capacity, and the fix is a larger holdout, a longer window, or accepting that you are below the scale where this method works.