Every marketer who has run a lift study or an incrementality test eventually runs into the same uncomfortable question. How big should the holdout group actually be? It sounds like a small technical detail, but the size of your holdout group quietly shapes almost everything about your test, from how much budget you’re effectively giving up, to how confident you can be in the results, to how quickly you’ll get a usable answer.
Get it wrong, and you either end up with a test that can’t tell you anything meaningful, or you sacrifice more revenue than necessary just to prove something you could have proven with a smaller sacrifice. This article walks through exactly how holdout groups work, why they create a genuine tradeoff rather than a simple checkbox decision, and how to build a practical framework for deciding how much budget to hold back in any given test.
If you’re just getting familiar with lift testing in general, our Complete Guide to Conversion Lift Measurement is worth reading first, since this article assumes some familiarity with the basics of incrementality testing.
Also Read: Incremental vs. Attributed Conversions
What a Holdout Group Actually Is
A holdout group is a portion of your eligible audience that is deliberately excluded from seeing a campaign, so it can serve as a control group for comparison against the exposed group that does see your ads. The entire logic of incrementality testing depends on this comparison. Without a holdout group, you have no baseline to measure against, and you’re left guessing whether conversions happened because of your marketing or simply because those customers were always going to convert.
Think of it like a clinical trial. Researchers don’t just give everyone the new medication and declare it works if people get better. They hold back a control group that doesn’t receive the treatment, so they can compare outcomes and isolate the actual effect of the drug from natural recovery, placebo effect, and random variation. A holdout group in marketing plays exactly the same role. It’s the only way to know what would have happened without your campaign.
Why Holdout Groups Create a Real Tradeoff
Here’s where things get genuinely difficult, and where a lot of marketers make decisions that hurt them later. Every person placed in a holdout group is a potential customer who doesn’t get exposed to your marketing during the test period. If your campaign is actually effective, that means real, measurable revenue is being deliberately left on the table for the sake of getting a clean read.
This creates a direct tension between two things marketers care about deeply. On one side is statistical confidence, since a larger holdout group generally produces a cleaner, more reliable comparison, especially in accounts with lower conversion volume. On the other side is revenue protection, since a larger holdout group means more people who don’t get marketed to, which translates directly into forgone conversions and, if the campaign truly works, real lost revenue during the test window.
This is the budget tradeoff framework at its core. You are never choosing between a “good” holdout size and a “bad” one in the abstract. You’re choosing a specific point along a spectrum, and every point on that spectrum trades some amount of statistical reliability for some amount of protected revenue, or vice versa.
The Two Ends of the Spectrum
It helps to think about this as a genuine spectrum rather than a single correct answer, because the right holdout size depends heavily on your specific situation.
A Small Holdout Group
At the small end of the spectrum, you might hold back just 5 to 10 percent of your eligible audience. The appeal here is obvious. Almost everyone still gets exposed to your marketing, which means you protect the vast majority of expected revenue during the test period.
The cost is statistical. A small holdout group means a small control sample, and small samples are noisier. Random week to week fluctuations in conversion behavior can easily be mistaken for a real lift effect, or a real lift effect can get buried in the noise and appear statistically insignificant even when it’s genuinely there. In accounts with low overall conversion volume, an undersized holdout group is one of the most common reasons a lift study comes back inconclusive.
A Large Holdout Group
At the other end, some tests use holdout groups as large as 20 to 50 percent of the eligible audience. This produces a much more statistically robust comparison, since the control group has enough volume to smooth out random noise and give a clearer signal of the true underlying effect.
The cost here is direct and often underestimated. If your campaign is genuinely effective and you’re holding back half your audience from seeing it, you’re potentially forgoing a substantial number of conversions for the duration of the test, purely in the name of measurement precision. For a large advertiser with significant baseline revenue, this can represent a meaningful dollar amount, even over a relatively short test window.
Neither end of this spectrum is inherently right or wrong. The correct choice depends on what you’re optimizing for and what your account can tolerate.
Building the Budget Tradeoff Framework
Rather than picking a holdout size arbitrarily or defaulting to whatever a platform suggests without question, it helps to work through a structured framework that weighs the specific factors relevant to your situation.
Step 1: Estimate Your Conversion Volume
Start by looking at your historical conversion volume for the audience and time period you plan to test. Accounts with high conversion volume can generally use a smaller holdout percentage and still reach statistical significance, because there’s enough raw data flowing through both groups to detect a real effect. Accounts with low conversion volume, such as those selling high consideration or high price products, often need a larger holdout percentage simply to accumulate enough control group conversions to make a valid comparison.
Step 2: Estimate Your Expected Lift Size
Consider how large an effect you actually expect the campaign to produce. Detecting a large lift, such as a 20 or 30 percent increase in conversions, requires much less statistical power than detecting a subtle lift of just a few percentage points. If you have reason to believe your campaign will produce a large, obvious effect, you can often get away with a smaller holdout group. If you’re testing something with a more modest expected impact, you’ll likely need a larger holdout to detect it reliably.
Step 3: Calculate the Real Cost of the Holdout
This is the step most marketers skip, and it’s the one that makes the framework actually useful. Take your average conversion value and your expected conversion rate, and calculate what a given holdout percentage actually costs in forgone revenue over the planned test duration. A 20 percent holdout on a campaign with strong existing performance might represent a genuinely significant dollar figure, and seeing that number in concrete terms often changes how comfortable a team feels with a particular holdout size.
Step 4: Weigh Confidence Against Cost
Once you know both the statistical requirement and the dollar cost, you can make an informed tradeoff decision rather than an arbitrary one. If the cost of a larger holdout is modest relative to your overall marketing budget and the campaign in question is a strategic priority you need to evaluate rigorously, leaning toward a larger holdout for better statistical confidence often makes sense. If the cost is substantial and the campaign is more routine or lower stakes, a smaller holdout paired with a longer test duration can sometimes achieve similar statistical reliability at a lower revenue cost, since duration and holdout size both feed into the same underlying power calculation.
Step 5: Set a Minimum Viable Confidence Level
Decide in advance what level of statistical confidence you actually need before the results will be considered actionable. If you’re making a major budget reallocation decision based on this test, you likely want a higher confidence threshold, which may justify a larger holdout despite the cost. If you’re running a lighter, more exploratory test just to get directional signal, a lower confidence threshold paired with a smaller holdout may be perfectly acceptable.
Common Mistakes in Holdout Sizing
A few patterns show up repeatedly when marketers get this wrong, and it’s worth naming them directly.
Defaulting to whatever the platform suggests without question. Automated recommendations are a reasonable starting point, but they’re generally optimized for statistical power alone, not for your specific revenue tolerance or business priorities.
Treating holdout size as a one time decision. The right holdout size can change over time as conversion volume grows, as expected lift shrinks due to campaign maturity, or as the strategic importance of a given test changes.
Ignoring the compounding cost of holdouts across simultaneous tests. If multiple campaigns are running holdout tests at the same time, the combined revenue impact can be larger than any single test suggests in isolation, and this is easy to miss when each test is evaluated separately.
Under sizing the holdout to protect revenue, then being surprised by inconclusive results. This is the most common mistake of all. Teams understandably want to protect as much revenue as possible, shrink the holdout accordingly, and then are disappointed when the test comes back statistically insignificant, having sacrificed some revenue for a result that ultimately can’t be trusted either way.
Putting the Framework Into Practice
A practical way to apply this framework is to run the cost and power calculations side by side before launching any test, rather than choosing a holdout size first and checking the implications afterward. Lay out two or three candidate holdout percentages, calculate the expected forgone revenue for each, and check the corresponding statistical power for each against your account’s conversion volume and expected lift size. This turns an abstract tradeoff into a concrete comparison your team can actually discuss and agree on, rather than a number picked somewhat arbitrarily under time pressure.
It’s also worth revisiting holdout sizing decisions after each test cycle. If a test consistently comes back underpowered, that’s a signal the holdout was too small relative to the account’s conversion volume. If a test consistently protects far more revenue than necessary while still reaching strong statistical confidence, that’s a signal the holdout could be trimmed in future tests without sacrificing reliability.
The Bottom Line
Holdout groups are not a minor configuration setting buried in a testing interface. They sit at the center of a real tradeoff between how confidently you can trust your results and how much revenue you’re willing to set aside to get there. There’s no universally correct holdout percentage, only the right percentage for your specific account, your conversion volume, your expected lift size, and how much statistical confidence a given decision actually requires. Building a deliberate framework around this tradeoff, rather than defaulting to a platform suggestion or an arbitrary round number, turns holdout sizing from a guessing game into a decision your team can defend and refine over time.
For more on how holdout based testing fits into the broader landscape of incrementality measurement, revisit our Complete Guide to Conversion Lift Measurement, which covers the full framework this article builds on.