If you’ve ever set up a conversion lift experiment in Google Ads and seen a small number sitting next to a label called Study Power, you’ve probably wondered what it actually means and why Google seems so insistent that you push it higher. It’s not a vanity metric. Study Power is arguably the single most misunderstood mechanic in lift testing, and understanding it is the difference between running a test that gives you a real answer and running one that just wastes budget.
This article breaks down exactly what Study Power calculates, how the 50 to 95 percent certainty scale works, which levers actually move that number, and what to do when your account comes back underpowered. By the end, you’ll know how to read Google’s automated budget guidance instead of just clicking accept and hoping for the best.
If you’re new to lift measurement in general, our Complete Guide to Conversion Lift Measurement is a good starting point before diving into this more tactical breakdown.
Also Read: Geo-Based vs. User-Based Conversion Lift
What Is Study Power, Actually?
Study Power is Google Ads’ feasibility calculator for conversion lift studies. Before you launch an experiment, Google runs a statistical simulation using your account’s historical conversion data to estimate the likelihood that your test will detect a real lift effect if one actually exists. That likelihood is expressed as a percentage, and it’s what shows up on screen as your Study Power score.
In plain terms, Study Power answers this question: if your ads truly are driving incremental conversions, how likely is this specific test setup to actually prove it with statistical confidence? A low Study Power score doesn’t mean your ads aren’t working. It means your test, as currently configured, doesn’t have enough data or a large enough sample to reliably tell the difference between a real effect and random noise.
This is a concept borrowed directly from clinical trial design and academic research, where it’s called statistical power. Google has simply adapted it for advertising experiments and given it a friendlier name.
The 50 to 95 Percent Certainty Scale
Google Ads presents Study Power on a scale that generally runs from around 50 percent up to 95 percent, and understanding what sits at each end of that range matters more than most advertisers realize.
Around 50 percent is essentially a coin flip. At this level, your test has roughly even odds of detecting a genuine lift effect versus missing it entirely, even if your campaign is truly driving incremental results. Launching a study at this power level is a bit like flipping a coin to decide whether your marketing worked. You might get a clean read, but you’re just as likely to get a false negative that tells you your ads aren’t working when they actually are.
Around 80 percent is the conventional minimum threshold used across most fields of applied statistics, including marketing measurement. At this level, you have a reasonably strong chance of detecting a real effect, though there’s still meaningful room for error.
90 percent and above is where Google Ads starts to consider a study reliably conclusive. At this level, the risk of a false negative, meaning missing a real lift effect that’s actually there, drops substantially. This is the range most experienced advertisers target before greenlighting a study, and it’s the benchmark referenced in the very concept of getting your lift study to 90 percent certainty.
95 percent sits near the top of what Google typically displays, representing a very high degree of statistical confidence, though pushing this high often requires either a very large budget, a long duration, or both.
The key takeaway is that certainty of lift percentage isn’t just a number to glance at. It directly determines whether the results you eventually see are trustworthy or essentially meaningless.
What Actually Moves the Study Power Number
This is where most advertisers get stuck. They see a low Study Power score, increase the budget by a small amount, and get frustrated when the number barely moves. Understanding the actual levers behind the calculation saves a lot of wasted trial and error.
1. Budget
Budget is the most obvious lever, and it works because more spend generally means more conversions, which gives the statistical model more data points to work with. However, budget alone won’t always solve an underpowered study, especially in accounts with naturally low conversion volume or a low baseline conversion rate. If your account converts rarely, doubling budget might increase impressions without proportionally increasing the conversion events the model actually needs.
2. Holdout Size
The holdout group is the portion of your eligible audience that’s deliberately excluded from seeing your ads so it can serve as a control group for comparison. Holdout size optimization is a genuinely underused lever. A larger holdout group can sometimes improve the statistical cleanliness of your comparison, but shrinking it too far in the other direction, trying to maximize exposed traffic, often backfires because it leaves too few control conversions to compare against. Google typically recommends a default holdout percentage, and deviating from it should be done deliberately, not accidentally.
3. Conversion Volume
This is the single biggest driver of Study Power, and it’s also the one advertisers have the least direct control over in the short term. The model needs enough historical and projected conversion events in both the exposed and holdout groups to distinguish a real lift from ordinary week to week fluctuation. Low conversion volume accounts, such as those selling high consideration or high price products, will almost always show lower Study Power scores regardless of how much budget gets thrown at the problem.
4. Expected Lift Size
Counterintuitively, the size of the lift effect you’re trying to detect changes how much power you need. Detecting a small lift, say a 2 percent increase in conversions, requires far more data and statistical power than detecting a large lift, like a 20 percent increase. If your campaign is only expected to move the needle slightly, you’ll need a much bigger sample to prove it statistically, which is part of why some very effective but subtle campaigns still show weak Study Power scores.
5. Duration
Extending the length of the study gives the experiment more time to accumulate conversions in both groups, which directly feeds back into the conversion volume lever above. A longer duration is often the most practical fix for accounts that don’t have the option to dramatically increase budget, though it does mean waiting longer for a usable result.
What to Do When Your Account Is Underpowered
Seeing a low Study Power score isn’t a dead end, but it does mean you need to make a deliberate adjustment rather than launching the study as is. Here’s a practical sequence to work through.
Start by extending duration before increasing budget. Duration is often the cheapest lever to pull, and it doesn’t require committing additional spend upfront. If your account has decent conversion volume but the study window is short, stretching the test from two weeks to four or six weeks can meaningfully improve power without touching your budget allocation.
Reconsider your conversion event. If you’re measuring a deep funnel event like a completed purchase and your Study Power is low, check whether a higher funnel event, such as add to cart or a qualified lead form, has enough volume to serve as a more statistically reliable proxy metric, provided it still reflects genuine business value.
Broaden the campaign or audience scope. Sometimes a single narrow campaign simply doesn’t generate enough conversions to power a standalone study. Consolidating similar campaigns into one broader experiment can pool enough conversion volume to reach a usable power level.
Adjust holdout percentage carefully. If Google’s default holdout size seems unnecessarily large for your account’s volume, a modest reduction can sometimes help, but this should be a considered decision, not a blind attempt to inflate exposed traffic.
Increase budget as a later step, not a first resort. Once duration, conversion event selection, and campaign scope have been reviewed, increasing budget becomes a more targeted decision rather than a blunt instrument.
How to Read Google’s Automated Budget Guidance
When your Study Power falls below the 90 percent threshold, Google Ads will often surface an automated recommendation suggesting a specific budget increase needed to reach that target. It’s worth understanding what’s actually happening behind that number before accepting it.
Google’s recommendation is generated by running the same statistical power calculation in reverse, essentially asking what level of spend, given your account’s historical conversion rate and current holdout configuration, would be needed to push Study Power to the 90 percent mark. This is useful, but it shouldn’t be treated as gospel.
First, the suggested budget assumes your future conversion rate and cost per conversion will resemble recent historical performance. If your account is seasonal, or if you’re about to launch new creative or landing pages, that assumption may not hold, which means the actual power achieved could differ from the estimate.
Second, the recommendation optimizes purely for statistical power, not for cost efficiency or overall campaign strategy. Accepting a large suggested budget increase without evaluating whether that spend fits your broader marketing plan can lead to an experiment that’s statistically sound but financially reckless.
Third, treat the guidance as one input among several. If the suggested increase is modest and fits comfortably within your existing budget, it’s usually worth accepting. If it represents a dramatic jump, it’s worth first exploring the duration and campaign scope adjustments covered earlier, since those often close some of the power gap without requiring as large a spend commitment.
Bringing It All Together
Study Power exists to protect you from drawing false conclusions about whether your Google Ads campaigns are actually working. A test launched with weak statistical power is essentially a coin flip dressed up as an experiment, and its results, whether positive or negative, shouldn’t be trusted enough to inform real budget decisions.
Getting your lift study to 90 percent certainty is rarely about pulling one single lever. It’s usually a combination of extending duration, selecting the right conversion event, ensuring adequate conversion volume, and only then considering a budget increase, ideally with a clear eyed view of what Google’s automated recommendation is actually optimizing for. Treat Study Power as a planning tool rather than an obstacle, and your lift studies will start producing results you can actually act on with confidence.
For the broader framework this fits into, including how lift studies compare to other incrementality testing methods, revisit our Complete Guide to Conversion Lift Measurement.
One Response