Fast, Confident, and Wrong: The Risk of Noisy Incrementality Tests

We simulated a year of marketing decisions 36 million times. Accuracy protected the business. Volume didn't.

Jul 23, 2026

"Any incrementality is better than no incrementality."

I hear some version of this almost every week, and it sounds reasonable: If you can't measure everything perfectly, start with whatever solution is available, get a directional read, keep testing. Enough results should eventually point you toward the truth.

But there's a step hidden inside that logic: acting on the results.

An experiment doesn't create value because it produces a lift estimate. It creates value when that estimate leads to a better decision – scale the channel, pull back, hold, retest. If the signal behind those decisions is noisy, doing more of them doesn't average you toward the truth. It means making more wrong turns, faster.

I've watched this play out anecdotally for years: rigorous measurement programs pulling away and brands that cut corners on their testing practice paying the price (often literally). But anecdotes aren't numbers. So I built a simulation that modeled a full year of experiment-driven budget decisions, varying only the two levers a team actually controls: how precise their measurement is, and how often they test.

36 million simulation runs later, here's what the data said.

The short version

  • Precise measurement is downside protection. The precise scenarios beat a frozen-budget baseline 82% of the time. The noisy ones, only 62%. Put differently: acting on noisy results left the business worse off than doing nothing 38% of the time, more than twice as often as the precise approaches. 
  • You can't test your way out of noisy signals. Brands running noisy experiments finished ahead no more often at 15 tests a year than at 6 (61.8% vs. 62.3%). More volume just meant bigger swings in both directions.
  • A strong experimentation program is worth double digits. The precise, high-cadence approach delivered a 13.2% average revenue lift with a 20% limit on budget shifts, more than double the noisy approach's payoff (5.6%) at the same cadence and limit.
  • A tight confidence interval can be manufactured. Some methods look precise by hand-picking a few well-matched markets, but the result only holds for those markets. In these cases, precise results aren’t necessarily accurate ones.

That's the story. Everything below is the evidence, and at the bottom, the questions to ask before you move budget on your next readout.

What we actually simulated (and what we didn't)

One objection to get ahead of: Of course a simulation built to vary measurement quality shows that measurement quality matters. The findings aren't the direction, they're the magnitudes: how much precision is worth in win rates and revenue, and whether volume can substitute for it.

One definition, because it matters when comparing solutions. Precision has several ingredients: random noise, systematic bias, and how well the test design represents your business. We model only the first, the random noise in each estimate (what statisticians call precision), so we'll say "precise" and "noisy" throughout. The gaps below come from noise alone; the other failure modes only add to them.

The setup: the true performance of every channel is held fixed. Each experiment returns the truth plus random noise, and after every readout, budget shifts toward whatever looks best, capped by a reallocation limit. Because each move changes the starting point for the next decision, errors compound. We crossed two precision levels with two cadences (15 vs. 6 tests per year), three spend levels, and three reallocation limits. That’s 36 scenarios, each run one million times. Full mechanics and noise parameters are in our methodology walkthrough.

Finding 1: Precision protected the downside 

The headline isn't just that more precise measurement created more upside. It's that noisy measurement created real, quantifiable downside.

Across every spend level and reallocation limit, the less precise approaches finished below the do-nothing baseline 38% of the time. Not “missed an aggressive target.” Worse than if the team had frozen budgets and gone on vacation. Precise measurement failed that way only 18% of the time. Same business, same budget, and the noisy signal more than doubled the odds of ending the year behind where it started.

That 2× is the number I'd put in front of any operator. Imagine two options with a similar shot at a good year, but one doubles your chance of going backward. There's no budget argument that makes that trade worth it.

And the damage isn’t only about how often you lose – it’s about how badly. Take a representative scenario: $30M in annual spend, budget moves capped at 20% per readout. The precise-and-fast and noisy-and-fast approaches post nearly identical best years (+24.9% vs. +23.2% at the 90th percentile). Their worst years are not close: the precise approach’s bad year is down 5.9%, the noisy one’s is down 11.8%. That’s double the drawdown, and it lands in the red far more often.

Precision is downside insurance. It doesn't guarantee a great year, but it makes it far less likely that your measurement system confidently sends you the wrong way.

When a solution is pitched as “directionally accurate,” we picture a wider range around roughly the same midpoint. But a wider range means the midpoint itself moves. A strong channel can look weak, a weak one can look strong, and the team confidently moves money on a misread. No single mistake is catastrophic. The business just drifts through a series of individually reasonable decisions made with more confidence than the evidence deserved.

Precision is downside insurance. It doesn't guarantee a great year, but it makes it far less likely that your measurement system confidently sends you the wrong way.

Finding 2: You can’t out-test a weak signal

The natural counterargument: can't you overcome noise with volume? Run more experiments, let the Law of Large Numbers do its thing?

In theory, yes. A navigation app that points you the right way 51% of the time will still get you home eventually. It will just take a thousand turns instead of fifteen. The problem is that marketers don’t get a thousand turns. Budget decisions run on annual cycles, and in any given year you can only run so many valid, well-powered tests across so many channels.

Moving faster helps when the heading happens to be right. When the heading is off, speed just gets you more lost, faster.

That's where the math turns against a weak signal. Combine two well-powered tests in a year and you're near peak confidence about a channel. To reach the same confidence from a noisy signal, the number of clean experiments you'd need simply doesn't fit in the time you have.

The simulation shows the same thing: the noisy approach running 15 experiments a year beat the baseline 61.8% of the time; running just 6 a year, 62.3%. Tripling the cadence widened the range of outcomes without improving the odds of finishing ahead. With 20% budget reallocation limits, the fast-and-noisy approach's worst-decile year (-11.8%) was substantially deeper than the slow-and-noisy one's (-7.8%).

Moving faster helps when the heading happens to be right. When the heading is off, speed just gets you more lost, faster. 

Finding 3: The precision premium is worth it to your bottom line

So far this has been about avoiding harm. The upside is just as concrete.

At a 20% budget reallocation limit, the precise, high-cadence approach delivered a 13.2% average revenue lift, more than double the noisy approach's payoff (5.6%). Same business, same budget, same number of tests. The only difference was the quality of the signal behind each decision.

One nuance worth stating plainly, because it changes how you read the number. That 7.6% gap isn’t a windfall you collect during the test year. It’s the position you’ve earned by the end of it and an advantage you carry into the next twelve months. Precise measurement doesn’t just win the year you’re in, it sets the starting line for the year after.

That reframes the “too expensive” conversation. A 7.6% gap in revenue dwarfs the difference in cost between a rigorous solution and a cheap one. The question isn’t whether precise measurement fits the budget. It’s whether you can afford to leave that much on the table every year, and start the next one from behind.

The Frontier is a sequence, not a shortcut

The best results came from combining both precise measurement and an ambitious testing cadence. But the takeaway isn't "run more tests." It's the order of operations.

First, build a measurement system that produces evidence you can responsibly act on. Then use automation, operational discipline, and an ambitious roadmap to increase the pace of learning. Cadence compounds the quality of whatever system it's pointed at. Pointed at a trustworthy one, it's how the best programs pull away. Pointed at a noisy one, it just automates wrong turns.

The strongest teams don't optimize for the number of experiments completed. They optimize for the number of decisions they can make with evidence strong enough to deserve the dollars behind them.

And to be clear, this is a multi-year journey, not an overnight turnaround. Our CMO Olivia Kory made this point on a recent Open Haus episode: Some of the biggest wins we see come from brands who take six to nine months to improve one channel. Rome wasn't built in a day. Neither is an elite incrementality practice.

The strongest teams don't optimize for the number of experiments completed. They optimize for the number of decisions they can make with evidence strong enough to deserve the dollars behind them.

Going deeper: how measurement really goes wrong

You now have the full argument. What follows is for readers who want to pressure-test it. If that's not you, skip to the conclusion.

Not all clean-looking numbers are trustworthy

As noted above, the simulation isolates random noise. Real-world precision also depends on systematic bias, design validity, data quality, and representativeness – these matter a lot when comparing solutions, because two providers can report similarly tight estimates while making very different design tradeoffs.

Here's a common one. Some methods improve their apparent statistical quality by selecting only a handful of markets with unusually strong statistical twins. It's like running a drug trial where 500 people volunteer and you keep only the 20 most similar across age, BMI, and health markers. Splitting those 20 into treatment and control makes the drug's effect easier to measure, but the result only generalizes to people who look like those 20.

If you sell nationally, an experiment run on 5, 10, or 15 hand-picked DMAs carries a similar problem: the tight confidence interval doesn't reflect the error introduced by unrepresentative selection. The number looks precise, but it only is for the markets you picked. 

The right question isn't "who reports the lowest MDE?" It's "what had to be sacrificed to produce that number?"

Six questions to ask before you move budget

  1. Is the estimate precise enough to distinguish between the actions we're actually considering?
  2. Are we using the full range of uncertainty, or reacting only to the midpoint?
  3. Does the experimental design represent the population where we'll apply the result?
  4. What happens when the evidence supports a hold or retest rather than a budget move?
  5. Is the proposed change proportionate to the strength of the signal?
  6. What did the provider trade away to make their confidence interval look that tight?

Sometimes the right outcome of an experiment is "learn more before moving." That isn't a failed test. That's the measurement system doing its job.

A note on scope

This simulation is a controlled thought experiment, not a forecast for any individual business or a performance claim about any named provider. It covers one year and does not fully simulate systematic bias, unrepresentative market selection, data-quality failures, external market changes, or more conservative decision processes such as holds, retests, or Bayesian updating. Every business is different; the point is to test for yours.

Earn the right to move quickly

"Any incrementality is better than no incrementality" is useful if it gets a team to question platform reporting. It becomes dangerous when every estimate is treated as a mandate to move money.

Across 36 million simulations, more precise measurement beat the do-nothing baseline more than 80% of the time. Less precise measurement left the business worse than doing nothing roughly 38% of the time, and testing more didn't change those odds.

The lesson isn't to slow down. It's to earn the right to move quickly by investing in a solution that captures the true value of your marketing. Then test, test again, and test some more. 

Subscribe to our newsletter

Article Authors

Patrick Hillery

Patrick is a Senior Solutions Consultant at Haus. With 18 years of experience spanning analytics, marketing, and B2B SaaS, Patrick has led multiple 0-to-1 product launches and founded two ventures that each ended in acquisition. He brings a rare combination of technical depth, working fluidly with data science and engineering teams, and a strong customer-facing instinct.