Incrementality Testing Is Maturing—Now What?
Haus CMO Olivia Kory and brand operators unpack how incrementality testing has evolved from a scorecard into a continuous decision-making system—and whether AI should run it.
Written by AI. Jonathan Park

Photo: AI. Eira Pendragon
The first stage of any measurement revolution is convincing people it matters. The second—harder—stage is figuring out what to actually do with it.
That's the tension at the center of a 73-minute Marketing Operators episode (per the 9operators.com episode listing) featuring Olivia Kory, CMO and Strategy Officer at Haus, alongside Ridge CMO Connor MacDonald and HexClad Head of Growth Connor Rolain. The conversation is less about whether incrementality testing works and more about something more operationally thorny: what happens after you have the numbers.
The Scorecard Problem
Kory describes Haus's story in two chapters. The first, running roughly through 2021 and 2022, was evangelism—convincing brands that incrementality measurement wasn't just for Airbnb-scale companies. That chapter, she says, is effectively closed.
The second chapter is harder to summarize in a single sentence, which is kind of the point. "They're thinking about it more as a culture of continuous experimentation than a report card," Kory says of where sophisticated brands have landed.
That reframing isn't cosmetic. A scorecard invites a binary judgment—channel passes or fails—and binary judgments in marketing tend to produce bad decisions. MacDonald describes this clearly: his team would run three-week geo-lift holdouts on Meta, get a conservative incremental ROAS reading, and be left with a number that was probably underrepresenting the channel's true contribution because the test window was too short, seasonality affected the read, and the customers in the holdout cells weren't blank slates—they'd been seeing Ridge's ads for years before the test ever started.
Kory's explanation for why that last wrinkle doesn't invalidate short holdout tests is worth understanding: if your customers have been exposed to your ads for years, that historical exposure exists in both the holdout and the treatment group. So it lifts both baselines simultaneously and doesn't contaminate the comparison. What it genuinely doesn't answer is whether the compounding long-term effects of advertising are worth measuring—and for upper-funnel channels like CTV, Kory argues you need a much longer test window to get at that question.
Optimize the Channel You're In Before You Expand
The most tactically useful section of the conversation covers sequencing: when to run a channel-wide holdout versus when to run optimization tests inside your existing channels.
Rolain's position is fairly direct: if Meta built your business, don't burn your first six months with a measurement platform running a binary "is Meta working?" holdout. You already know Meta is working—it got you here. The more valuable early tests, he argues, are intra-channel comparisons: view content versus purchase conversion campaigns, video versus image, upper funnel versus lower funnel spend. HexClad now allocates millions to view content on Meta as a direct result of this kind of directional testing.
"The best wins and learnings we've gotten in the last year and a half" came from that three-cell optimization framework, Rolain says. "I look at our account and we're spending millions on view content now. We never would have done that if we didn't have that directional data."
Kory adds the counterargument for why channel-level holdouts do have a place: when you want to launch YouTube or CTV, you need something to compare against. And if your MMM has never seen a clean holdout on your biggest channel, it's essentially running blind on its most important input.
Both things can be true. The channel holdout gives you a reference point for new channel comparisons and calibrates your model. The optimization tests inside an existing channel give you the tactical wins that actually move spend allocation from week to week. The disagreement in the room isn't really about which is right—it's about which to prioritize first.
What "Debiasing" Actually Means
The more technically interesting thread involves what Kory describes as debiasing attribution data. The practical problem with day-to-day attribution tools—your Northbeam, your GA4, your in-platform dashboards—is that they report ROAS using models that can't distinguish between incremental sales and sales you would have gotten anyway. Holdout experiments can isolate that incrementality, but they're not running in real time.
Haus's answer to this is a machine learning model built on top of a pixel that pulls real-time event data. The model takes the historical experiment database—across all Haus customers, not just your own—and uses it to debias the attribution signal. The goal: when you're looking at a campaign ROAS in your dashboard at 9 a.m. on a Tuesday, the number you're seeing has been adjusted based on what experiments have shown about how much of that reported return is actually incremental.
MacDonald frames it as a description of what his team is already attempting manually: "Okay, we're getting this result from this Meta campaign and the ROAS is really low, but based on all these previous tests, we're going to debias and say we're comfortable spending on this campaign." What Haus is building automates that judgment and applies it at a granularity no human team is realistically tracking across hundreds of campaigns simultaneously.
The incrementality index—Haus's aggregate database of experiment results across its customer base—also addresses what happens when you've never run a holdout on a given channel. A brand that's never tested Pinterest can still get a prior estimate based on how Pinterest has performed for structurally similar brands. Kory is careful here: brand-specific results always take precedence when available. But for a cold start, cross-customer data is meaningfully better than nothing.
The Agentic Question
The episode's most forward-looking section is also the one with the fewest resolved answers. Kory asks the question directly: if you have experiments, an MMM, and debiased attribution all pointing in the same direction, why is a human doing the budget allocation?
MacDonald's answer amounts to: agreed in principle, not yet in practice. The context problem is real—promo calendars, product launch schedules, inventory constraints, the CEO's emotional attachment to billboard advertising on the 405. These are hard to encode. But he doesn't treat them as permanent blockers.
"There isn't a doubt in my mind that a system could do that better than a person," MacDonald says, talking about the challenge of holding all the dimensions of incrementality data in your head simultaneously while making campaign-level budget calls.
Kory describes what Haus is actually building on this front: the system currently produces budget recommendations that brands can approve or reject. Full autopilot—where the system executes the buy directly via API into the ad platforms—isn't live yet. A handful of brands are accepting the recommendations. Trust, Kory notes, took time to build.
The honest version of this picture is that automated budget decisioning is closer than most brands' internal processes assume, further than the most excited vendors suggest, and entirely dependent on the quality of the measurement signal underneath it. If your ROAS data isn't debiased, you're automating noise.
Bad Results Aren't the Problem
One thing Kory says that deserves more attention than it gets in the conversation: her user research surfaced brands describing a bad Haus result as "the biggest fire in their organization."
That's a symptom of using measurement as a report card—which is what the entire conversation argues against. A negative result that prompts quick reallocation isn't a bad outcome. HexClad's discovery that non-purchase conversion objectives weren't more incremental than purchase conversion campaigns is the example Rolain offers: "Is it a bad result if you action it quickly?"
The organizations that struggle with measurement aren't usually struggling because their numbers are wrong. They're struggling because leadership hasn't built a culture that knows what to do with an unflattering number. That's an org design problem dressed up as a measurement problem, and no geo-lift holdout will fix it.
The more useful question isn't whether your channels are passing or failing their incremental ROAS hurdles. It's whether you're consistently reallocating from lower-performing tactics to higher-performing ones, and whether the pace of that reallocation is faster than your competition's.
— Jonathan Park, Business Desk Editor
We Watch Tech YouTube So You Don't Have To
Get the week's best tech insights, summarized and delivered to your inbox. No fluff, no spam.
More Like This
How APL Turned an NBA Ban Into a Brand Identity
APL's NJ Falk breaks down the six storytelling threads behind one of e-commerce's most unlikely brand-building moves—starting with an NBA ban.
Bill Ackman's Art of Finding Invisible Money
Explore Bill Ackman's strategy of finding hidden value in companies, beyond standard metrics.
USMCA Won't Be Renewed: What Annual Reviews Mean
The Trump administration has blocked USMCA's 16-year renewal, opting for annual reviews. Here's what that means for trade, business, and North America.
Meta Pixel Misconfiguration Is Quietly Draining Ad Budgets
How a single tracking setup decision caused Ridge's travel ads to fund wallet sales instead. What it means for brands, consumers, and Meta's growing power over both.
Email Isn't a Revenue Channel—It's an Ad Cost Problem
Sammy Tran argues e-commerce brands are measuring email all wrong. The real job? Making your paid ads cheaper. Here's what that actually looks like.
How Parakeet Chat Made $1.5M by Serving a Niche Market
Discover how Parakeet Chat generated $1.5M by targeting an overlooked niche: incarcerated individuals seeking legal knowledge.
RAG·vector embedding
2026-08-15This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.