

The Brandformance Podcast • Ep 61
Can AI predict incrementality without running experiments?

In this episode of Brandformance's marketing-science segment, co-hosts Pranav Piyush (Paramark) and Sundar Swaminathan (ex-Uber, creator of ExperiMENTAL) do their first paper breakdown — “Predicted Incrementality by Experimentation (PIE)” by Brett Gordon, Robert Moakler and Florian Zettelmeyer. They unpack how a model trained on ~2,200 Meta randomized controlled trials can predict the incrementality of campaigns that were never tested, hitting an R-squared of 0.88 versus 0.19 for last-click — and why “post-determined” metrics like clicks aren't as useless as marketers think. They explain why this doesn't free brands from running RCTs (it mostly reveals how Meta's “incremental attribution” works), where the model breaks down for new advertisers and verticals, and why benchmarks deserve real skepticism. The takeaway: predicting incrementality is impressive, but it's a triangulation input, not a replacement for the test.
Episode details
Transcript
Your hosts
This is the marketing-science leg of Brandformance, co-hosted by Pranav Piyush and Sundar Swaminathan. Every two weeks they break down the concepts behind modern marketing measurement — incrementality, marketing mix modeling, brand versus performance — in plain, practical terms, and bust a few myths along the way.
Pranav Piyush is the co-founder and CEO of Paramark, where he’s building modern marketing measurement — incrementality testing and marketing mix modeling — for growth teams. A long-time marketing leader, he hosts Brandformance.
Sundar Swaminathan is a marketing data science and experimentation advisor to consumer-tech scaleups and the creator of the ExperiMENTAL newsletter and podcast. He previously built and led brand data science at Uber — measuring over a billion dollars of brand spend — and earlier worked on the debt desk at the US Treasury.
The gist
This episode breaks down a single paper: “Predicted Incrementality by Experimentation (PIE) for Ad Measurement” by Brett Gordon, Robert Moakler and Florian Zettelmeyer.
PIE uses ~2,200 Meta randomized controlled trials to learn a model that predicts the incrementality of campaigns that were never run as experiments.
It hits an out-of-sample R² of 0.88 — versus 0.19 for seven-day last-click — partly by using “post-determined” metrics like clicks and last-click conversions as features.
It doesn’t free brands from running RCTs. It mostly explains how ad platforms (like Meta’s “incremental attribution”) build their models — and the rich get richer.
It generalizes within an advertiser and vertical, breaks down across new ones, and is a triangulation input — not a replacement for experiments.
The paper: can you predict incrementality without the test?
Pranav brought a paper that kept showing up in his LinkedIn feed: “Predicted Incrementality by Experimentation,” or PIE. The setup is bold — reframe ad measurement as a prediction problem. The authors took roughly 2,200 Meta RCTs (randomized controlled trials, i.e. incrementality tests), learned the relationship between campaign-level features and the measured causal lift, and then used that model to predict incrementality for campaigns that had no control group at all.
Sundar’s practitioner read: strip away the marketing-science packaging and it’s a familiar idea — use the data you have to predict the things you can’t test. That’s common and doable; what’s exciting is the application to incrementality specifically, and the fact that the authors are careful about exactly what it is and isn’t.
Last-click isn’t useless after all
The headline result: an out-of-sample R² of 0.88, versus about 0.19 if you lean on seven-day last-click alone. The surprise for measurement people is what powers it — “post-determined” features like exposure rates, clicks and last-click conversions, the very metrics marketing scientists love to throw away. Used in combination, last-click as a feature meaningfully lifts the model’s predictive power.
Sundar’s caution keeps it honest: the arrow only points one way. Incremental lower-funnel tests tend to drive a lot of clicks — but the inverse isn’t true. Driving clicks doesn’t mean you were incremental. The paper teases that relationship out cleanly; the danger is people reading “last click matters” and jumping straight back to judging campaigns on clicks.
It doesn’t free brands from RCTs
Predictability kept improving as the training set grew from 50 to 1,500-plus RCTs — which means you need a very large corpus before the model is trustworthy. So this isn’t a brand running a handful of tests and never testing again. As both hosts stress, and as the paper’s first line says, RCTs remain the gold standard.
Where it really matters is the platforms. PIE essentially explains how Meta built “incremental attribution” — using twenty years of first-party data and brand-lift studies as ground truth to predict, on a graded scale, how incremental a given campaign is. It’s a genuine advantage of owning the data. The catch is that only an elite tier of brands — the Bookings, Airbnbs and Ubers with the RCT volume and the culture — can do this in-house. As Sundar puts it, the rich get richer.
Where it breaks: new advertisers and new verticals
The most important caveat: the model generalizes well to new campaigns by the same advertiser, and much less well to brand-new advertisers. A booking.com with deep RCT history gets good predictions; a company that has never run an RCT gets a far weaker signal. The same holds across verticals — there are orders of magnitude more RCTs in e-commerce than in B2B SaaS, so the model is far better at one than the other.
The practical proxy Sundar offers: look at how often your category and your peers run RCTs. If you’re a scaled consumer brand in a vertical full of experiments, you can lean on this more; if you’re a Series A still finding product-market fit, the model has nothing to go on, so just run your own tests.
The benchmark trap
Pranav used a reading tool to interrogate the PDF, and one line stuck: marketers shouldn’t blindly import calibration factors from other categories or years. It’s why Paramark has deliberately never published incrementality benchmarks — the numbers swing so much by vertical and require so much data that a tidy benchmark invites exactly the wrong behavior.
Sundar agrees: benchmarks strip away the context that matters — margins, AOV, goals, whether a company is even VC-backed. A benchmark can be a rough calibration check (“are we roughly in line, or way off?”), but most people treat it as a goal and never question it. Both land where the authors do: treat PIE as one more input in your triangulation, not a replacement for judgment or experiments.
What it means for ad platforms
The hopeful takeaway is about the platforms. Every ad network — Meta, Google, Reddit, Pinterest, TikTok — could make RCTs easier and encourage far more of them. Many will run “shadow RCTs” behind the scenes and feed the results into their attribution, which means accepting some short-term revenue hit in exchange for better measurement and more advertiser trust over time.
Meta has already done this with incremental attribution, and Amazon appears to be doing something similar inside its attribution system. For Pranav, that’s the exciting part — the paper doesn’t just describe a model, it demystifies how the incremental-attribution numbers in your ad account are actually produced.
Quote snacks
“RCTs are the gold standard. That’ll never change.” — Sundar
“Just because you’ve driven clicks doesn’t mean they’re incremental.” — Sundar
“This doesn’t change much for a brand — it gives ad platforms a way to improve their attribution.” — Pranav
“With all marketing science, there’s an elite group that can do this. The rich get richer.” — Sundar
“Don’t think of PIE as a replacement — think of it as another part of your triangulation.” — Pranav
“Don’t blindly import calibration factors from other categories or years.” — from the paper
Why it matters
This is the show getting concrete: a real paper, real numbers, and a clear-eyed read on what they do and don’t mean. The big unlock is conceptual — the “incremental attribution” column quietly reshaping budgets inside Meta and Amazon isn’t magic; it’s a model trained on a huge corpus of experiments, with all the strengths and limits that implies.
It also reinforces the series’ throughline. Models are tools to be triangulated, not oracles to be obeyed; experiments stay essential; and context — your vertical, your stage, your RCT history — decides how much you can trust any predicted number. Predicting incrementality is impressive, but it doesn’t retire the test.
Practical next steps
Don’t read this as “stop running RCTs.” They remain the gold standard; predictive models extend them, they don’t replace them.
Treat platform “incremental attribution” as a modeled estimate. Meta’s and Amazon’s numbers are predictions trained on their own experiments — useful, but worth validating with your own tests.
Judge whether modeling applies to you. Gauge how many RCTs your category and peers run; the more experimentation in your vertical, the more you can lean on predicted incrementality.
Resist the benchmark trap. Use external numbers as a rough calibration check, never as an imported goal, and always question the context behind them.
Keep triangulating. Combine RCTs, MMM and attribution; let PIE-style predictions be one input among several, not the answer.
More sharp conversations
The sharpest marketing conversations, straight to your inbox
By providing your contact info, you agree to receive communications from Paramark. You can opt-out at any time. For details, refer to our Privacy Policy



