Insights
The MMM terms every marketing leader should know
Every CMO has nodded along to terms like endogeneity, ad stock and credible intervals. A plain-English guide to the MMM concepts marketing leaders actually need, so you can challenge assumptions and get more from your model.


Sundar Swaminathan
Marketing Science Advisor
The MMM terms every marketing leader should know
Every CMO has sat through an MMM presentation where a vendor or a data scientist drops a term like endogeneity, ad stock, or credible interval.
Silence.
Nobody wants to be the one asking what it means.
We've been on both sides of that table. First, as the people nodding along until we eventually had to learn what we'd been nodding at. Today, we're often the ones explaining these concepts in plain English.
So, we put together this guide to save other marketing leaders from those awkward moments and help them ask better questions, challenge assumptions, and get more value from their MMMs.
Baseline
Baseline is one of the most fundamental concepts in MMM, and also one of the most debated. After all, the goal of an MMM is to estimate what marketing is driving and, by extension, what it isn't. There are two common ways people define it.
Definition 1: What's left if you turned off marketing tomorrow
The working definition most teams use is: How much revenue would you have if you turned off all marketing today? That's your baseline. Everything on top of it is incremental and driven by marketing.
The problem is that this definition is only partially right.
Baseline isn't a fixed number. It's the result of all the prior work (marketing, brand building, and product investment) that got you to where you are today.
Let’s look at two examples.
After 30 years of building its brand, Nike decided to cut brand marketing and lean into performance marketing. Even though total marketing spend stayed the same, the mix changed, and its baseline started to erode. Five years later, asking "What happens if we turn off marketing?" leads to a very different answer because brand strength, competition, and retail availability have all changed.
Now imagine you’re a cereal box.
Simply being on the shelf generates sales without additional marketing. That shelf presence is the baseline. But retail stores stock the best-performing products at eye level. If you cut spending and sell less, you get less shelf visibility, which means you sell even less. Over time, your baseline starts trending toward zero.
Definition 2: Whatever the model couldn't explain.
Statisticians often define baseline differently. To them, baseline is whatever the model couldn't explain. The residual. After accounting for all marketing channels, seasonality, pricing, and control variables, whatever is left over becomes the baseline.
The problem with this definition is that it's tempting to keep adding variables in an attempt to explain more of the residual. But that can hurt the model, resulting in a problem called overfitting.
Sometimes there are things in the real world that are simply not modelable. In those cases, the baseline genuinely captures the unexplained. Other times, however, what's left unexplained is simply a relationship the model hasn't identified yet.
Both definitions are partially right, but neither is complete.
The important thing to remember is that baseline isn't static. It naturally trends toward zero over time. You're either investing in it and building it up, or you're watching it decay.
Control variables
Control variables are factors that influence your business but have nothing to do with your marketing decisions.
Seasonality is the obvious example, but there are many others, including:
Price changes,
FX rates,
Product outages,
Crypto sentiment (if you're in that space),
Or even a Cloudflare outage that takes your site down for a day.
Without accounting for these factors, your MMM might attribute changes in performance to marketing when they were actually caused by something else. That's why control variables matter: they help the model isolate marketing's true impact by accounting for everything else.
One important warning: don't throw every variable you can think of into the model. More isn't better. Too many control variables create their own problems, and the model starts picking up misleading correlations. Be thoughtful about which variables truly drive your business at scale.
Ad stock and carryover
These are two ways of describing the same phenomenon: carryover is the psychological concept, while ad stock is its statistical representation.
Carryover describes how people remember your brand over time. If Booking.com stopped advertising tomorrow, you'd still remember to use them next month. But what about the month after? Or the month after that? Carryover is simply the "inventory" of your brand in consumers' minds, slowly decaying over time.
Think of carryover like a store that only sells one product. Now, imagine you stop restocking the shelves. Customers don't immediately notice. There's still product there.
But every day there's a little less, and eventually the shelves are empty. That's the decay curve. How fast the shelves empty depends on how strong your brand is to begin with. A well-known brand with high mental availability will have a longer carryover than a challenger brand that most people haven't heard of.
Ad stock is the same idea, but from the model's perspective. It measures how much impact your marketing continues to have after you stop spending, including its decay and retention over time.
The practical implication is this: when you run a test, don't assume your ad stock drops to zero the day you cut spend. If you run an eight-week holdout but had been spending heavily beforehand, the first few weeks of that test are still benefiting from residual ad stock. End the test too early, and you'll underestimate how long your previous marketing continued to have an effect. Always build an additional measurement window at the end of your test to capture that tail.
Saturation curves and diminishing returns
Again, these are two ways of describing the same concept: diminishing returns is the phenomenon, while the saturation curve is its statistical representation.
Diminishing returns describe what happens as you increase your marketing spend: each additional dollar generates less incremental value than the one before it.
Your first $10k in marketing reaches the people most primed to buy: early adopters, high-intent users, and people who already have the problem and are actively looking for a solution. As you keep spending, you start reaching people with a slightly less urgent version of that problem. Then, eventually, people who didn't even know they had the problem. Conversion gets harder, and your cost per acquisition creeps up.
The saturation curve describes how quickly those diminishing returns happen. The shape of that curve varies by category. For example, AI products currently tend to have relatively flatter curves as there's still a lot of room to acquire new users efficiently. Retail apparel is the opposite: fierce competition, saturated audiences, and steep diminishing returns.
Frequentist vs. Bayesian
As a Marketing leader, you don't need to become an expert in statistical methods. But understanding the difference between Frequentist and Bayesian models will help you ask better questions when evaluating an MMM.
At a high level, Frequentist statistics uses observed data to estimate what's true based on the assumption that the world follows predictable distributions. Bayesian statistics starts with what you already know (called a prior) and updates that belief as new information becomes available.
In the context of MMM, most modern vendors use Bayesian methods because they let you feed in priors from your incrementality tests. You run a geo test, get a real causal estimate of how incremental your Meta spend is, and use that result as the model's starting point. The model then updates its estimates around that anchor. That's much better than starting from generic assumptions.
You don't need to understand the statistical mechanics behind this. What you do need to know is: Does your MMM vendor accept priors? And have you given them real incrementality data to work with? If the answer is no to both, the model is starting from scratch instead of building on evidence.
Statistical significance, confidence intervals, and credible intervals
Statistical significance
Statistical significance has more mysticism around it than it deserves.
The 95% threshold isn't a law of nature. It became popular after a statistician in the early 1900s printed statistical tables using that threshold, and the industry gradually adopted it. It represents how much risk you're willing to accept when making a decision under uncertainty, and that threshold varies by context.
Pharmaceutical trials need high thresholds because the stakes involve human lives. For most marketing decisions, you can be more pragmatic. Choose your significance threshold based on the business risk of being wrong, not an arbitrary convention.
At Uber, we ran email tests at 95% two-tailed significance for years. Zero marginal cost, high volume, fast iteration. We eventually moved to 80%, and nobody cared. Decisions got faster, we tested more, we learned more.
While statistical significance tells you whether an effect is likely to exist, confidence and credible intervals take that one step further by showing how precise your estimate is.
Confidence vs. credible intervals
If you truly want to get nerdy, confidence intervals and credible intervals are not the same thing, even though most people use them interchangeably.
Confidence Interval (Frequentist): If I repeated this test many times, how often would the interval contain the true value?
Credible Interval (Bayesian): Given the data I have today, what's the range within which the true value is likely to fall?
As a Marketing leader, you don't need to memorize the statistical definitions. But when vendors throw these terms around, ask which one they mean, and bring your Data Science team into the conversation if needed.
Multicollinearity and endogeneity
These are two terms that sound intimidating, but they both describe situations where an MMM struggles to identify true cause and effect. In other words: is the model actually measuring what you think it's measuring?
Multicollinearity
Multicollinearity happens when two channels move together so consistently that the model can't separate them.
A common example is a marketing team that manages both Meta and Google and moves budgets proportionally every month — up 20% on each, down 10% on each. The model sees two channels that always move in lockstep and has no way to attribute results to one versus the other. There's simply not enough variation to work with.
TV and branded search are another good example. TV drives awareness, awareness drives branded search, and branded search appears effective in the model when it's capturing demand that TV created.
Endogeneity
Endogeneity is trickier. It's when an input and the outcome you're trying to measure influence each other, creating a circular relationship that breaks the model's logic.
Branded search is the poster child. Most teams include branded search as a marketing input in their MMM to predict sales. But demand also drives branded search: when more people want your product, they search for it. That means the model is using something that's partly a consequence of sales to explain sales. The causal direction runs both ways, and the model can't cleanly separate them.
The practical takeaway is simple: think carefully about what you're including as an input. If a variable tends to increase when sales increase for reasons unrelated to your marketing decisions, it's endogenous and needs to be treated differently.
Model validation: in-sample vs. out-of-sample
Any model is good at fitting the data it was trained on. The real test is whether it can accurately predict data it's never seen. That's the difference between in-sample and out-of-sample performance.
In-sample: how well does the model fit the historical data you gave it?
Out-of-sample: how well does it predict a hold-out period it wasn't trained on?
Out-of-sample performance is what really matters, because you're using the model to make future decisions. Backtesting the model on a hold-out period and seeing how well it predicts that unseen data is the most honest way to assess its performance.
Your job isn't to run these tests. It's to ask your data science team: Did you backtest the model? What was its out-of-sample error?
Across all of these concepts, don't get caught going too deep into the statistics. It can quickly become a rabbit hole, and there will almost always be multiple schools of thought. Instead, focus on understanding the assumptions, limitations, and potential pitfalls of your MMM.
One final thing: treat your MMM as an estimation tool, not a prediction machine. If your vendor is promising 98% predictive accuracy, that's a red flag, not a selling point. Nobody can predict the future.
The goal isn't perfect predictions; it's making directionally better decisions with more evidence behind them. That's what a good MMM delivers.
If you'd like to explore these concepts in more depth, we discuss them in this episode of the Brandformance podcast.
While you're here
Find your edge
Book a demo now and you'll get a live, expert-run walk through of how Paramark can help you with incrementality measurement and experimentation.




