Podcast page image

The Brandformance Podcast • Ep 65

How Intercom built an in-house marketing measurement engine

With Raunak Kumar

With Raunak Kumar

Senior Manager, GTM Analytics, Intercom

Senior Manager, GTM Analytics, Intercom

In this episode, Raunak Kumar — Senior Manager of GTM Analytics at Intercom (previously Stripe and Atlassian) — walks through how to actually measure B2B marketing when clickbased attribution stops telling the truth. He argues that the hard part of incrementality and marketing mix modeling isn't the machine learning; it's the decisions — what outcome you model toward, how you segment self-serve versus enterprise motions, and how you communicate uncertainty to leadership. Raunak breaks down a full-funnel geo holdout at Intercom that revealed a roughly 20% halo lift from brand media (with 18% flowing to organic and direct, plus SEM spillover), and the data-quality bugs the experiment surfaced along the way. He shares a pragmatic build-vs-buy framework — build MTA and MMM, get help on testing logistics — and where Claude Code fits (and doesn't). It's a candid, tactical playbook for anyone doing measurement in a B2B, PLG-plus-sales-led business.

Episode details

Transcript

Behind the expert Raunak Kumar is Senior Manager of GTM Analytics at Intercom, where he leads marketing measurement for a company mid-pivot from SaaS to AI-first. He has spent his career in B2B analytics — a computer science background, then customer data, then marketing analytics — across Atlassian, Stripe and now Intercom. Over the last few years he and Pranav have been on parallel journeys exploring incrementality and marketing mix modeling in a B2B world, which behaves very differently from the CPG and consumer playbooks most MMM advice is written for. This Office Hours conversation is a candid, tactical walk through how he actually does it: why click-based attribution stops being enough, the decisions that matter more than the modeling, a full-funnel geo holdout that surprised everyone, and where tools like Claude Code fit. The gist • The hard part of incrementality and MMM isn’t the machine learning — it’s the decisions: what you model toward, how you segment, and how you communicate uncertainty. • Pick an outcome metric that’s both responsive to marketing and predictive of revenue — and segment self-serve from sales-led, because their ACVs and journeys differ wildly. • A full-funnel geo holdout at Intercom revealed a ~20% halo traffic lift from brand media — most of it in organic, direct and even branded search. • Experiments don’t just measure lift; they expose data-quality bugs an MMM would quietly smooth over. • Build your MTA and MMM in-house for flexibility; get help on testing logistics. And feed Claude Code the business context, or it can’t do the nuanced work. Why incrementality: when cutting spend changed nothing Raunak works in growth companies where marketing is in the driver’s seat, budgets are big, and it’s easy to lose sight of where the next dollar should go. The traditional leading indicator is MTA — click-based attribution — which tells you a lot and isn’t wrong. But the aha moment came when they cut spend on a major channel and expected pipeline to drop… and nothing happened. That, plus brand work at Atlassian that was hard to measure through clicks, and post-COVID tracking limitations that inflated a mysterious “direct / catch-all” channel, pushed him toward incrementality and eventually MMM. At Intercom the buy-in was mixed until he found a champion (a fully bought-in head of brand) — and, crucially, until he had a real test result in his back pocket to rally leadership around. The decisions that matter more than the data science Raunak’s central theme: the modeling is the easy part; the judgment calls are what make or break it. Start with data discovery, not the model. His first instinct — one output variable, one MMM — was a mistake, because a self-serve signup worth $50–100 in ACV and a sales-led deal worth $50,000 aren’t the same thing. Treat them equally and the model happily optimizes for fast, non-sticky signups. At a minimum, segment self-serve versus sales-led; at Stripe he settled on signups as a clean quality metric, and at Intercom the enterprise model targets S2 (post-MQL) pipeline. The rule of thumb he and Pranav share: pick a metric that’s responsive to marketing and predictive of downstream revenue — ideally within about a 30-day window from touch to pipeline. If MQL-to-pipeline takes longer than a month or two, move the outcome up-funnel to MQL. And a byproduct of doing this well: you learn that conversion rates differ by channel because the intent of the audience differs. Ad stock, uncertainty, and geography Three more decisions round it out. Ad stock — how long a touch’s effect lasts — is the input everyone struggles to explain; Raunak had historically pegged brand’s lag at around eight weeks from imperfect click data, and his holdout gave him a definitive, and very different, number to feed the model. On uncertainty, he pushed his team away from single definitive ROI figures on board slides (a number shown once becomes a target you have to defend forever) toward tighter, confidence-interval-based ranges. And geography matters: Intercom is global, so region is a key segment. A NAMER lead isn’t worth the same as an APAC lead, and you can’t blindly apply a US ad-stock and response curve to an under-penetrated APAC market with more headroom. He’s starting global-by-motion and may ladder up to a hierarchical model — while staying wary of running too many models at a mid-sized company. The geo holdout: a 20% brand halo The centerpiece is a full-funnel geo holdout run at the start of the year — holding out every brand channel (display, programmatic, YouTube, Meta, LinkedIn, Reddit). Rather than a risky 50/50 split that might darken California or New York, they built six US state pairs (Ohio– Pennsylvania, and so on) using a composite similarity score that went beyond trend-matching to volume, demographics, audience composition and even brand exposure from local offices and events. Designed to detect a 6% lift over eight weeks, the test came back at roughly 20x the minimum detectable effect: a ~20% halo traffic lift with a p-value under 0.001 (he re-ran the analysis three times to be sure). Decomposed, the lift was about 80% direct/paid clicks, 18% halo in organic and direct, and 3% SEM spillover — brand awareness measurably feeding branded search. A counter-intuitive twist: the control group actually had a higher visit-to-lead conversion rate, because the test states pulled in lower-intent, newly brand-exposed visitors — which is exactly the point. Experiments expose what models smooth over One of Raunak’s favorite arguments for testing is that it surfaces problems an MMM would paper over. Mid-test, Ohio (a control state) spiked — a bot/spam issue where logging was misclassifying data-center traffic to Ohio. Because he was watching a live experiment, he caught it and worked with data engineering to fix an actual bug; an MMM would have averaged it away. Same with a paid-display UTM that had been silently misclassified as direct after a midflight UTM change. His discipline now: every one of those learnings — the halo number, the three-week ad stock, the Ohio bug — goes into a memory.md and skills.md file so the next analysis, whether he runs it or Claude Code does, knows to look for them. Which sets up his stance on AI: Claude Code can do the work if you’re deliberate, but the business context (does Ohio behave like Indiana? which states have brand exposure from events?) still has to be fed in. Build vs. buy (and the maturity ladder) Raunak’s answer is hybrid, organized as a three-layer ladder. MTA at the bottom — build it inhouse for the flexibility to exclude junk touchpoints (auto-responder emails he “nukes” first thing) and to choose your weighting (W-shape at Intercom, last-touch elsewhere). MMM in the middle — also build, leaning on open source like Google Meridian with a dedicated data scientist for tight, weekly iteration on things like Hill-function ad-stock curves; the underlying statistics haven’t changed much since the 1980s. Testing on top — run it internally but get agency help on the logistics of negating campaigns and capturing clean spend and impression data. That data problem is his recurring warning: good weekly time-series spend and impression data is genuinely hard to get, especially for YouTube, Meta, Reddit and CTV — he’s currently scratching his head trying to source CTV impression data for a new incrementality test. Partners help here, but it’s the unglamorous foundation everything else sits on. Quote snacks • “We stopped spend on a big channel and pipeline didn’t move. That was the aha moment.” • “What’s the outcome you’re modeling towards? Couple it with the data nuances of your business.” • “We saw upwards of a 20% halo traffic lift, with a p-value under 0.001. I redid the analysis three times.” • “Present ranges, not points. Putting out a single ROI number is misleading.” • “Claude can do it if you’re deliberate about it — but the business context still has to be fed in.” • “MMM thrives on variability. If you’re marketing at a steady state, it’s very hard to pick up signal.” Why it matters This is the brand-versus-performance debate translated into a B2B measurement problem. Click-based attribution flatters the last-touch channels and misses the brand media that, in Raunak’s holdout, drove a 20% halo — including a measurable lift in the very branded search that attribution loves to credit. The performance and brand stories aren’t opposed; they’re just measured on different tools and time horizons. His deeper point is a corrective to the AI-and-MMM hype cycle. The models and the agents are commoditizing fast, but the durable skill is judgment — choosing the right outcome, segmenting honestly, communicating uncertainty, and running the messy experiments that expose the truth. And a sober caveat: MMM only works with variability and clean data, so B2B teams should make sure there’s real appetite and real signal before they jump in. Practical next steps Start with the data, not the model. Clean weekly time-series spend and impression data is the bottleneck; one year is workable, and you can segment creatively if history is short. Pick your outcome variable deliberately. Choose something responsive to marketing and predictive of revenue, and don’t over-credit fast, non-sticky signups. Segment by motion. Model self-serve and sales-led separately (and consider region), because ACVs and journeys differ. Run experiments to bulletproof the model. A geo holdout gives credible data points — halo lift, real ad stock — and surfaces data bugs your MMM would smooth over. Don’t confuse ad-stock decay with decision horizons. Effects can converge in weeks but compound later; pair decay with time-to-pipeline analysis. Present ranges, not points, and feed your learnings to your AI. Document data quirks in memory/skills files so the next analysis — human or Claude Code — carries the context forward.

The sharpest marketing conversations, straight to your inbox

By providing your contact info, you agree to receive communications from Paramark. You can opt-out at any time. For details, refer to our Privacy Policy