

The Brandformance Podcast • Ep 67
What thousands of A/B tests reveal about websites that convert

In this episode, Casey Hill, CMO of DoWhatWorks, a platform that tracks the A/B tests brands run across the web, shares what thousands of real experiments reveal about the websites that convert. He argues that in the AI era, building is easy, so the hard questions are what to test and how to measure it, and offers a litmus test for a meaningful experiment (does it introduce new information or new value?), since most variants 'lose' by simply doing nothing. Casey unpacks why specificity beats generic benefits, why logos actually lose in testing and how social proof should change in an age of fakes, and why letting users self-select can beat dynamic personalization. He closes on the 'are websites dead?' debate, making the case that your site is still the hub LLMs draw from, with real examples of companies getting cited by ChatGPT within 30 days. It's a tactical, data-backed CRO masterclass for anyone optimizing a B2B website
Episode details
Transcript
Behind the expert
Casey Hill is the CMO of DoWhatWorks, a platform that monitors the A/B tests brands run across the web — SaaS, e-commerce, fintech and more — and turns them into recommendations tailored to your situation.
He’s spent 15 years in software: his first company, Tech Validate, was acquired by SurveyMonkey; he’s consulted for the likes of McKinsey and BlackRock, guest-lectured at Stanford, and built a large following writing on Quora and then LinkedIn. He came up through sales into growth (Bonjoro), spent time at ActiveCampaign, and now runs marketing at DoWhatWorks — where he has a rare window into what actually converts across the best B2B websites.
This Office Hours conversation is a practical CRO masterclass: how to think about experimentation in the AI era, a litmus test for whether a test is worth running, and a stack of counter-intuitive findings from the data.
The gist
• In the AI era, building has never been easier — so the hard, valuable questions are what to build, how to think about it, and how to measure it.
• Most A/B tests don’t fail so much as do nothing: “lose” usually means neutral. A good test introduces new information or new value.
• Specificity wins. Generic benefits (“drive more revenue”) get scrolled past; concrete capabilities and numbers convert and get remembered.
• Logos lose in testing at a high rate. Proof now hinges on how hard it is to fake — interactive case studies and video beat grayed-out logo bars.
• Websites aren’t dead in the AI era — your site is still the hub LLMs draw from, and clear structure serves humans and agents alike.
Inside a database of thousands of A/B tests
DoWhatWorks started about five years ago with two ex-Meetup executives who’d run a lot of experiments and built technology to find split tests across the open web — the tell-tale public URLs (site.com/a, /b, /c) — then use computer vision to see what’s being tested and track what companies keep in production. Nothing black-hat; it’s public information. The signal is the kept variant: if 2,854 B2B SaaS brands tested one versus two CTAs in the hero and two CTAs won 89% of the time, that’s worth knowing.
Two things make it more than a scraper. First, first-party data from hundreds of top SaaS customers — so they see what was statistically significant, on mobile, after the fact. Second, Casey interviews the teams behind the best sites (Lovable, Replit, Ahrefs, MongoDB, Glean) for his Substack, pairing qualitative signals with quantitative results. And the whole thing lives or dies on nuance: industry, company size, audience and traffic source change everything, which is why their new agent returns tailored recommendations with confidence intervals rather than one-size-fits-all “best practices.”
The pre-launch sanity check
Before you test anything, Casey runs three checks. First, understand your traffic — source, device and, above all, intent. If 80% of your traffic is retargeting ads on mobile, the visitor’s mindset is completely different than cold search. His Bonjoro example: whether to sell “video email” as a category or “why Bonjoro beats Loom” depends entirely on whether visitors already know the category exists.
Second, test toward a single clear problem. The best CRO teams doing a revamp ladder every change up to one objective — a “unified variable set.” When Fin (formerly Intercom) overhauled its pricing page, every change (fewer plans, simpler colors, less copy) tied back to one word: simplicity. Third, and biggest: test something meaningful. Optimizely found 89% of variants “lose,” but “lose” mostly means neutral — the test just didn’t move the needle. Ask whether a variant actually introduces new information or new value; “get started” vs. “get free trial” at least adds the word “free.”
Make every test specific
Specificity is the through-line of what converts. Pranav’s own Dropbox story illustrates the trap: changing “start free trial” to “try it free” lifted trial starts 17% — but those signups converted far worse, because the copy implied the whole product was free. The lesson is to pair expectation with reality and never optimize click-through in isolation; tie every test back to the real goal.
From there, the winning patterns are all about concreteness. Reassurance subtext under a CTA that removes a real barrier — “no credit card required,” Kit’s “free migrations,” Zapier’s SOC 2 messaging — works because it adds new information. Mercury ran ~15 tests leaning into hard numbers (“zero minimums, 1.5% cash back, 3.8% yield”) and got more engagement and recall than generic benefits. And on AI positioning specifically, the data is blunt: “deploy and orchestrate fleets of specialized agents” means nothing, while Docket simply saying it handles “qualification, discovery and booking” lands. Connect your AI to a capability, not a vague outcome — a reversal of the old “sell benefits, not features” advice.
Let users self-select
Personalization has been promised for 15 years, but Casey’s data favors a humbler version: let people choose. Sage lets visitors pick company size and industry and tailors the page; MongoDB’s experience selector sends developers to documentation and business leaders straight to pricing. As brands scale and chase a wider ICP, a single home page can’t speak to everyone — and self-selection has the bonus of AEO utility from all those tailored pages.
He contrasts this with dynamic, enrichment-based personalization (the RB2B/Clay/Apollo/Warmly world), which can work but is much harder to do well and more error-prone — misidentify a visitor and you’ve served the wrong page. Self-selection is the more conservative bet, and it turns a passive scroller into an engaged participant who has curated their own experience.
Why logos lose: proof in an age of fakes
One of Casey’s biggest surprises on joining: logos lose in A/B testing at a high rate, despite “plaster customer proof everywhere” being career dogma. His litmus test for social proof now is simple — how easy is this to fake? A row of grayed-out, non-clickable logos tells a visitor nothing; you can’t tell a two-person branch from the parent conglomerate.
The fix is proof that’s hard to fake and easy to validate: interactive logos that link to real case studies (the signal helps even if nobody clicks), video testimonials (a named exec on camera reads as real), and specific, textured context — “it took three months to deploy; our demo flow is up 32%; we hired two AEs for the inbound.” That messiness is more believable than “set up in a day,” especially in enterprise. And specific framing beats scale: Gorgias’ “used by 47% of Shopify Plus” and Asana’s “used by 81% of the Fortune 100” outperform “trusted by 300,000.”
Benchmarks, confidence, and earning trust
Marketers love to ask for the benchmark, but Casey’s point is that a benchmark means nothing without tailoring — and once you tailor it to your industry and situation, you might only have three data points. DoWhatWorks’ answer is confidence intervals: 4,000 single-variable tests on one thing yields real conviction; four multi-variable tests yields very little. There’s always noise (sunk-cost creative that ships anyway, first-party bias from teams who want their launch to look good), but more data means more confidence.
The deeper truth is that trust is earned by results, not percentages. Casey’s viral post about Clay attaching case studies to its logos drove the single biggest wave of enterprise interest since he joined — because people implemented it and saw it work. With colder prospects he leads with the highest-density findings (like two CTAs for sites with varied traffic sources) so they get a fast, credible win, and the relationship grows from there.
Are websites dead in the AI era?
With one camp declaring websites dead (“just optimize for agents”), Casey pushes back with evidence. Lovable and Replit launched deliberately minimalist, activation-first sites — then both relaunched this year with far more content, because a competitor with a comprehensive site (Bolt) was showing up more in LLMs despite weaker domain reputation. The reason was old-school SEO: the fuller site actually used the keywords (“no-code,” “vibe code”) that matched intent.
His takeaway: your website is still the hub LLMs draw from. When Glean and Framer added competitor comparisons to their sites, ChatGPT began citing them within 30 days — even noting the source might be biased, then using it anyway as context. So keywords, layout, header/footer and clear structure still matter. He points to Lightfield, an AI-native CRM that names the enemy, defines exactly what “AI-native” means, and links a clarifying article — content that’s useful for humans and, because it’s well-structured, especially useful for agents. The caution: don’t go radically minimalist just because you can build fast; there’s a transition period, and for now what’s on your site still shapes the narrative.
Quote snacks
• “The majority of tests just don’t move the needle one way or another — that’s what ‘lose’ usually means.”
• “A benchmark means nothing if it’s not super tailored to your industry and situation.”
• “Don’t just say your agents drive revenue — that’s fluff. Tie it to a specific capability.”
• “A good litmus test for social proof is: how easy is this thing to fake?”
• “Building has never been easier — knowing what to build, how to think about it, and how to measure it, those are the questions now.”
• “Your core website still matters — within 30 days, these companies show up as an LLM citation.”
Why it matters
This is the brand-versus-performance conversation zoomed all the way into the website — the place where media spend either converts or doesn’t. Casey’s throughline is a corrective to both hype and habit: most tests do nothing, most social proof is ignorable, and most AI copy is fluff, because they all fail the same test — do they give a real human (or agent) something specific and new to act on?
It’s also a measured take on the AI panic. Building is cheap now, so judgment is the moat: knowing what to test, tailoring to real intent, demanding enough data before you trust a number, and keeping a substantive website even as LLMs reshape discovery. Do what works — but prove it, and stay specific.
Practical next steps
Start with intent. Map where your traffic comes from and what visitors already believe, and design the page for that mindset.
Ladder every change to one problem. On a revamp, make sure each variable boils up to a single clear objective rather than a grab-bag of opinions.
Only test meaningful changes. Ask whether a variant introduces new information or new value; skip cosmetic copy swaps that can’t move the needle.
Get specific everywhere. Replace generic benefits with concrete numbers and capabilities, and add reassurance subtext that removes a real barrier.
Rethink social proof. Favor interactive case studies, video and specific stats (“used by X% of Shopify Plus”) over grayed-out logo bars.
Keep a substantive website. Use clear structure, real keywords and comparison content so both humans and LLMs can draw from your site — and validate with confidence intervals, not benchmarks.
More sharp conversations
The sharpest marketing conversations, straight to your inbox
By providing your contact info, you agree to receive communications from Paramark. You can opt-out at any time. For details, refer to our Privacy Policy



