The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Marketing/Brandformance
Brandformance artwork

Why most A/B tests fail - and what actually improves conversion rates

Brandformance · 2026-08-17 · 42 min

0:00--:--

Key moments - from our scoring

Substance score

66 / 100

Five dimensions, 20 points each

Insight Density14 / 20
Originality11 / 20
Guest Caliber15 / 20
Specificity & Evidence13 / 20
Conversational Craft13 / 20

Casey Hill discusses why most A/B tests produce neutral results rather than improvements, and outlines a framework for designing experiments that matter. His platform, Do What Works, monitors public split tests across websites using computer vision to identify patterns - finding, for example, that 89% of B2B SaaS companies using dual CTAs saw improvements. The core problem isn't testing itself but testing the wrong things: copy tweaks that introduce no new information, generic value propositions, and changes that don't address a single clear problem. Hill walks through real examples from HubSpot, Fin (formerly Intercom), Zendesk, Mercury, and Zapier to show what works - reassurance subtext like "no credit card required" or "free migrations," specificity over generic benefits (e.g., "zero minimums, 1.5% cash back" beats "optimize workflows"), and personalization that tailors pages by company size or industry. The conversation also touches on how traffic sources and user intent shape testing priorities, and why enterprise products need different metrics than simple conversion rate.

Key takeaways

  • →Tests fail because they lack meaningful differentiation - copy tweaks and generic value propositions don't introduce new information, so they move the needle nowhere rather than down.
  • →Understanding where your traffic comes from and the user's existing knowledge (are they sold on the category or do you need to sell that first?) is the foundation for prioritizing what to test.
  • →Reassurance subtext that removes specific barriers - "no credit card required," "free migrations," "SOC2 compliance" - drives measurable improvements when it addresses real objections.
  • →Specificity in value propositions (e.g., "zero minimums, 1.5% cash back") outperforms generic benefits claims that every competitor can make.
  • →Everything in a test should ladder up to one clear problem; unified variable sets (like Fin's pricing page overhaul focused entirely on simplicity) avoid confounding results.

Guests

Casey Hill

Topics in this episode

A/B testing frameworkCRO (Conversion Rate Optimization)Do What WorksUnified variable sets (UVS)Reassurance subtext on CTAsSpecificity in value propositionsPersonalization by company size and industryComputer vision for web testing detectionSplit test monitoringMercury fintech

Questions this episode answers

Why do 89% of A/B test variants lose if most companies are testing?

Loss doesn't mean variants perform worse - it means they're neutral and don't move the needle. Most tests fail because they lack meaningful differentiation, like testing "book a demo" vs. "get a demo," which introduces no new information to the user.

How do I know if my test is actually meaningful?

Ask whether you're introducing new information or new value to the user. Testing specificity ("1.5% cash back") over generic claims ("optimize workflows") introduces real signals; simple copy swaps or feel-good branding claims do not.

What's the difference between understanding traffic and designing a test?

Understanding traffic source and user intent shapes your entire testing strategy - retargeting mobile users have different intent than organic visitors, so your test priorities (differentiation vs. category education) should reflect that.

Do What Works monitors which companies' A/B tests and how?

The platform searches the public web for split test variants (like HubSpot.com variant URLs) using computer vision to detect differences, then tracks which variant wins and stays in production; it also combines this with first-party data from customer companies and interviews with top SaaS leaders.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

14 / 20

The episode delivers solid, actionable frameworks for A/B testing with concrete examples (Zendesk's language specificity, Mercury's numerical precision, Sage's personalization selector), but relies heavily on surface-level pattern-spotting from observed test results rather than explaining *why* these patterns work or exploring deeper mechanisms. The guest repeats several intuitive ideas (test meaningful things, avoid generic copy, add specificity) without novel depth that would surprise an experienced operator.

the majority of tests, they just don't move the needle one way or another
Specificity matters. You can absolutely talk about your agent, but tie it to a specific capability set. Don't just tie it to an outcome

Originality

11 / 20

The core thesis - that generic AI positioning underperforms and specificity wins - is sensible but not novel. The observation about logos losing tests contradicts conventional wisdom slightly, but the guest doesn't provide a rigorous mechanism or systematic analysis; it's anecdotal. The analysis of LLM-driven future relevance of websites is timely but speculative. Most frameworks (reassurance subtext, unified variable sets) are established CRO principles repackaged.

we're going to move more in the future towards this LLM style
instead of just saying, we work with this company and we got a result, I tell people, take a step back, add those couple extra sentences of context

Guest Caliber

15 / 20

Casey Hill is CMO of a relevant platform (Do What Works) with real data access and direct relationships with practitioners at scale (Replit, Ahrefs, MongoDB, Dropbox). His background spans 15 years in SaaS, includes an acquisition (TechValidate → SurveyMonkey), consulting at tier-one firms, and teaching at Stanford. However, he's primarily a data analyst and consultant rather than an operator who scaled conversion at a single company to significant magnitude; his most concrete operational experience cited is older (Bonjoro, ActiveCampaign).

I've had the privilege to talk to lovable Replit, Ahrefs, MongoDB, glean tons of the top B2B SaaS companies
I've been in the software industry for 15 years. My first company, Tech Validate, was acquired by Survey Monkey

Specificity & Evidence

13 / 20

The episode includes named companies and specific test results (Clay's case studies, Sage's industry selector, Zendesk's 140 languages, 68% resolution rate, Mercury's 1.5% cash back and 3.8% yield), but most examples lack critical detail: win rates, lift magnitude, sample size, duration, and statistical confidence. Many claims are qualitative observations ('we see people finding success') without attached metrics. The MongoDB and Dropbox anecdotes from the host add color but limited novelty.

they said, you know, we have 24.7 coverage. We speak in 140 languages. We have a 68% average resolute
Mercury...they actually have run a ton of tests. I think they ran like 15 tests in the last 12 months

Conversational Craft

13 / 20

The host (Speaker B) asks solid clarification questions and pushes back thoughtfully (the dynamic vs. reactive personalization question, the Dropbox trial anecdote adding nuance). However, follow-ups are often brief and don't excavate disagreement or contradiction. The guest's long monologues go largely unchallenged; the host affirms ideas ('I have an immediate idea for paramount.com') rather than stress-testing claims. There's genuine rapport but limited intellectual friction.

I have a clarification question. So when your platform looks at, uh, let's take the HubSpot example, they just completed a test
What's your take? Like I don't know if you have like any test type of results, but like, what's your personal take on like a um, more proactive approach to personalization versus actually a reactive approach based on actual human input

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A80%
  • Speaker B20%

Most-used words

website22different21important18data17first17back17test17tests16testing16example16information14page13works12interesting12point12specific12

Episode notes

Why do most A/B tests fail to produce any meaningful improvement - and what separates a useful experiment from another inconclusive result? In this episode of Brandformance, Pranav Piyush speaks with Casey Hill, CMO at Do What Works, about what thousands of real-world website experiments reveal about conversion rate optimization. They explore: Why most A/B test variants fail to move the needle The three checks to make before launching an experiment Why specific messaging consistently beats generic benefits How reassurance copy can remove conversion barriers What marketers misunderstand about benchmarks and confidence Why customer logos alone may no longer build trust How to position AI products without relying on empty language Whether websites still matter in a world of AI agents and LLMs For CMOs, growth leaders and B2B marketers looking to make smarter website and experimentation decisions, this conversation offers practical lessons grounded in actual testing data.

Full transcript

42 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: I've had the privilege to talk to lovable Replit, Ahrefs, MongoDB, glean tons of the top B2B SaaS companies and sit down and say, what's working on your site.

Speaker B: A benchmark means nothing if it's not super tailored to your industry and situation.

Speaker A: The majority of tests, they just don't move the needle one way or another, where the data tells a story that's gotta be told. It's the art and the science getting the numbers right.

Speaker B: Hey, everyone, and welcome to another episode of brandformance. Today is a special session. We are reviewing the office hours that we did live with Casey Hill on all topics around conversion rate optimization. Casey is the CMO of Do what Works. It's a platform that monitors all the AB tests that lots of brands do across all types of industries, from SaaS to E Comm to FinTech, and then comes up with recommendations that are tailored to your brand that based on what others are doing. So let's dive straight in. How are you doing, Kasey?

Speaker A: I'm doing great. I'm doing great. Uh, it's been, uh, following a lot of World cup, you know, it's been fun.

Speaker B: Fantastic. Well, here we are. We're going to kick it off. Thank you for joining me on these office hours. Um, this is an interactive session we do. Kasey, I've been following you on LinkedIn for, I don't know, like three years, something crazy like that. And I really appreciate the work that you all do at do what Works. I love the name, it's just right there, do what Works. Um, and you know, we have a little bit of a adjacent worldview, right? Like we at Paramark do a whole bunch of media measurements. So this is like out there, you know, with TV and out of home and meta and LinkedIn and what have you. And you all then sort of come in at the website, um, and sort of do a whole bunch of experimentation there. So I'm sure we'll have a rich conversation and debate about all the things. But thank you for joining me. This is going to go fantastic. Fantastic.

Speaker A: Yeah, I'm looking forward to it. It's been awesome. Being a Do It Works and having access to some of the data has been kind of eye opening and so really excited to share some of the learnings with folks today.

Speaker B: Fantastic. Before we do that, let's talk about you. So you have a really interesting background. You want to give us a little bit of a, uh, rundown of, like, how you ended up at do what Works and what what got you sort of to. To even join the company?

Speaker A: Yeah, for sure. So I'll be kind of quick. I've been in the software industry for 15 years. My first company, Tech Validate, was acquired by Survey Monkey. I started all the way back at Berkeley when I was going to college and have kind of since worked primarily in the Martech space. Done a good amount of consulting with McKinsey, BlackRock, Coleman's, a bunch of different folks in that kind of domain have done some guest lecturing with Stanford Business 112 Modern Approaches to Creating Customer Demand. And really, one of the main avenues that I share information is across LinkedIn. I actually cut my teeth on Quora, funny enough, when it was back in 2016 and 2017, when it was kind of a more relevant channel. I actually started in sales before marketing, and, and I used Quora actually as a way to drive business. So this was, you know, me as a hungry AE back in the day, trying to find new ways to draw up business. And then I was having so much fun writing, I was like, oh, I kind of like the storytelling aspect of this. And that's what eventually led me to run Growth, which was kind of a hybrid of marketing and sales back at Bonjoro. From there I took a, uh, job with a large organization, activecampaign, and finally ended up as CMO here at do what Works. So that's a little bit of the kind of career journey.

Speaker B: Fantastic. Okay, so let's talk about sort of what we're going to cover today. I'm actually going to just hand over to you, walk us through what you've prepared, and then I will pepper you with questions as I have them.

Speaker A: Yeah, that sounds wonderful. So I think what's really important anytime we talk about testing is to acknowledge that there's a ton of nuance in it. Right. There's so many different variables that go in. So I'm going to try to share some best practices, some ways of kind of thinking about testing. But I think it's really important to preface that with just how should you even think about experimentation in the first place? Especially now in an age of AI, we're at a very interesting inflection point where building has never been easier. Right. People can build incredibly fast with the tooling that we provided, but knowing what to build, knowing how to think about, knowing how to measure, these become these really important questions that are kind of connected. And so I, I'd love to start out by having kind of a conversation around the litmus test that I use to Think about whether something is a. Is a good experiment, an impactful experiment. Um, and then I'll share some major examples, some different trends I'm seeing, and we can kind of open it up to questions from there. So that's kind of the agenda that I have.

Speaker B: Sounds great. Let's do this. And I mean, again, I have to give kudos to whoever came up with, like, the name do what Works. Tell us a little bit about how did you even come to this point of view? You guys have some proprietary data, if I'm not wrong. Like, tell us about of how the platform works and what powers the point of views that you have developed.

Speaker A: Yeah, so about five years ago, we had two of our founders who were executives at meetup, who had basically exited and they were running a lot of experiments there. And essentially they had come across this ability, this technology that we developed that allows them to search across the entire web and see where people are running split tests. And so it's really basic, actually. Like HubSpot just wrapped up like a split test where you could see it's like HubSpot.com A, B, C. Right. So it's basically public information. So we're not doing anything black hat. We're not infiltrating people's websites. We're looking at public information. And then we're basically using computer vision technology to essentially look at the differences, what is being tested, and then tracking what people keep. So, for example, if we see 2854 B2B SaaS brands on their homepages, tested 1 versus 2 CTA buttons in that hero section at the top, we could say, okay, that's interesting. 89% of the time, 2 CTA is won. So that is a signal. And we can use those signals to try to help people make better testing decisions, prioritize what to test, think about the ordering. And as I kind of said in the very beginning, the nuance here is everything. So industry, company size, audience type, traffic source, like all of these variables inform our best practice recommendations. And what works in B2B SaaS might be completely different, different than what works in streaming or banking or travel. So everything is broken down with a lot of granularity to kind of think about that problem. And we actually just recently are introducing an agent that allows people to kind of talk in layman's terms, hey, I'm thinking about this website experiment. Is this a good idea? And we'll process and give that intelligence automatically through that, through that agent. So that's essentially the lens or the data source. That is informing some of the stuff we're going to talk about today.

Speaker B: I have a clarification question. So when your platform looks at, uh, let's take the HubSpot example, they just completed a test. So because you're monitoring it on an ongoing basis, you're essentially making the decision that they had three variants and then one month later they just have one variant left, so that one variant must be the winner. Is that how I should think about it?

Speaker A: Yeah. Correct. So we look at what people put into production. It's a little bit more complicated for a couple reasons. Number one is we work with hundreds of the top SaaS. And so one of the benefits is we're not only getting third party data, we're actually getting first party data too.

Speaker B: Nice.

Speaker A: So we see like, hey, Salesforce ran a test. Hey, this company ran a test. And we see what happens afterwards on mobile, on whatever, what is statistically significant. So we get this really great combination between the two. Also, because we're really active on social media, I interview a lot of the top teams. So even people that aren't customers, I've had the privilege to talk to lovable REPLIT hrefs MongoDB, glean tons of the top B2B SaaS companies and sit down and say, what's working on your site? What's converting the best, what is the most impactful test you did this month? And I published those to my substack. So it's kind of a combination of, um, looking at some of these qualitative signals, right? Saying, okay, what do they keep in production? That is a signal to push us in the direction. But then let's look at the quantitative data behind the scenes and let's use that to strengthen our confidence behind certain recommendations.

Speaker B: Okay. It makes a ton of sense. All right, let's talk about. I thought this slide was great. So let's, let's talk about what do you. What is the sanity check that you do before you even launch something?

Speaker A: Yeah. Awesome. So first thing we do is we want to understand traffic. This is incredibly important. I was chatting with a large organization at scale company yesterday and they said 80% of our traffic comes from paid ads and the vast majority is on mobile. And so that was a huge piece of information. Okay. They're doing ads, a lot of them, retargeting ads, and they're doing it on mobile. It's a very, very specific type of intent and understanding that the user has when they hit the website. And so one of the examples I use that I think really drills this home is. Back in my Bonjoro days where it was a video email tool, we had to make a decision. Are we selling video email as a category you should be doing video email or are we selling why Bonjoro is the best choice for video email vs. Loom vs. Vidyard vs. All the names people are familiar with. And that question is fundamentally answered by when people arrive at the site, what is their understanding if they're already sold on video email? Right. Then it doesn't make sense for us to spend all that time. We need to focus on the differentiation. But if they're not even familiar with that, if their focus is on the use case, the output, we're just trying to increase our onboarding success. We're trying to increase activation. I need to sell video emails to channel. So this I think is really, really important and something we think about a lot when we're trying to make best practice recommendations. Where's your traffic coming from? The second point is are you testing towards a single clear problem? This is a huge miss. We find a lot of times when we see website revamps, people change a million variables. They do a complete overhaul, often based on a lot of very subjective opinions across the organization and whatnot. We don't have time to unpack all the challenges of expensive website revamps. But I think one of the things that the best CRO folk conversion rate optimization folks do when they're doing an overhaul is everything boils down to a clear problem that everything addresses. So a great example is fin, formerly Intercom. And when they did a pricing page revamp, one of the pieces of feedback they had is that this was too complex, it was too busy. They had a modular thing at the top that said, you know, x dollars per transaction from fin. And they had four different plans and a bunch of color schemes. And so they made all these changes. They completely overhauled it. If you go to the page today, it's like two big plans. It's black instead of white. Like I made all these shifts, but if you go through and you look at all the changes they made, one of the things you'll find is everything tied back to simplicity. Simpler colors, less plans, less copy, less like. So everything was going towards and laddering up to that single clear objective. Some of them call it like a unified variable set, a uvs. Is everything boiling up to that? So that's a really important part. The third piece that I think is super, super important is are you actually testing something meaningful? This is probably one of the biggest problems with testing, people get in there, they're like, I'm going to trust 20 different versions of my headline, I'm going to test 20 different versions of my CTA button. And then they say, nothing moves the needle. It doesn't change. A lot of companies, especially startups, one of the pieces of feedback I get is we just stopped A B testing because it didn't matter. We kept AB testing and just wasn't significant. Right. And so it's funny, Optimizely published a study that said 89% of variants in A B tests lose. But what I think often gets lost in the nuance of lose is lose doesn't mean that it does way worse than the control. Lose can just mean it does nothing, it's neutral. And so this is what happens a lot. Like the majority of tests, they just don't move the needle one way or another. It's not like they cause some huge drop. Right? So when you design a test, I think it's really important to ask yourself these questions of are you introducing new information or are you providing new value? And I'm going to show a couple examples in the decks and the slides that follow of this, but I think it's really, really important. So first is, if you're just doing some basic copy tweak, book a demo, get a demo. I'm not learning anything different, right? So if you said, for example, get started or get free trial, get free trials a little bit more specific, you've said free. So I now have a new piece of information, like there's something new introduced in that dynamic. But this first example, I'm not getting anything new. And the same thing with this, the second one, there's so many different tests where people are trying to be clever, they're trying to get the brand, whatever, ethos across. Sometimes I genuinely feel like it is a play to investors where they want to show they have the biggest tam. So they are. We're the AI app for everything. We handle orchestration across all. And it's these very, very generic statements. But what we found, by the way, looking at a lot of tests around AI positioning is specificity matters. You can absolutely talk about your agent, but tie it to a specific capability set. Don't just tie it to an outcome. Don't just say that our agents drive revenue. That's too generic. That's just a generic statement. It's fluff. You just see that, you scroll past. But if you are like, I think this is one run by Zendesk where they said, you know, we have 24. 7 coverage. We speak in 140 languages. We have a 68% average resolute. Like they're giving me very tangible things that are connected. And as a team I can look at that, I can be like, okay, that actually seems really valuable to be able to have coverage because we have a bunch of people over in emea, we have a bunch of people in apac and this is going to, you can start to connect those dots. So, so I would m. Think about.

Speaker B: I have a, I have an interesting sort of story here. So I used to work at Dropbox like a lifetime ago and I was uh, kind of doing a whole bunch of conversion rate optimization on the growth team. And one of the first tests we launched, it was very surprising to me. It was the early days, it was like 2013. We were trying a whole bunch of different things. The product was very new, like Dropbox business was a new thing compared to the core Dropbox offering. And we had this trial, a 30 day trial. And the website call to action was start uh, free trial. Okay. And just for the heck of it, I changed it to try it free. And I didn't think like anything big will come out of it but like that resulted in a 17% increase in trial starts and, and, and then you know, we looked at it 30 days later and those didn't really convert. So the actual sort of conversion rate was closer to 4 or 5% higher than sort of the baseline. And I learned something meaningful because at the first instant I was like, oh my God, 17%, great. But it sounds like it was just a uh, hack in the sense that you went from start free trial which is like there is some information there that it is a time bound free trial versus try it free which might indicate that the whole product is free and that you know, even though it's like a very subtle difference, uh, it did have a material impact on the buyer psychology. So sometimes you really have to think deeply about the way you sort of use copy and these tests. But I completely agree with your. I'm not refuting what you have on the slide here, am I like adding a nuance there that sometimes it's like not obvious when you even design the test what the actual buyer psychology might be a hundred.

Speaker A: I mean, I think there's two things nested there. Number one is I think expectation to reality. So many things boil down to this which is like, is there a good pairing of those two things? Right. Like a lot of times when we're looking at improving CTA is it's like when I click this button, what do I think is going to happen? What actually happens and is there dissonance? And the example you shared there was obviously a little bit of a gap of expectation which led to the increased click. But it's kind of like having an email with a great subject line that gets tons of opens but doesn't actually drive conversions. The second thing you said, which I think is really important, is that you don't want to just look at click through rate, right? Because you can do things that drive up. Like we think about this all the time in signup flows, right? You can do things that jack up your amount of people that start that process. But if they're all poor quality leads, if they're not converting. So it's really important, whenever we talk about best practices again, to think back to what is ultimately your goal, what are you trying to accomplish? Conversions is maybe the goal in a lot of cases, but sometimes there'll be other metrics. Sometimes, especially with enterprise products, there's education goals. You want them to click or play through a thing. Like it's not always a direct route. So definitely important. This is another one I thought was was good. So one of the things we've seen work well in testing is reassurance subtext. We've seen quite a few people that have found that when you put that little text below a CTA that removes specific barriers, um, that helps the most common ones can be something like no credit card required is commonly mentioned. But I thought Kit did something kind of clever here with free migrations, so they understood. And I've been in this Martech world for a long time, so I also understand that migration is one of the biggest barriers. Right? When you're using HubSpot and you're thinking about going to some other tool, the biggest thing is like the migration is a pain in the neck. So Kit was smart, they realized that and they said, let's call that out up front. Zapier did a website overhaul recently and they're trying to move more enterprise and so they use reassurance subtext tied to SOC2 compliance tied to those specific protocols. And that also has been effective for them talking with the Zapier team. And so I think there, there's value here. And again it goes back to meaningful test. They're running a test saying, hey, what if we just said no credit card required? Or what if we added free migrations? Does the free migrations help? And I think that checks the meaningful box because you're introducing a new piece of information for the individual. So no surprise that they've kind of kept this in place.

Speaker B: I have an immediate idea for paramount.com, so. I love this. Fantastic.

Speaker A: Awesome. Uh, that's always my goal, by the way, whenever I do these, when I hop on podcasts or workshops or webinars, I want to give people a couple things they can think in their head. Hey, maybe that applies. Go, go. Put it into action. So that's what I love to hear. This one is also something I think is really valuable, which is the importance of specifics. So this kind of goes back a little bit to what I was talking about with subheaders and headlines, but so many times. And this is coming from someone who looks at websites all day. Like, I look at so many websites. I can't tell you guys, but one of the biggest gaps is just general language. Uh, you see so much general language, and it's gotten way worse in the AI era, where you're just not giving people something really material to link onto. So this is an example from a fintech company, Mercury, and they actually have run a ton of tests. I think they ran like 15 tests in the last 12 months around different product sections and around the language, both. Both in kind of the top header as well as that subhead. And one of the things they really leaned into is specificity. Right? And so zero minimums, 1.5% cash back, 3.8% yield on their Treasuries. What they found is they were getting way more engagement, way better recall when they just started being way more specific with those actual value props. They realize that even if they're not directly driving a click at that point in the page, they're signaling the key pieces of information that keep someone moving along. So the takeaway I think there is to really encourage folks to go back to your product sections and ask yourself, are you using language that is just saying, you know, hey, optimize campaigns, optimize workflows to. To increase, you know, reduce time, increase. Like I call those generic benefits. So I call all those statements, generate more leads, generate more revenue. Like they're generic benefits. Every single company essentially can claim those things. So try to drill down, give some differentiation, give some signal to the exact thing that you're delivering. And that has a lot of utility. So, um, no surprise to see brands testing into that.

Speaker B: That's a great.

Speaker A: This next one. Yeah, this next one is something. Another trend that I've seen, which is this idea of personalization. I know we've Been talking about personalization for 15 years. Since I started in SaaS, they've been talking about personalization. But there's a couple interesting tests that some larger brands have kind of rolled into that I think are just again bringing people back to this fact. So this one is from Sage and Sage essentially allows people to select their company size and industry and it then tailors the page. I actually experimented and tried this. I like selected uh, SaaS and selected startup and the page that they served up actually was quite good. And so I think there's this utility that says especially as brands grow, especially as brands scale, they often are trying to serve a wider and wider icp. And so you making your website really, really tailored I think has a lot of valuable value and utility to taking people through. When you're this multi product suite and you're trying to speak to everyone on a homepage, it's very, very difficult. Right. And there's also I think kind of AEO utility to even just building all those individual pages in the first place and having those very tailored experiences that speak to those exact use cases that have the verbiage and all the other pieces that are tied together. So it was cool to see them execute this and that it was successful for them.

Speaker B: I have a question for you on this. There was a whole trend three or four years ago where you would dynamically create your page based on intent signals and we can talk about that. This one is kind of taking a slightly opposite point of view, which is like, hey user, you tell me how you want to customize this and I will go do that. What's your take? Like I don't know if you have like any test type of results, but like, what's your personal take on like a um, more proactive approach to personalization versus actually a reactive approach based on actual human input.

Speaker A: Yeah, I think the reality is that it's a lot harder to do it well than people. Like I work with a lot of the teams, I've talked to RB2B, I work closely with Clay, I work closely with Apollo, I work closely with warmly like I work with tons of people that are in this world. Right. And so the idea behind those is kind of that data enrichment angle.

Speaker B: Right.

Speaker A: People come through, you have uh, data enrichment, you tie it in execution. And first off, I want to say it's not that there isn't necessarily some utility to that and there are people that are serving up personalized experiences dynamically that probably have some success. But I think that right now at the current stage that we exist in. Allowing people to self select is probably the more conservative, less error prone way to approach it. Because you can imagine like if I land on this page and it for whatever reason misidentifies and sends me to the E commerce, uh, personalized page, like that's a pretty bad experience. And that kind of just like puts me versus me actively taking an action, which also is a level of engagement on my part. Right. So my. By me choosing to do this thing, I'm now more engaged. I'm not just a passive scroller on a website. I've chosen to provide this personalized information and curate my experience. So I think there's something to both the intent side as well as the risk reduction that still makes something like this. And I think I might share also an example from MongoDB which is similar with the experience selector. I still think there's some utility to that, or maybe not. Uh, but MongoDB does the same thing.

Speaker B: Awesome. No, that's great. And another idea for us, because we have a similar issue where we sell to B2B SaaS, we sell to FinTech, we sell to E Com and consumer companies and retail companies. And it's hard to talk to like such a wide audience with just one page. So I think this is very clever. Hey folks, thanks for listening to this podcast today. If you're enjoying the show and if you're getting value out of it, we'd really appreciate if you'd drop us a five star rating on your favorite podcasting app. All right, well, let's talk about AI. This is obviously sort of top of mind for pretty much everybody who will be watching this, whether you're here live or you're listening to it after the fact. Yeah, I'm very curious to hear your take on this.

Speaker A: Yeah, for sure. So the story behind this, the genesis of this story, is our research team was compiling a report for a major Fortune 500 company. And it was around AI positioning. This company was like doing some overhauls. They were like, I want you to look at all the testing, all the information around headlines. And one of the things that they had done as part of this report was to go look at top AI companies and just see what were the headlines, what were the subheaders. And I was blown away as I was looking through the research at how insanely generic.

Speaker B: Right.

Speaker A: Deploy and orchestrate fleets of specialized agents that work with your team. Like so many of these things gave me absolutely no clarity as to what this thing fundamentally did. And I think it was like number eight or nine on the list where I got to this, this company called Docket that just like said what they did, hey, we handle qualification, discovery and booking and you know, basically gave me that, that information. And so I think that one of the takeaways there is again to connect your AI to the capability. I know that a while back in marketing they kind of took the ops attack. They said, don't talk about features, it's not about features, you need to talk about the benefits. Right. Well I will tell you from looking at the data and seeing what people keep and seeing what is converting and seeing what is driving engagement, it appears to be the capabilities, it appears to be the specifics of what you can provide that is driving the best engagement. So I encourage brands to try to embrace that.

Speaker B: AI does that does everything. That's the, that's the headline.

Speaker A: Yeah, exactly. And uh, this is a specific, I think it's like a product page or a sales page that they were uh, driving some Traffic to for GoDaddy for their new arrow. But again, this is just driving home the point, right? Instead of saying buy a domain and experience the magic of AI, what the heck does that mean? Right. They had more success with instantly build your website, your logo, the actual components of the thing. So I don't think we need to belabor this one. I think it kind of makes intuitive sense, uh, being more direct. Oh, and this is the MongoDB example again. They had an experience selector where you could choose between developer and when you chose developer, it showed you documentation. It was very kind of technical in its layout. And when you chose business leader instead of documentation and went straight to pricing, it went to the things that an executive at the organization would be more concerned with. So talking with the Mongo team and seeing their success here, I thought this was really cool to see. And by the way, like I had a lot of people, when I did some posting about this while back, a lot of people came in, they said, uh, you know, I don't think this is a good idea. People aren't going to see this. This is not going to be, you know, there was a lot of pushback. And what I tell people is like, look, what we do, right, is we try to look at the data trends, we try to understand what companies are testing, we try to look at those results, we then try to talk to those teams and get their firsthand results. There's always going to be noise in these subsets. There's always going to be, hey, this team invested a ton of money in this video so they're going to put it up anyways, even if it sucks. Uh, they spent $30,000 on it, they're going to put it up. Uh, right. And there's also noise, even first party noise. I talk to a team who implements a new talkdoor agent. They want people to perceive that as successful. So can I validate every time. Am I inside their looker dashboard seeing the back end? No. So there obviously is a margin of error, but the more people you add to the data set, the more conviction we get. And that's, I think, a really important part of what we do is, and our new agent that we're released as well, we give confidence intervals. So we say, looking at all these data trends, we have this percentage confidence interval based on all these factors, based on all of these components. And that helps kind of understand. If we only have four tests that have been done on a thing, our confidence intervals are going to be fairly low. And if those weren't single variable tests, they were multiple variable tests, then it's way lower. Whereas if we have 4,000 single variable tests on a very specific thing, we're going to have a lot of conviction in that thing because we have a lot of data to pull from.

Speaker B: So this is an interesting point. Like when you have that conversation, we both, you know, our audiences are marketers, CMOs, VPs of marketing, what have you. How well do they understand the statistics and even this idea of a conference interval? Because I see a lot of marketers being like, hey, what's the benchmark? Like, can you share the benchmarks with me? And then, you know, to your point, a benchmark means nothing if it's not super tailored to your industry and situation. And then if you tailor it to your industry and situation, you only have three data points. Is that really a benchmark? Right. So how do you navigate that conversation? Like, how do you help people sort of have that, have that perspective when they're, uh, thinking about benchmarks?

Speaker A: Yeah, no, it's a really great question. I mean, I think there's kind of two ways I'll approach, I'll approach it. The first is I'll say that the best possible thing I can do is to have someone try a thing, see a result, and have that trust build up front. So, like a funny example is I started a series that was really successful looking at unique things that websites were doing. One of the first viral posts I did was about clay, who attached case studies to their logos. At the time, basically nobody was doing this. Now if you look so many companies are doing it. But what's crazy is we had so many new enterprise logos that reached out to me and said, kasey, we did this thing, we're driving tons of click through, like, this is awesome. What else can you do for us? And it was a huge source, probably since I joined, the single most impactful from a business revenue standpoint. But I think it was an important lesson to say, why was that so successful? Because people implemented a thing. I earned their trust up front. Right. And then, uh, then from there, um, that kind of carried over. What I often do with people that are coming in a little colder is I try to focus on those things that we have the highest density of data around. So, like a perfect example is I'll sit down with people, I'll learn about their business a little bit. A lot of people benefit from two CTAs. Not everyone benefits from two CTAs. If you have a very monolithic source of where your traffic is coming from, there's a very specific user path. Not always, but for people that are a little bit larger, that draw traffic from a lot of different sources, they're going to have variable levels of intent. And serving those variable levels of intent will provide a lift at a pretty high confidence interval. So that's an example of something where I'll say, hey, let's do this test, right? I put that in place, they get a win within that first month, and again, I've earned that confidence. So I think it's tricky. There isn't necessarily universal benchmarks. The percentages don't necessarily. What does a 72% versus a 91% exactly mean? I think it takes a little bit of experience and interaction to earn that trust. And ultimately our results. I mean, I'm sure it's the same in your business, like ultimate. Your results are what builds trust. Someone takes the dive, totally. They start seeing that you're driving them impact packed and then they're hooked and you, you kind of expand and grow with them.

Speaker B: Makes a lot of sense. Fantastic.

Speaker A: Yep. So this is one that I think has always been interesting and it's been funny and almost kind of controversial, I think a little bit for me. So when I first joined in, there was a couple things that really surprised me, things that I did not expect. One of those things was I saw all these different brands, including Dropbox, where you used to work, that were testing logos in different places on the homepage, on the pricing page, all different things. And I kept seeing logos lose and I was like, that's not hot. Like, this is like one of the universal things that we all learn, uh, is that customer proof is important. Put customer proof across your website as much as possible. That is what I had been habituated my entire career in doing. And if you look at almost all the top software companies and B2B SaaS and FinTech, they have customer proof all over. What I learned from spending a lot of time doing a lot of deep dives is that we're in a little bit of a new age when it comes to customer proof. And I think that a good litmus test for people is how easy is this thing to fake, right? People can go on their website, they can throw up a bunch of logos partially grayed out sometimes, that have no ability for someone to click, engage, learn anything from. And someone doesn't know. Do you work with State Farm, the massive conglomerate, or did you work with one small branch of two people in North Dakota? And now you're saying that, like, State Farm is one of your customers? I think that there's a lot of nuance. And so what we're taking away from that is, number one, if you're going to talk about customers, the more interactive you make it, the better you allow people to validate, even if they don't click through, by the way. So I'll say this, even if people don't click through to your OpenAI case study, but if they see that you're linking directly to an OpenAI case study, there's a signal that exists within that. I think that is valuable. The next thing is something like video. Video feels very hard to fake, right? Not that people at some point can't do it, but if you have the CPO of OpenAI talking about your product, that's a pretty strong signal, right? Again, whether people click on it or not, they're going to assume that that's probably an actual customer that you have. So we see people finding success with these harder to fake things with case studies, with video testimonials, with things that have a high degree of confidence. So instead of just saying, we work with this company and we got a result, I tell people, take a step back, add those couple extra sentences of context. Hey, it took us three months to fully deploy this. In the first two months, this, this, and this happened. But now where we're at is our demo flow is up by 32%. We actually hired two new AES to deal with this inbound flow, and yada, yada, yada, when I read that, I'm like, oh, okay, that actually sounds like Something that actually happened.

Speaker B: Yeah, it's kind of like, you know, people, people assume that it has to be like really buttoned up and perfect from the very beginning. Where at least in B2B SaaS, if you're selling something that's like high value, showing that it was actually hard can sometimes be more believable that oh yeah, it's not click to click and like I solved world hunger. Um, it's. It's hard. That's why you exist. And let's talk about that hard stuff.

Speaker A: That makes a lot of sense 100%. Especially in enterprise. Like enterprise takes a long time to implement. Right. So if you're an enterprise buyer, you're not dumb. Like if someone's like, oh yeah, in your first day you'll be fully set up on this like full suite thing that ties in with your er. Like people, people are just going to look at. Any sophisticated buyer is going to look at that and say that's nonsense.

Speaker B: Yeah.

Speaker A: Uh, so yeah, I agree with that. And this is just an example. It's funny actually, Clay just relaunched a website like two days ago, but they still kept the same premise where you linked out case studies. And I think that does have utility. One other thing I'll notice that we've seen testing from too. I think recently a test from Gorgis, which is an E commerce focused help desk tool. In this headline, instead of saying like trusted by 300,000 leaders of all sizes, I actually think that's kind of weak. What Gorgis did is they said used by 47% of Shopify plus or Shopify Enterprise or whatever their exact cohort is. That felt like a very specific signal. And we've seen companies like Asana and a handful of others that Asana actually tested completely out of logos. They remove them, then they re put them on. When they did their enterprise focus, but they said used by 81% of the Fortune 100 or something like that. Right. Very clear targeting of like this is who we're going after. And all the logos that we're going to include are these like massive companies. It was a very concerted focus of like this is for a very specific purpose that performs better. We see from testing makes a lot of sense.

Speaker B: Okay. This was an absolute masterclass. So thank you for going through all of this with me. I want to just hear your take on what's happening with AI in general. I don't think we can ever have a conversation without talking about AI these days. And specific to what you do and how you think about the world. There's one camp and this is especially in the VC ecosystem, websites are dead. There's no more websites. You just need to optimize for agents. And that, that's a camp. There's another camp which is like, you know, the webflow co, uh, founder just launched a whole startup called Ploy. I don't know if you heard of Ploy, they just went through IC and they're like taking a different approach, like no, no, no. Websites are very much alive. In fact we're going to double down and he started a new website, you know, platform called Ploy to make the whole process of, you know, running a website easier. And so I'm curious, like what's your take? Are websites dead? Are they not? Do you think about humans versus agents? Like do you have to do testing for agents? I'm just curious like if you've thought about this whole space.

Speaker A: Yes, I've thought about this a ton. So I'll start with a little bit of an interesting example that kind of ties in. So if people are familiar with some of the vibe coding tools, the lovables, the replets of the world, right? They launched with extremely minimalist sites. They basically. And I've, I've spent time with these teams, which has been amazing and kind of learn. And they were like, we're all about the activation. That's what's important. We want someone to hop in, activate right away. And they had these sites that had like two blocks on them, right? Like super, super basic. But something very interesting happened. Both those companies relaunched their websites this year and actually added tons of different content sections. And I was very curious about this and I kind of went a little bit investigative like what is, what is happening? One of the things I discovered was that I think they were called Bolt. I hope I'm not misremembering that. But there were another, there were another person in that space who had a much more comprehensive website, had been around for a while and they were showing up way more in LLMs, even though they had a lower domain reputation. Even though like all these other factors should have pointed to replit and these up M and lovable, uh, showing that more. But one of the things that was interesting is when you looked at their actual website, there was very little mention of no code of vibe, code of AI builder, like the actual terminology going back to old school SEO. Are you using the keywords? Right. And so when I talked to the team, they phrased it a little bit differently, but they essentially were saying the exact thing they Were saying we're expanding our market, so we need to make sure that we include verbiage that speaks to our larger icp. But translation of what was actually happening is they were editing their meta descriptions, they were editing their subheaders, all these different things. Adding in no code, adding in Vibe code, all these different pieces. Why do I mention this? I mentioned this to say that I think we are going to move more in the future towards this LLM style. Ask your intent. It's going to serve up dynamic content based on it. I think we're moving in that direction. But it's a cautionary note to say in this world where everything feels like you should just change everything right away. Right. Claude? And these tools have changed things so fast. People are like, I want to be on the cutting edge. Right. I'm just going to vibe code everything, Vibe in. And I think the cautionary note is there's going to be a transitionary period. Right now. It still matters what you put on your website. And we can see that evidenced by that example and many others that that I've kind of looked at that your website still serves as a hub to tee up information. It is still important as a source that is drawn from. And I've actually shared a lot of really interesting stories recently with people introducing our product versus Claude or our product versus ChatGPT Enterprise. And I shared one recently from Glean because I thought it was so interesting. Glean added these comparisons and I shared ChatGPT's thinking notes. You can actually see ChatGPT thinking. ChatGPT said, Hmm, I see that Glean has a comparison on their website to ChatGPT and to Claude. This might be a little bit biased because it's coming from the company, but I should include it anyways because it's important context. And then they cite it.

Speaker B: Wow.

Speaker A: Right. And I shared. I can share that post with you, but you can actually see in the thinking notes they just decide to use it. Same thing in Framer. They decide to use it. These companies who introduce things on their website are shaping the narrative and in a very short period. Amazing. Like within 30 days they're showing up as a citation. And so even though third party still gets cited more. Right. Even though there's a lot of other. I won't go down the AEO rabbit hole. I spent a lot of time there. But my point is to say that your core website still matters. The keywords, the layout, the conversations you have, the framing, having a clear structure, putting things in your header and your footer, like all of these pieces still matter. And we're going to keep evolving, we're going to keep progressing. But I would caution folks to go so radical and kind of try to just cut everything out and go minimalist. I think that is a mistake. And I would actually point to one more example. There's a AI native CRM called Lightfield, right? Light Field is a new entrance. And Lightfield does something that I love, which is right below their hero. They have a thing that basically says, here's what all the other CRMs are doing. Here's the problem. They set up the enemy very clearly. Here's the problem and here's how we. We're saying AI native CRM. Here's exactly what that means. Here's exactly what we're going to do. And then they have an article linked down on the page which is like, uh, navigating AI CRM. And it defines all the terms. Everyone's saying AI native. What does that actually mean? What does that tangibly mean? Not only is that useful for humans, but it's super useful for agents because you're giving a clear structure. Here's three different ways that people essentially use AI native, and here's the differences and here's the nuance, blah, blah, blah. So having information that's clear, structured like that, I think has super high utility.

Speaker B: You've given me another idea because our industry deals with a lot of jargon and so, and you know, there's new, new stuff coming up, like causal AI and causal attributions. I'm like, yeah, this is a good idea. I like it all. Ah, right. Kasey, this was fantastic. Thank you so much for joining us today. Um, I have two asks for everybody who's listening in. One is go follow Casey Hill on LinkedIn. He shares these types of test results all the time, and I learn a lot from that. And then, Kasey, you mentioned you have a substack. What's your substack?

Speaker A: Yeah, just do what works. I can send it as a link for folks if they want to go check it out, but it's literally just do it work substack. And in there we share a lot of these types of conversation points. So I just put it in the chat. But essentially this is where we do a lot of, like, kind of deeper analysis. So if you're interested in some of the content and some of the findings that I post on social, we essentially take that and often go deeper and we also do all the interviews of those top brands. So if you're interested in the interviews with a lot of these top websites, the people running a lot of these top websites and looking at their data and learnings. That's also all discussed on the substack.

Speaker B: Fantastic. All right, thank you. And we'll share this recording with everybody. And, uh, thanks again for joining me, Kasey.

Speaker A: Yeah, thanks for having me.

Speaker B: All right, folks, that was Casey Hill, the CMO of do what Works. Go check out his substack. Go connect with him on LinkedIn. He shares some amazing insights. And tune in again next week for another great episode of Brandformants.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • How to Stop Guessing and Start Testing with Casey Hill, DoWhatWorksThe Content Cocktail Hour · features Casey Hill78 / 100
  • Mastering SaaS Marketing - Shane Murphy-Reuter @WebflowMasters of SaaS · on CRO (Conversion Rate Optimization)83 / 100
  • The Customer Journey is Your Key to Marketing Success | With Becky SimmsThe Strategic Marketing Show · on CRO (Conversion Rate Optimization)78 / 100
  • The Right Way to A/B Test Landing Pages on Meta Ads for EcommerceMarketing Operators · on CRO (Conversion Rate Optimization)77 / 100
  • 218: Five Years of No Hacks - The Guest Host TakeoverNo Hacks · on CRO (Conversion Rate Optimization)76 / 100
  • Stop Guessing on Launch Day: Meet Your New Message Testing Framework w/Shannon KearnsProduct Marketing for You · on A/B testing framework69 / 100

More from Brandformance

All episodes →
  • Why the performance marketing era is over88 / 100
  • Why the smartest brands are investing in affiliate marketing90 / 100
  • How Figma's CMO thinks about community as a channel68 / 100
  • Why the best CMOs think like CEOs68 / 100
  • 7 Marketing Measurement Myths Costing Companies Millions76 / 100
Explore the best B2B Marketing podcasts →
All Brandformance episodes →