The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Sales/Leading PreSales
Leading PreSales artwork

Garbage In, Confidence Out [34]

Leading PreSales · 2026-07-06 · 6 min

0:00--:--

Key moments - from our scoring

Substance score

68 / 100

Five dimensions, 20 points each

Insight Density16 / 20
Originality15 / 20
Guest Caliber12 / 20
Specificity & Evidence11 / 20
Conversational Craft14 / 20

This episode tackles a counterintuitive AI trap: when bad work is automated, it becomes polished bad work that downstream stakeholders trust without question. The host automated discovery-to-demo workflows using call transcripts as input, faithfully reproducing weak discovery practices at scale - just wrapped in professional formatting that bypassed human scrutiny. The core insight is that AI doesn't improve your baseline; it industrializes whatever baseline exists. Before automating any SE workflow (discovery, demo prep, value hypothesis), leaders must first extract and articulate their standard by studying their strongest performers. Rather than asking AI what good looks like, decode how your top SEs actually think and use that as the specification. Once the standard is crisp enough to hand to a machine, automation scales something real instead of confident mediocrity. Implementation requires upfront investment in definition, followed by human spot-checking of output during the early phase to catch drift before it compounds.

Key takeaways

  • →AI makes good work faster and bad work faster - if your discovery baseline is weak, automation just industrializes the weakness and wraps it in credibility-earning polish.
  • →Before automating any SE workflow, explicitly define what good looks like by extracting thinking from your 2-3 strongest performers, not by asking the model for generic best practices.
  • →The standard should define non-negotiable floor principles that apply universally, but leave room above the floor for context-specific craft that fits different verticals and deal sizes.
  • →Automation will feel slower at first because you must invest in articulating the standard upfront, but once real, the scaling payoff justifies the investment.
  • →Keep a human reading AI-generated output for an extended period to confirm the machine hits your stated bar and catches drift while it's still small, because polish obscures degradation.

Guests

NateEva

Topics in this episode

AI workflow automationCall transcript analysishuman-in-the-loop validationDiscovery call transcriptsDemo flow generationSE performance standardsValue hypothesis definitionWorkflow automation frameworksBaseline quality assessmentSE best practice extraction

Questions this episode answers

Why did automating discovery summaries and demo flows make demos worse?

The AI faithfully reproduced the mediocre discovery and thinking in the source transcripts, then polished the output into professional-looking briefs that downstream stakeholders trusted without question, scaling weakness instead of fixing it.

How should a leader define what good discovery actually looks like?

Extract the thinking from your 2-3 best SEs by sitting with them, understanding why they ask the questions they ask, and using their instinctive approach as the standard - not by asking AI for generic definitions from the internet.

Should the standard be the same across all team members and verticals?

No; the standard should define universal non-negotiables that always apply, but above that floor allow room for context-specific craft so different verticals and deal sizes aren't flattened into one style.

When should you implement AI automation in SE workflows?

Only after you have explicitly defined what good looks like and extracted that standard from your best performers; automating before you have a clear definition creates efficient machines for producing confident garbage.

How long should a human monitor AI-generated output after automation goes live?

Long enough to confirm the machine is hitting the bar you set and not drifting in ways that look fine on the surface; stop checking too early and degradation slides unnoticed because the output looks polished and complete.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

16 / 20

The episode delivers a concrete, actionable insight about the pitfall of automating before standardizing - 'garbage in, confidence out' - which most SE leaders would not have fully articulated. The core realization (that AI polishes bad discovery rather than fixing it) is novel and densely packed for a 6-minute format, though some time is spent on scene-setting and recap that could be tighter.

It just made the mediocrity look polished.
The polish buys credibility. The content didn't earn.

Originality

15 / 20

The framing of AI as an amplifier of existing quality (good or bad) is not entirely novel in tech discourse, but the specific application to SE discovery and the 'polished mediocrity' trap is fresh and contrarian to the typical 'AI will save your workflow' narrative. The insight about extracting standards from best performers rather than asking AI what good looks like shows original thinking.

Because if you ask the model what a great discovery looks like, you get the average of the Internet generic. What you want is your best encoded.
When bad output looked bad, people caught it. A sloppy handwritten note. You knew to double check it. Now the sloppy thinking arrives in a gorgeous three page brief with icons.

Guest Caliber

12 / 20

Nate and Eva are presented as solution engineering practitioners with hands-on experience (Eva explicitly ran the experiment with auto-generated summaries), and the show is positioned on real coaching situations from 350+ SEs. However, the speakers are voice actors/characters created by SE Rockstars rather than named independent practitioners, which limits credibility and verification of their operational depth. The host Tim co-founded SE Rockstars but does not personally participate in the substantive discussion.

I built an AI workflow to auto generate discovery summaries and suggested demo flows for my team. Pointed at the call transcript. Out comes a beautiful brief. And within two weeks, my demos got measurably worse.
Every conversation you hear on this show is based on real coaching situations, real challenges, real problems that SE leaders like you are dealing with right now. None of this is made up.

Specificity & Evidence

11 / 20

The episode provides one concrete metric (demos got 'measurably worse' within two weeks) but lacks specific numbers, company names, or detailed deal/team sizes. The advice to extract standards from 'two or three best SEs' is actionable but vague; no examples of what a strong discovery call actually contains are given. The playbook is clear but evidence-light.

Within two weeks, my demos got measurably worse.
I took my two or three best SEs, the ones who instinctively do it right, and I basically extracted their thinking, sat with them, pulled apart. Why?

Conversational Craft

14 / 20

Nate pushes back constructively on the one-standard assumption (pointing out 300-person organizations need flexibility), and there's genuine dialogue around the temporal sequencing of standardization before automation. However, the conversation is relatively brief and lacks deeper follow-up on implementation challenges, measurement, or how to identify the 'two or three best' performers objectively. The format constrains depth but the host does avoid softball questions.

Though I'd push on One thing at, uh, two or 300 people, our best Ses brain isn't one brain.
And here's the part people resist. This makes AI slower before it makes you faster.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B44%
  • Speaker C32%
  • Speaker A24%

Most-used words

discovery8makes6standard6leading5real5faster5part4output4automate4machine4best4presales3show3leaders3nate3team3

Episode notes

Before you automate discovery, define what good actually looks like. AI makes good work faster-good and bad work faster-bad - and now the bad work looks polished. Before you automate discovery or demo prep, you have to define what good looks like. What Nate and Ava discuss Why automating a weak baseline just industrializes the weakness Extracting 'what good looks like' from your two or three best SEs Keeping a human on the output long enough to catch drift The move Before automating any part of the SE workflow, write down what good looks like - sourced from your strongest people. Then, and only then, let the machine scale it. Resources & Links: paths.to/presales Book a Discovery Call: calendly.com/serockstars-tim/discovery-call

Full transcript

6 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Hey there and welcome to Leading Presales, the show for solution engineering leaders who want to build teams that drive revenue and not just demos. My name is Tim and I'm the co founder of SE Rockstars. And together with Jan, we've coached over 350 solution engineers and their leaders across several dozens of companies. Every conversation you hear on this show is based on real coaching situations, real challenges, real problems that SE leaders like you are dealing with right now. None of this is made up. We use AI to bring these stories to life through our, uh, two hosts, Nate and Eva. But the insights come straight from the trenches. Each episode gives you one actionable takeaway you can apply immediately, whether you're already leading a team or working your way into that role. All right, let's get into it.

Speaker B: So I did a thing last quarter that backfired in the most useful way. I built an AI workflow to auto generate discovery summaries and suggested demo flows for my team. Pointed at the call transcript. Out comes a beautiful brief. And within two weeks, my demos got measurably worse.

Speaker A: Hm.

Speaker C: Worse. You automated your way backwards.

Speaker B: I automated my way backwards. And it took me a minute to see why. The AI wasn't wrong. It faithfully scaled exactly what was in the transcripts, including the mediocre discovery. It just made the mediocrity look polished.

Speaker C: Welcome to Leading Presales. I'm Nate.

Speaker B: And I'm Ava. Uh, part two of our AI series. And this is the one I want every leader to sit with because it's seductive. AI makes things faster, but it makes bad things faster too.

Speaker C: I say a version of this to my managers constantly. AI makes good work faster. Good. And it makes bad work faster. Bad. If your baseline discovery is weak, you haven't fixed anything. You've just industrialized the weakness and wrapped it in a font that looks like it cost money.

Speaker B: And that last part is the trap. When bad output looked bad, people caught it. A sloppy handwritten note. You knew to double check it. Now the sloppy thinking arrives in a gorgeous three page brief with icons. And everyone downstream just trusts.

Speaker C: Launders the mistake. That's the danger. The polish buys credibility. The content didn't earn.

Speaker B: So the question I had to answer was, before I automate anything, do I actually know what good looks like? Could I point at a discovery call and say, that one? That's the standard. And honestly, I wasn't sure I'd defined it crisply enough to hand to a machine.

Speaker C: And that's the real work. Everybody wants to skip to the Automation. But you can't automate a standard you haven't articulated. When I've seen this go well, the leader did the unglamorous thing first. They defined the bar. What does a strong discovery actually contain? What separates a uh3 from a uh1?

Speaker B: Here's what worked for me. I didn't write the standard from scratch. I took my two or three best SEs, the ones who instinctively do it right, and I basically extracted their thinking, sat with them, pulled apart. Why? They ask what they ask. That became the definition of good.

Speaker C: So you didn't ask the AI what good looks like. You told it using your best people as the source.

Speaker B: Exactly. Because if you ask the model what a great discovery looks like, you get the average of the Internet generic. What you want is your best encoded.

Speaker C: Though I'd push on One thing at, uh, two or 300 people, our best Ses brain isn't one brain. Different verticals, different deal sizes. The standard has to have enough room that it doesn't flatten everyone into one style.

Speaker B: That's fair. The floor should be universal. These things are always true. But above the floor, there's craft that's context specific. You're defining the non negotiables, not a uh script.

Speaker C: And here's the part people resist. This makes AI slower before it makes you faster. You have to invest in the standard up front, but once it's real, the automation is worth having. Because now you're scaling something good.

Speaker B: Yes. And the ordering is everything. Define good, then automate, do it backwards, and you've built a very efficient machine for producing confident garbage.

Speaker C: I'd also add, keep a human reading the output for a while. Not forever, but long enough to confirm the machine is actually hitting the bar you set, not drifting past it in a way that looks fine on the surface.

Speaker B: The output looks so done that you stop checking. And that's exactly when it slides.

Speaker C: What's the move?

Speaker B: Before you automate any part of the SE workflow, discovery, demo, prep value, hypothesis, stop and write down what good looks like and don't invent it. Sit with your two or three strongest people and extract how they actually think that's your standard.

Speaker C: Then, and only then, let the machine scale it and read the first stretch of output yourself so you catch drift while it's small. Polish is not proof. I'm Nate.

Speaker B: And I'm Eva. Uh, see you next episode.

Speaker A: Thanks for listening to leading presales. If you've got a question, a topic you'd like us to cover, or you just want to connect, reach out to me@timrockstars.com and if what you heard today hit home and you want to talk about how to develop your SE team, feel free to book a discovery call@serockstars.com no strings attached. The link is of course also in the show Notes until next time and keep leading from the front.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • The Current State of Agentic Retrieval - Qdrant RoundtableMLOps.community · features Eva59 / 100
  • What Makes Us Go: Retracing America’s First Road TripSteel Stories by U. S. Steel · features Nate52 / 100
  • Data, AI, and Knowing When to Let Go - with Tommy CotterDefinitely, Maybe Agile · on human-in-the-loop validation81 / 100
  • AI Pricing & Demand Forecasting Today: Insights from a Decade in Applied AI with Jayadeep ShitolePricing Heroes · on human-in-the-loop validation81 / 100
  • AI for the Glass IndustryIndustrial AI Podcast · on human-in-the-loop validation76 / 100
  • S02E11 Embracing Change in the Workplace | Hallie Condon | Wired for WonderWired for Wonder · on AI workflow automation71 / 100

More from Leading PreSales

All episodes →
  • I Don't Know - And That's the Right Answer [30]54 / 100
  • Stop Hiring Resumes [29]52 / 100
  • Your AE Won't Brief You - Now What? [28]60 / 100
  • The Octopus Problem [33]
  • Hire or Pass - The Scoring Problem [32]
Explore the best B2B Sales podcasts →
All Leading PreSales episodes →