The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/HR/Graeme Codrington's Future of Work
Graeme Codrington's Future of Work artwork

Should We Use AI for Medical Advice? It's Complicated - ThrowForward Thursday 187

Graeme Codrington's Future of Work · 2026-05-07 · 8 min

0:00--:--

Key moments - from our scoring

Substance score

30 / 100

Five dimensions, 20 points each

Insight Density8 / 20
Originality6 / 20
Guest Caliber4 / 20
Specificity & Evidence9 / 20
Conversational Craft3 / 20

Graeme Codrington examines when generative AI is safe for medical diagnosis, drawing on Oxford University's Reasoning With Machines Laboratory research. The study found that when medical professionals used AI diagnostic tools, accuracy reached 95% - promising on the surface. However, when untrained members of the public used identical AI systems with the same symptom descriptions, accuracy plummeted to just 34%, creating dangerous outcomes. Codrington explains this gap through the Dunning-Kruger effect: experts possess domain-specific language and knowledge that lets them prompt AI more precisely and critically evaluate responses, while non-experts lack the filter to catch AI's confident but incorrect answers. He argues large language models cannot self-check or recognize the limits of their own knowledge. Codrington then introduces his firm Tomorrow Today's 5T AI Impact Model - spanning personal productivity, team augmentation, workflow redesign, transformative innovation, and trust - emphasizing that AI amplifies existing expertise rather than replacing human judgment. This episode serves operators, healthcare leaders, and enterprise decision-makers weighing AI adoption risks.

Key takeaways

  • →Medical AI tools require existing expert knowledge to use safely - untrained users get dangerously inaccurate results because they lack the domain language to properly prompt the system and validate responses.
  • →Large language models confidently generate responses without self-checking mechanisms, making them particularly dangerous for non-experts who fall victim to the Dunning-Kruger effect.
  • →The 5T AI Impact Model identifies workflow redesign (the transversal level) as where companies see real bottom-line profit improvements, yet 95% of clients are still stuck at productivity and team-member levels.
  • →AI augments and supercharges existing human expertise rather than replacing it - real value comes from experts using AI as a tool to improve their work, not from non-experts using it independently.
  • →Most organizations pursuing transformative innovation with AI have failed to build the foundational levels first, dooming their efforts to hype rather than measurable impact.

In this episode

  1. 1Oxford University's AI Medical Accuracy Study
  2. 2Why Doctors Achieve 95% Accuracy with AI Medical Diagnosis
  3. 3Why Untrained Users Only Achieve 34% Accuracy and the Dangers
  4. 4The Dunning-Kruger Effect and AI's Overconfidence Problem
  5. 5The 5T AI Impact Model for Organizations
  6. 6Building Trust in AI Systems Through Human Expertise

Mentioned

Graeme CodringtonUniversity of OxfordReasoning With Machines LaboratoryTomorrow Today

Topics in this episode

Large language modelsPrompt engineeringgenerative AIReasoning With Machines LaboratoryUniversity of OxfordDunning-Kruger effect5T AI Impact ModelTomorrow TodayMedical diagnosis accuracyWorkflow redesignAI trust and safetyMedical diagnosis

Questions this episode answers

What did the Oxford University study find about AI medical diagnosis accuracy between doctors and lay people?

Medical professionals achieved 95% accuracy using AI diagnostic systems, but untrained people using identical AI tools with the same symptom information saw accuracy drop to only 34%, proving the systems are dangerous without expert knowledge to validate responses.

Why do medical experts get better results from generative AI than non-experts for diagnosis?

Doctors use precise medical terminology and have baseline knowledge that enables better prompting - describing 'vision blur at edges' rather than 'blurry vision' - and can critically evaluate and refine AI responses, while non-experts lack this filtering ability and accept confident but incorrect answers.

What is Codrington's 5T AI Impact Model and what are its five levels?

The model outlines five implementation levels: personal productivity automation, AI as team member for experts, workflow redesign (where most real profit emerges), transformative innovation, and trust running throughout - with most companies stuck at levels one and two.

Can large language models self-check or apply the Dunning-Kruger effect to their own responses?

No, large language models cannot yet be programmed to self-check or recognize the limits of their own knowledge; they are designed to sound confident and authoritative regardless of accuracy.

How does Codrington say AI should be used most effectively?

AI should augment and supercharge existing human expertise, not replace humans; experts get measurable bottom-line impact by using AI as a bionic tool to enhance their domain knowledge, not by relying on AI for domains where they lack expertise.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

8 / 20

The episode contains one genuinely useful and data-backed insight - the dramatic accuracy gap between medical professionals (95%) and laypeople (34%) using AI for diagnosis - but the remaining content is thin filler, Dunning-Kruger recaps, and a self-promotional framework pitch that adds little instructional value for a B2B operator.

the accuracy and validity of the responses of these AI platforms plummeted to only about a third 34% accuracy, which is a disaster
large language models have no way of self checking themselves. They have no way of uh, applying the Dunning Kruger effect filter to themselves

Originality

6 / 20

The medical-AI framing using real experimental data is a mildly fresh angle, but the core conclusions - expertise amplifies AI, AI augments rather than replaces - are recycled talking points heard constantly in AI commentary; the proprietary 5T model is presented without any non-obvious reasoning behind its structure.

AI is never going to replace us as humans. AI is going to augment existing human expertise supercharged and make us bionic in what we do
The way you prompt it will enable it to get to a better answer for you

Guest Caliber

4 / 20

This is a solo monologue by the host, a futurist consultant who references external research he did not conduct and spends significant airtime promoting his firm's discovery calls; there is no practitioner guest, no demonstrated at-scale operational experience, and the episode functions more as a thought-leadership marketing piece than expert testimony.

Our team at Tomorrow Today has built a model that we call the 5T AI, uh, impact model that helps organizations to think about where and when AI can be best used
make sure you contact our team for a discovery call. We'd love to chat to you and take you into tomorrow's world today

Specificity & Evidence

9 / 20

The Oxford study data provides genuine specificity - named institution, sample sizes of ~100 professionals and 1,300 laypeople, and the 95% vs. 34% accuracy figures - but the 5T model levels are described in vague, unverifiable terms, and the claim that '95% of our clients' are stuck at early stages is an unsupported internal stat.

The Reasoning With Machines Laboratory, located at the University of Oxford in the uk
they took these scenarios firstly to a group of about a hundred medical professionals

Conversational Craft

3 / 20

This is a solo monologue with no guest, no interviewer questions, no pushback, and no dialogue of any kind; the episode closes with an overt sales pitch for a discovery call, which undermines any pretense of neutral educational intent.

Slightly different in the throw Forward Thursday studio this week. Maybe more of a PSA this week than a future focus.
make sure you contact our team for a discovery call. We'd love to chat to you and take you into tomorrow's world today

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

medical11level11language8answer7large6expertise6already5models5scenarios4model4first4expert4team4call3professionals3headache3

Episode notes

Here's something that should make you think twice before asking ChatGPT about that headache. Oxford University ran a study earlier this year. They gave the same medical scenarios to AI chatbots, twice. First with 100 doctors at the keyboard. Then with 1,300 ordinary people. The doctors got 95% accuracy. The ordinary people got 34%. Same AI. Same scenarios. Different humans. Why the gap? And what does it tell us about how the rest of us should be using AI at work? That's what this week's ThrowForward Thursday is all about. I get into the Dunning-Kruger problem with large language models, and why AI in the hands of a non-expert can be a liability rather than a tool. I also introduce the 5T AI Impact Model, our team at TomorrowToday Global uses to help organisations get past the productivity hype and into the workflow redesign where AI starts to pay for itself. More details about the 5T AI Impact Masterclass: And the links to the research: ⁠ ⁠

Full transcript

8 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Should you use generative AI for medical advice? The short answer is no, although the medium answer is yes if you are already a doctor. The longer answer is this. My name is Graham Codrington. This is Throw Forward Thursday AI as a medical doctor, yes or no? The Reasoning With Machines Laboratory, located at the University of Oxford in the uk, uh, did a remarkable piece of research earlier this year. Uh, there are links in the notes if you're interested in reading up more about it. They put together a number of scenarios, um, scenarios, uh, of people who might be having medical problems, some of them minor, some of them incredibly serious, with the need to call an ambulance immediately. And they took these scenarios firstly to a group of about a hundred medical professionals and asked the medical professionals to put the symptoms in. It's things like, um, I've got a striking headache, I'm losing vision in my left eye, my arm is tingling, that sort of thing. Uh, and then ask the system any one of the various, uh, large language models, uh, and generative AI platforms. Ask them, you know, give me a diagnosis, tell me what to do. When those hundred medical professionals put that information into these systems, they got about a, uh, 95% accuracy rate. And 95% would be good enough to say we can use these systems for medical assistance. But, but that's not the end of the story because they then took the exact same scenarios to 1,300 randomly chosen normal people without medical training and asked them to do the same thing. And there the accuracy and validity of the responses of these AI platforms plummeted to only about a third 34% accuracy, which is a disaster. Um, and basically was medically, not, I want to say medically problematic, but it was actually dangerous. And here I think you can work out, uh, what we should be doing, what the answer is to should we use AI for medical advice. The answer is, if you are already a medical professional with medical expertise, you will have the language of, of medicine. You. Instead of saying, my vision is going a bit blurry, I have a headache, you might be able to say, the vision in my left eye has begun to blur at the edges. I have a headache that is pounding in the back of my head, towards the top of my neck at vertebra number this or that, and you would describe the tingling in your arm in a very particular way because you already have a base level of knowledge and an of the language that you might use. Tingling, for example, might not be the right word. What happens is the large language model, which probably does have the right answer. The way you prompt it will enable it to get to a better answer for you and you will also be able to push back and have some level of filter to the first response to be able to ask an additional question or realize it didn't quite understand what you had said, give it a bit more information and get to 95%. In the hands of somebody uh, who doesn't have that expert level knowledge already, it becomes a disaster. This is a wonderful example of the Dunning Kruger effect. I'm sure you know this. The effect of knowing that when you don't know what you don't know, it's very easy to think that you are an expert. As soon as you begin to learn something uh, about a topic, you realize how little you know and you underestimate your ability until you develop a little bit more expert knowledge and then you've got uh, a little bit more confidence but also a healthy dose of humility at the same time. The problem with large language models is they are designed to give very confident, very specific, very authoritative sounding responses even when they are not actually um, as clear and authoritative and factual as they should be. And the bottom line is large language models have no way of self checking themselves. They have no way of uh, applying the Dunning Kruger effect filter to themselves. That's a human thing that we can do in large language models. Literally can't even be programmed to do that yet. Maybe the future, but for now, not the future for now. Uh, if you have expertise in a field, large language models become a fantastic tool to improve and increase your expertise. If you are not an expert, they can be dangerous. Our team at Tomorrow Today has built a model that we call the 5T AI, uh, impact model that helps organizations to think about where and when AI can be best used and how to get the best out of AI. The five T's refer to five different levels at which AI can uh, be implemented very, very simply. Uh, the first level is about productivity, personal uh, productivity gains with uh, automating, um, and improving your tasks. The second is when AI becomes a team member. This is when you, who already have expertise, uh, allow AI, uh, to be a tool as part of your team to improve and increase your expertise. The third level is by the way where most businesses will see real bottom line profit improvements. This is why most companies are not seeing it. Because this together level, um, or sometimes called the transversal level, is where you look for workflow redesign and we believe that that's where you start to see the real value. And we also know that about 95% of our clients haven't even got there yet. They still at these first two levels, maybe even only the first level. The fourth level is what everybody wants, which is transformative innovation. But unless you've built the foundation of these other three things, you're not going to get that fourth level. And the fifth level is more of a pillar that runs throughout this and that is we have to trust the systems. And that comes back, uh, to that Oxford University research that we can only trust the system if we know the sorts of things the system should be producing. AI is never going to replace us as humans. AI is going to augment existing human expertise supercharged and make us bionic in what we do. That is how you really get impact. Measurable, bottom line impact from AI. Slightly different in the throw Forward Thursday studio this week. Maybe more of a PSA this week than a future focus. But if you'd like to know more about our 5T model and if you'd like to know more about how to unlock real AI value, getting beyond the hype, make sure you contact our team for a discovery call. We'd love to chat to you and take you into tomorrow's world today. I'll see you next week in our studio.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Utilizing AI internally to iterate faster and empower smaller teams to upskill w/ Vivek Raghunathan #263The Engineering Leadership Podcast · on generative AI96 / 100
  • Ep 90: AI Pioneer Jürgen Schmidhuber on the State of AI TodayUnsupervised Learning with Jacob Effron · on Large language models85 / 100
  • Why B2B Brands Are Using AI to Write Sales ProposalsThe Growth Operator with Fexingo · on Prompt engineering85 / 100
  • Episode 018: Season 2, the $75 Consult and the Frankenstein StackAI Tools for Practicing Lawyers · on Prompt engineering84 / 100
  • AI Is Ready for Government. Is Government Ready?The So What from BCG · on generative AI84 / 100
  • How B2B Marketers Use AI to Personalize at Scale for EnterpriseB2B Marketing with Fexingo · on Large language models82 / 100

More from Graeme Codrington's Future of Work

All episodes →
  • Smartphones predict earthquakes - Tour Guide to the Future Ep 11640 / 100
  • $80 billion. Gone - ThrowForward Thursday 18636 / 100
  • What if... we banned phones in restaurants (like we banned smoking) - ThrowForward Thursday 18527 / 100
  • Smartphones predict earthquakes - Tour Guide to the Future Ep 116
  • Why everyone wants a piece of Greenland - ThrowForward Thursday 184
Explore the best B2B HR podcasts →
All Graeme Codrington's Future of Work episodes →