The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/HR/Between You and AI
Between You and AI artwork

The METR chart that shrinks your job every quarter - and is keeping Wall Street up at night.

Between You and AI · 2026-05-14 · 11 min

0:00--:--

Key moments - from our scoring

Substance score

48 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality11 / 20
Guest Caliber10 / 20
Specificity & Evidence11 / 20
Conversational Craft5 / 20

Andrea Lloro examines the METR chart - a benchmark tracking AI agent autonomy created by Beth Barnes' 30-person Berkeley nonprofit - that shows AI task-completion capability doubling every three to four months. In March 2025, the best models could handle one-hour tasks autonomously; Claude Opus 4.6 now sustains 12-hour work sessions, with projections reaching 24 - 48 hours by late 2025. The chart reveals not just speed acceleration but category shift: from answering questions (4 seconds of human work) to implementing complex systems (10+ hours of senior engineer work) in seven years. Lloro frames this through Amara's Law and Daniel Kahneman's System One cognitive bias - humans instinctively underestimate exponential curves. Rather than panic, leaders should adopt a Quarterly Delegation Audit: every three months, identify tasks AI now outperforms, anticipate what becomes automatable next, and reallocate human hours to irreplaceable strategic work. This framework applies to anyone with calendars full of planning, analysis, research, and reporting - roles most vulnerable to the next doubling wave.

Key takeaways

  • →The time AI agents can work autonomously on complex tasks is doubling every 3-4 months, meaning capabilities that seem impressive today will be routine in 6-12 months.
  • →Most of your current work tasks - planning meetings, reports, market analysis, competitive research - could already be 70-80% automated by Claude Opus 4.6 today.
  • →Amara's Law explains why most professionals are still in the overhype phase of AI; the real disruption happens in the understimated second half when capabilities quietly become mainstream.
  • →Humans cannot cognitively process exponential change because System One thinking evolved linearly; you must consciously audit your work quarterly to stay ahead of automation.
  • →The Quarterly Delegation Audit - 30 minutes every three months answering what AI now does better, what will be automated next, and where to invest human hours - is the cheapest career insurance available.

In this episode

  1. 1The METR Chart: How AI Task Duration is Doubling Every Quarter
  2. 2Evolution from Benchmark Tests to Autonomous Agent Measurement
  3. 3Seven Years of Progress: From Question-Answering to Senior Engineering Work
  4. 4Amara's Law and Human Cognitive Blind Spots to Exponential Change
  5. 5The Quarterly Delegation Audit: A Practical Framework for Career Survival
  6. 6Identifying Tasks Below the Claude Opus Line in Your Calendar

Mentioned

METRClaude Opus 4.6GPT-5.2OpenAIAnthropicO1Beth BarnesAndrea AiorioDaniel KahnemanRoy AmaraNvidiaWiley

Topics in this episode

AI agentsBehavioral economicsOpenAIAnthropicClaude Opus 4.6Daniel KahnemanGPT-5METR (Model Evaluation and Threat Research)Amara's LawSystem One and System Two thinking

Questions this episode answers

What is the METR chart and who created it?

METR (Model Evaluation and Threat Research) is a benchmark created by Beth Barnes, co-founder and CEO, at a 30-person nonprofit in Berkeley, California. It measures how long AI agents can work autonomously on software engineering tasks like debugging code, setting up servers, and training models before getting stuck, plotting the results on a trend line to track AI progress over time.

How fast is AI autonomy capability doubling according to the METR chart?

AI task-completion capability was doubling roughly every seven months historically, but with newer models like Claude Opus 4.5 and GPT-5.2, the doubling rate has accelerated to every three to four months.

What specific tasks can Claude Opus 4.6 handle autonomously today?

Claude Opus 4.6 can work 12 hours straight on complex tasks such as implementing complex protocols for multiple technical specifications - work that typically takes a senior software engineer a full working day.

What is the Quarterly Delegation Audit and how should professionals use it?

It's a 30-minute exercise every three months where you answer three questions: (1) what tasks did I do three months ago that AI now does better, (2) what do I do today that will likely become automatable in three months, and (3) where should I invest my human hours to stay ahead of the curve.

Why does Amara's Law explain why most people underestimate AI's impact?

Amara's Law states we overestimate technology's short-term impact but dramatically underestimate long-term impact. Most people dismiss AI based on current mediocre outputs, missing that three doublings later - 18 months ahead - the world reorganizes around capabilities they were dismissing today.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

The METR task-length doubling concept and the 'quarterly delegation audit' offer genuinely useful framing for an operator, but the episode pads it heavily with hype, repetition, and motivational throat-clearing rather than dense, layered analysis.

The length in human hours of a task an AI agent was able to complete reliably was doubling roughly every seven months
every three months, you sit down with yourself. 30 minutes, a coffee, a notebook, and you answer three questions

Originality

11 / 20

The METR chart angle and delegation-audit prescription are moderately fresh, but Amara's Law and Kahneman's System 1/System 2 are extremely well-circulated tropes that add little new thinking.

there is a principle in technology forecasting called Amara's Law
humans operate on two distinct cognitive systems. System One is fast, instinctive, intuitive

Guest Caliber

10 / 20

No guest; a solo monologue by a host with relevant credentials (ex-Tinder LatAm, ex-L'Oreal CDO, MIT Tech Review columnist) but the content leans on secondhand reporting of METR rather than firsthand operating experience with the topic.

As a former head of Tinder in Latin America for five years, Chief Digital Officer at l' Oreal Brazil
Beth Barnes, co founder and CEO of Meter, asked, how long can I work without getting lost?

Specificity & Evidence

11 / 20

Cites named sources (METR, Beth Barnes) and concrete task-time examples across years, but several figures appear exaggerated or fabricated (12-hour Claude runs, confident 24/48-hour future dates), undermining the evidential weight.

in March 2025, the best model in the world could handle one hour on a single task. Today, in April 20, Claude Opus 4.6 works 12 hours straight
In October this year, that number will probably be 24. And by January next year, 48

Conversational Craft

5 / 20

A one-way monologue with no guest, no questions, and no challenge to any claim; rhetorically structured but there is no interviewing craft, pushback, or productive disagreement to evaluate.

do this exercise with me. Pause for a moment, please
How much of your calendar next week is already below the cloud? Opus 4.6 line. Sit with that.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

three10hours10human10chart9task7today7tasks7system7model6minutes6single5opus5models5answer5question5claude4

Episode notes

In this episode of "Between You and AI", Andrea Iorio (USA Today bestselling author and keynote speaker) explores the remarkable shifts in AI's ability to work autonomously and continuously. Drawing on data from exponential progress charts, host Andrea Iorio discusses how the acceleration of AI's evolution could transform careers, daily tasks, and the job market in the coming years. If you want to understand how to prepare for the next technological doubling, this episode is essential listening. Key topics: The evolution of artificial intelligence: from one hour of autonomous work to twelve, with projections reaching forty-eight hours by January 2027. The chart from the Berkeley-based nonprofit METR, which tracks how long AI can work without getting lost - revealing a sharp acceleration over the last few years. How AI's ability to perform engineer- and analyst-level work is undergoing a rapid transformation, moving from simple tasks to high-level cognitive work. Amara's Law: the overestimation of technology in the short run versus the underestimation of its long-run impact - and why a long-term view matters now more than ever.

Full transcript

11 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Every three months, something incredible happens. And no, it's not Nvidia's latest earnings report. It's not the Fed cutting rates again, not even Taylor Swift announcing another tour. But it's the fact that an UM AI doubles the time it can work on its own without getting lost and without needing us humans. Which means that if you think AI is already amazing at what it does now, well then wait for next quarter. See, in March 2025, the best model in the world could handle one hour on a single task. Today, in April 20, Claude Opus 4.6 works 12 hours straight on a single task. And here's the part I really need you to feel. In October this year, that number will probably be 24. And by January next year, 48 and so on. And unless you've been living in a cave, you've likely seen the AI chart. There is the talk in Silicon Valley, and that's moving the global stock market right now. Not the S and P chart, but a chart made by a 30 person nonprofit, um, in Berkeley, California, called Meter Model Evaluation and Threat Research. This is the chart that's deciding your professional future without asking your permission. And today in this episode of between youn and AI, we're going to understand what it's telling you about the evolution of AI so you don't get caught off guard when the next doubling hits. Here's your host, Andrea Lloro speaking. I'm an Italian KeyNote speaker and USA Today best selling author of the book between youn and AI. As a former head of Tinder in Latin America for five years, Chief Digital Officer at l' Oreal Brazil, MA mba, Professor at Funda Sanon Cabrao, and columnist at the MIT Technology Review. My latest book is between youn and AI, published by Wiley. And you can get to know me and everything about my work@AndreAYorio.com For a number of years, AI intelligence was measured as human intelligence, basically in tests, a little bit like the sat. You'd run a model through a math test, a law test, a reading test, and see how it performed. And companies would run their models through batteries of standardized exams, accessing how they stacked up against rival models at solving math problems, answering legal questions, or summarizing text accurately. Those were useful measurements, but they didn't work well when it came to AI agents. The autonomous systems that are designed to work for minutes or hours at a time. Because what you really wanted to know in this case if, uh, you're interested in these systems, was how long they could work before getting stuck. Could they handle simple tasks that would take a human a few minutes, or a more complex task that would take someone a few hours. So Beth Barnes, co founder and CEO of Meter, asked, how long can I work without getting lost? And that's how she came up with Meter. Meter's researchers attempted to track this by creating a benchmark of software engineering tasks like debugging code, setting up servers, and training small AI models. They hired expert software developers to do the tasks well humans. And then they had AI agents attempt the same tasks. And when an agent succeeded at a task well, they logged the time it had taken the human expert to do the same work. And they plotted the results on a single chart. Task length on one axis, time on the other, and produced a trend line across years of AI progress. What they found out was surprising at first. The length in human hours of a task an AI agent was able to complete reliably was doubling roughly every seven months. Now, more recently with models like anthropics, Claude Opus 4.5 on, and OpenAI's GPT 5.2, the line took a sharp upward turn. The test length is now doubling every three to four months. But I know numbers in the abstract don't land easily, so look at the chart for a second and follow the label with me. So in 2019, the best AI in the world could answer a single question. 4 seconds of human attention. Now in 2022, GPT 3.5 could count the words in a passage. About 36 seconds of clerical work. In 2023, well, GPT now managed to actually search something on the Internet, something that would take on average six minutes, which is roughly the work of a curious internal. Now, by late 2024, the O1 preview model was training a classifier 30 minutes, the work of a junior analyst. In 2025, GPT5 was training adversarially robust image models. Four hours on average, which is the afternoon of a competent machine learning engineer. And Today, Claude Opus 4.6 is implementing complex protocols for multiple technical specifications, which usually takes 10 hours. What a senior engineer does in a working day. Notice what just happened. The chart isn't only showing more time, it's showing a category change. We went from answer a question to implement a system from specifications, from reactive to architectural, from intern work to senior engineering work, just in seven years. And that's not just an acceleration of speed, but an acceleration of what kind of cognitive work counts as automatable. And if you think about your own work and if you know you think it's safe because it's more strategic More complex, more human. Well, I'd ask you to sit with this curve a little bit longer. This kind of curve is where my training as an economist starts screaming at me. Because there is a principle in technology forecasting called Amara's Law, named after Roy Amada, the Stanford trained futurist. And Amara's Law says we tend to overestimate the impact of technology in the short run and dramatically underestimated in the long run. So most people right now are still in the first half of Amara's Law. And AI, they overhyped phase of. And they watched AI write a mediocre poem and decided it was just another tech cycle. What they're missing is the second half. The part where the underestimation kicks in. The part where three doublings later, the word has quietly been reorganizing itself around a capability they were dismissing 18 months ago. And this chart isn't science fiction. It's Amara's Law plotted in real time. But hold on. Before you go and fire your whole team, like Jack Dorsey and Mark Zuckerberg have been doing recently, or more likely, spiral into an existential dread about your own job, let's pause for a second, because there's something that needs to be said, and it's kind of uncomfortable. Cause look, your brain was literally not built to read this chart. Daniel Kahneman, the Nobel laureate who basically invented a behavioral economics, spent his career mapping a single insight. That humans operate on two distinct cognitive systems. System One is fast, instinctive, intuitive. And System two is slow, deliberate and analytical. Almost of all our daily decisions, including how we read information, happen in System One. And System One has one weakness above all others. It cannot process exponential change. We, as a species, evolved in this savanna. We hunted antelope, we counted fruit. We measured distance with our own eyes, all linear. One more antelope, one less fruit. And when System One looks at an exponential graph, it instinctively flattens into a steeply rising striped lane, because that's the only shape it knows. Look, if you told someone in 1995 that within 15 years, almost every adult on Earth would carry a device with the entire knowledge of humanity in their pocket free, instantly accessible, well, they'd have called you delusional. And yet, that's exactly what happened. And not because it was unpredictable, but because System One cannot project compound curves. You can't. Which is fine, because I don't either. Nobody does. So what to do about it? Well, uh, do this exercise with me. Pause for a moment, please. Just don't do it if you're driving. But do it later. Open your calendar from last week and list five tasks that took more than an hour each. Planning meeting, a report, a market analysis, a deck, a long email response, a competitor research piece, five real tasks. And now, answer me. Not for me, but for yourself. Of these five, how many could Claude Opus 4.6, working 12 hours straight, could do today? Not 100% maybe, but 70, 80% of the work, probably more than you'd like to admit right now. The painful question. In six months, when that same model or its Successor is working 24, to wait 48 hours autonomously, how many of those tasks still make sense for you to do yourself? That's the question that will define the next decade of your career. And that's exactly what I wrote about, uh, in the second pillar of, uh, my book between you and AI called because there's a chapter, uh, that I call Augmentation. It's a skill. And this principle is, you know, simple to say and brutal to practice. Automate the routine and elevate the human part of it. But to do that, you need to know what counts as routine today. And what was human yesterday is becoming routine tomorrow. So what do we do with this? Well, uh, the answer is, first of all, not to panic, but to install a discipline. I call it the Quarterly Delegation Audit. And it's simple. Three months is roughly the speed of the next doubling. So every three months, you sit down with yourself. 30 minutes, a coffee, a notebook, and you answer three questions. One, what was I doing three months ago that AI now does better than me? Two, what do I still do today that will probably become an AI task three months from now? And three, given that, where should I be investing my human hours right now to be on the right side of this curve? 30 minutes every three months. It's the cheapest investment you can make in your career. And listen, I'll be direct, because if you don't do this, someone else will do it for you. And a decision made about you by someone else is rarely a good one for you. So here's the question I leave you with. How much of your calendar next week is already below the cloud? Opus 4.6 line. Sit with that. And when the next model drops 3, 4 months from now, with 24 autonomous hours, sit with it again. And with this, I wanted to thank you so much for your attention. This was between you and AI, and if you like this episode, please share it with colleagues and friends. Uh, hit, follow, drop a review. Tell me on LinkedIn what you thought and check out my book between youn and AI published by Wiley. That's where I unpack the whole framework of nine human skills for navigating this wave. You can find me@Andreayario.com for anything else. Keynotes, content, conversations. See you in the next episode of between you and AI.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • 115: Rethinking AI Governance for Enterprise Adoption with Dr. Markus SchmidbergerUsing AI at Work · on AI agents92 / 100
  • Eric Ries on Why Good Companies Go BadPodcast Archives · on Anthropic92 / 100
  • The 18x Midas Lister Betting $3B on AI (and calling most of it fake) | Navin Chaddha, MayfieldThe Peel with Turner Novak · on OpenAI91 / 100
  • 183: Why Trusted Data is the New AI Moat (w/ Rick Kranz @ AI Marketing Automation Lab)Move The Needle · on Anthropic91 / 100
  • Why Ploy.ai is more than another website tool - 60 MINUTES with Bryant Chou20 MINUTES by Noco · on AI agents87 / 100
  • Agentic Engineering for Testers: How to Automate Your Way to the Top with Amit RawatTestGuild Automation Podcast · on OpenAI82 / 100

More from Between You and AI

All episodes →
  • Tokenmaxxing: Why AI Keeps Getting More Expensive - and What It Means for Leaders.72 / 100
  • AI-xhaustion: The Double-Edged Sword of AI's Impact on Our Mental Health72 / 100
  • The Pope, Anthropic, and the battle for "trust" in the AI industry.64 / 100
  • The Boomerang Effect: Are AI Layoffs for Real, or Just an Alibi?73 / 100
  • FOBO: Why is Gen Z sabotaging AI’s transformation at work.59 / 100
Explore the best B2B HR podcasts →
All Between You and AI episodes →