Between You and AI · 2026-05-14 · 11 min
Key moments - from our scoring
Substance score
48 / 100
Five dimensions, 20 points each
Andrea Lloro examines the METR chart - a benchmark tracking AI agent autonomy created by Beth Barnes' 30-person Berkeley nonprofit - that shows AI task-completion capability doubling every three to four months. In March 2025, the best models could handle one-hour tasks autonomously; Claude Opus 4.6 now sustains 12-hour work sessions, with projections reaching 24 - 48 hours by late 2025. The chart reveals not just speed acceleration but category shift: from answering questions (4 seconds of human work) to implementing complex systems (10+ hours of senior engineer work) in seven years. Lloro frames this through Amara's Law and Daniel Kahneman's System One cognitive bias - humans instinctively underestimate exponential curves. Rather than panic, leaders should adopt a Quarterly Delegation Audit: every three months, identify tasks AI now outperforms, anticipate what becomes automatable next, and reallocate human hours to irreplaceable strategic work. This framework applies to anyone with calendars full of planning, analysis, research, and reporting - roles most vulnerable to the next doubling wave.
METR (Model Evaluation and Threat Research) is a benchmark created by Beth Barnes, co-founder and CEO, at a 30-person nonprofit in Berkeley, California. It measures how long AI agents can work autonomously on software engineering tasks like debugging code, setting up servers, and training models before getting stuck, plotting the results on a trend line to track AI progress over time.
AI task-completion capability was doubling roughly every seven months historically, but with newer models like Claude Opus 4.5 and GPT-5.2, the doubling rate has accelerated to every three to four months.
Claude Opus 4.6 can work 12 hours straight on complex tasks such as implementing complex protocols for multiple technical specifications - work that typically takes a senior software engineer a full working day.
It's a 30-minute exercise every three months where you answer three questions: (1) what tasks did I do three months ago that AI now does better, (2) what do I do today that will likely become automatable in three months, and (3) where should I invest my human hours to stay ahead of the curve.
Amara's Law states we overestimate technology's short-term impact but dramatically underestimate long-term impact. Most people dismiss AI based on current mediocre outputs, missing that three doublings later - 18 months ahead - the world reorganizes around capabilities they were dismissing today.
Our reviewer’s read on each dimension, with quotes from the episode.
The METR task-length doubling concept and the 'quarterly delegation audit' offer genuinely useful framing for an operator, but the episode pads it heavily with hype, repetition, and motivational throat-clearing rather than dense, layered analysis.
The length in human hours of a task an AI agent was able to complete reliably was doubling roughly every seven months
every three months, you sit down with yourself. 30 minutes, a coffee, a notebook, and you answer three questions
The METR chart angle and delegation-audit prescription are moderately fresh, but Amara's Law and Kahneman's System 1/System 2 are extremely well-circulated tropes that add little new thinking.
there is a principle in technology forecasting called Amara's Law
humans operate on two distinct cognitive systems. System One is fast, instinctive, intuitive
No guest; a solo monologue by a host with relevant credentials (ex-Tinder LatAm, ex-L'Oreal CDO, MIT Tech Review columnist) but the content leans on secondhand reporting of METR rather than firsthand operating experience with the topic.
As a former head of Tinder in Latin America for five years, Chief Digital Officer at l' Oreal Brazil
Beth Barnes, co founder and CEO of Meter, asked, how long can I work without getting lost?
Cites named sources (METR, Beth Barnes) and concrete task-time examples across years, but several figures appear exaggerated or fabricated (12-hour Claude runs, confident 24/48-hour future dates), undermining the evidential weight.
in March 2025, the best model in the world could handle one hour on a single task. Today, in April 20, Claude Opus 4.6 works 12 hours straight
In October this year, that number will probably be 24. And by January next year, 48
A one-way monologue with no guest, no questions, and no challenge to any claim; rhetorically structured but there is no interviewing craft, pushback, or productive disagreement to evaluate.
do this exercise with me. Pause for a moment, please
How much of your calendar next week is already below the cloud? Opus 4.6 line. Sit with that.
Computed from the transcript - who did the talking, and the words that came up most.
In this episode of "Between You and AI", Andrea Iorio (USA Today bestselling author and keynote speaker) explores the remarkable shifts in AI's ability to work autonomously and continuously. Drawing on data from exponential progress charts, host Andrea Iorio discusses how the acceleration of AI's evolution could transform careers, daily tasks, and the job market in the coming years. If you want to understand how to prepare for the next technological doubling, this episode is essential listening. Key topics: The evolution of artificial intelligence: from one hour of autonomous work to twelve, with projections reaching forty-eight hours by January 2027. The chart from the Berkeley-based nonprofit METR, which tracks how long AI can work without getting lost - revealing a sharp acceleration over the last few years. How AI's ability to perform engineer- and analyst-level work is undergoing a rapid transformation, moving from simple tasks to high-level cognitive work. Amara's Law: the overestimation of technology in the short run versus the underestimation of its long-run impact - and why a long-term view matters now more than ever.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Every three months, something incredible happens. And no, it's not Nvidia's latest earnings report. It's not the Fed cutting rates again, not even Taylor Swift announcing another tour. But it's the fact that an UM AI doubles the time it can work on its own without getting lost and without needing us humans. Which means that if you think AI is already amazing at what it does now, well then wait for next quarter. See, in March 2025, the best model in the world could handle one hour on a single task. Today, in April 20, Claude Opus 4.6 works 12 hours straight on a single task. And here's the part I really need you to feel. In October this year, that number will probably be 24. And by January next year, 48 and so on. And unless you've been living in a cave, you've likely seen the AI chart. There is the talk in Silicon Valley, and that's moving the global stock market right now. Not the S and P chart, but a chart made by a 30 person nonprofit, um, in Berkeley, California, called Meter Model Evaluation and Threat Research. This is the chart that's deciding your professional future without asking your permission. And today in this episode of between youn and AI, we're going to understand what it's telling you about the evolution of AI so you don't get caught off guard when the next doubling hits. Here's your host, Andrea Lloro speaking. I'm an Italian KeyNote speaker and USA Today best selling author of the book between youn and AI. As a former head of Tinder in Latin America for five years, Chief Digital Officer at l' Oreal Brazil, MA mba, Professor at Funda Sanon Cabrao, and columnist at the MIT Technology Review. My latest book is between youn and AI, published by Wiley. And you can get to know me and everything about my work@AndreAYorio.com For a number of years, AI intelligence was measured as human intelligence, basically in tests, a little bit like the sat. You'd run a model through a math test, a law test, a reading test, and see how it performed. And companies would run their models through batteries of standardized exams, accessing how they stacked up against rival models at solving math problems, answering legal questions, or summarizing text accurately. Those were useful measurements, but they didn't work well when it came to AI agents. The autonomous systems that are designed to work for minutes or hours at a time. Because what you really wanted to know in this case if, uh, you're interested in these systems, was how long they could work before getting stuck. Could they handle simple tasks that would take a human a few minutes, or a more complex task that would take someone a few hours. So Beth Barnes, co founder and CEO of Meter, asked, how long can I work without getting lost? And that's how she came up with Meter. Meter's researchers attempted to track this by creating a benchmark of software engineering tasks like debugging code, setting up servers, and training small AI models. They hired expert software developers to do the tasks well humans. And then they had AI agents attempt the same tasks. And when an agent succeeded at a task well, they logged the time it had taken the human expert to do the same work. And they plotted the results on a single chart. Task length on one axis, time on the other, and produced a trend line across years of AI progress. What they found out was surprising at first. The length in human hours of a task an AI agent was able to complete reliably was doubling roughly every seven months. Now, more recently with models like anthropics, Claude Opus 4.5 on, and OpenAI's GPT 5.2, the line took a sharp upward turn. The test length is now doubling every three to four months. But I know numbers in the abstract don't land easily, so look at the chart for a second and follow the label with me. So in 2019, the best AI in the world could answer a single question. 4 seconds of human attention. Now in 2022, GPT 3.5 could count the words in a passage. About 36 seconds of clerical work. In 2023, well, GPT now managed to actually search something on the Internet, something that would take on average six minutes, which is roughly the work of a curious internal. Now, by late 2024, the O1 preview model was training a classifier 30 minutes, the work of a junior analyst. In 2025, GPT5 was training adversarially robust image models. Four hours on average, which is the afternoon of a competent machine learning engineer. And Today, Claude Opus 4.6 is implementing complex protocols for multiple technical specifications, which usually takes 10 hours. What a senior engineer does in a working day. Notice what just happened. The chart isn't only showing more time, it's showing a category change. We went from answer a question to implement a system from specifications, from reactive to architectural, from intern work to senior engineering work, just in seven years. And that's not just an acceleration of speed, but an acceleration of what kind of cognitive work counts as automatable. And if you think about your own work and if you know you think it's safe because it's more strategic More complex, more human. Well, I'd ask you to sit with this curve a little bit longer. This kind of curve is where my training as an economist starts screaming at me. Because there is a principle in technology forecasting called Amara's Law, named after Roy Amada, the Stanford trained futurist. And Amara's Law says we tend to overestimate the impact of technology in the short run and dramatically underestimated in the long run. So most people right now are still in the first half of Amara's Law. And AI, they overhyped phase of. And they watched AI write a mediocre poem and decided it was just another tech cycle. What they're missing is the second half. The part where the underestimation kicks in. The part where three doublings later, the word has quietly been reorganizing itself around a capability they were dismissing 18 months ago. And this chart isn't science fiction. It's Amara's Law plotted in real time. But hold on. Before you go and fire your whole team, like Jack Dorsey and Mark Zuckerberg have been doing recently, or more likely, spiral into an existential dread about your own job, let's pause for a second, because there's something that needs to be said, and it's kind of uncomfortable. Cause look, your brain was literally not built to read this chart. Daniel Kahneman, the Nobel laureate who basically invented a behavioral economics, spent his career mapping a single insight. That humans operate on two distinct cognitive systems. System One is fast, instinctive, intuitive. And System two is slow, deliberate and analytical. Almost of all our daily decisions, including how we read information, happen in System One. And System One has one weakness above all others. It cannot process exponential change. We, as a species, evolved in this savanna. We hunted antelope, we counted fruit. We measured distance with our own eyes, all linear. One more antelope, one less fruit. And when System One looks at an exponential graph, it instinctively flattens into a steeply rising striped lane, because that's the only shape it knows. Look, if you told someone in 1995 that within 15 years, almost every adult on Earth would carry a device with the entire knowledge of humanity in their pocket free, instantly accessible, well, they'd have called you delusional. And yet, that's exactly what happened. And not because it was unpredictable, but because System One cannot project compound curves. You can't. Which is fine, because I don't either. Nobody does. So what to do about it? Well, uh, do this exercise with me. Pause for a moment, please. Just don't do it if you're driving. But do it later. Open your calendar from last week and list five tasks that took more than an hour each. Planning meeting, a report, a market analysis, a deck, a long email response, a competitor research piece, five real tasks. And now, answer me. Not for me, but for yourself. Of these five, how many could Claude Opus 4.6, working 12 hours straight, could do today? Not 100% maybe, but 70, 80% of the work, probably more than you'd like to admit right now. The painful question. In six months, when that same model or its Successor is working 24, to wait 48 hours autonomously, how many of those tasks still make sense for you to do yourself? That's the question that will define the next decade of your career. And that's exactly what I wrote about, uh, in the second pillar of, uh, my book between you and AI called because there's a chapter, uh, that I call Augmentation. It's a skill. And this principle is, you know, simple to say and brutal to practice. Automate the routine and elevate the human part of it. But to do that, you need to know what counts as routine today. And what was human yesterday is becoming routine tomorrow. So what do we do with this? Well, uh, the answer is, first of all, not to panic, but to install a discipline. I call it the Quarterly Delegation Audit. And it's simple. Three months is roughly the speed of the next doubling. So every three months, you sit down with yourself. 30 minutes, a coffee, a notebook, and you answer three questions. One, what was I doing three months ago that AI now does better than me? Two, what do I still do today that will probably become an AI task three months from now? And three, given that, where should I be investing my human hours right now to be on the right side of this curve? 30 minutes every three months. It's the cheapest investment you can make in your career. And listen, I'll be direct, because if you don't do this, someone else will do it for you. And a decision made about you by someone else is rarely a good one for you. So here's the question I leave you with. How much of your calendar next week is already below the cloud? Opus 4.6 line. Sit with that. And when the next model drops 3, 4 months from now, with 24 autonomous hours, sit with it again. And with this, I wanted to thank you so much for your attention. This was between you and AI, and if you like this episode, please share it with colleagues and friends. Uh, hit, follow, drop a review. Tell me on LinkedIn what you thought and check out my book between youn and AI published by Wiley. That's where I unpack the whole framework of nine human skills for navigating this wave. You can find me@Andreayario.com for anything else. Keynotes, content, conversations. See you in the next episode of between you and AI.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.