
Leveraging AI · 2026-06-27 · 46 min
Key moments - from our scoring
Substance score
49 / 100
Five dimensions, 20 points each
The AI industry has undergone a dramatic pivot from chat-based interfaces to agentic systems in just months. OpenAI's internal data reveals that Codex accounts for 99.8% of weekly output tokens across the entire organization - not just engineering but legal, finance, recruiting, and business operations - with token usage surging 56x in research, 32x in consumer support, and 27x in engineering since November 2025. This shift reflects a fundamental transformation: 25.6% of sampled users generated single Codex requests exceeding eight hours of equivalent human work, and heavy users at the 99th percentile ran over 60 hours of parallel agent turns daily. OpenAI's announcement that custom GPTs will sunset by August 2026 accelerates this transition. Anthropic's Claude Tag - an always-on Slack agent with access to codebases and multiple data sources - represents the emerging pattern of seamless human-AI collaboration in communication channels. Google DeepMind's TraitR framework (Taxonomy of Rogue AI Tactics and Routines) introduces threat modeling for autonomous agents, addressing real risks including loss of control, work sabotage, and direct harm. McKinsey projects agentic AI will unlock $2.9 trillion in US economic value by 2030. The EU AI Act's high-risk provision mandates robust oversight mechanisms starting August 2026. Early enterprise adopters at Rippling and Anthropic report agents handling email drafting, data synthesis, and code generation - with Anthropic's internal team generating 65% of product code through Claude Tag. For B2B operators, this represents an immediate shift requiring budget planning for token surges, security redesign for agents spanning multiple domains, and workforce reskilling in agentic orchestration.
By June 2026, Codex (OpenAI's agentic tool) accounts for 99.8% of weekly output tokens across the entire organization, with token usage surging 56x in research, 32x in consumer support, and 27x in engineering since November 2025 - affecting legal, finance, recruiting, and all other departments.
Claude Tag is an always-on agent available in Slack that can be called with @Claude to take actions across channels; it tracks information, pulls data from multiple sources, creates pull requests, analyzes data, solves bugs, and participates in team discussions while having access to codebases and company data.
OpenAI will sunset custom GPTs for business users by August 2026, leaving approximately two months for enterprises to convert existing custom GPTs into projects or agentic solutions.
TraitR (Taxonomy of Rogue AI Tactics and Routines) is a threat modeling framework identifying three AI agent risk categories - loss of control, work sabotage, and direct harm - with oversight mechanisms including real-time access control, chain-of-thought monitoring, and supervisor agents that can kill rogue agents.
Non-developer organizational usage of Codex has grown 189x in less than a year (since August 2025), and over 25% of non-coder tasks now involve code generation, showing how agentic AI democratizes technical capabilities across business functions.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode delivers substantial data points and concrete figures (99.8% token usage, 56x surge in research tokens, 189x growth in non-developer usage) that provide real insight into agentic AI adoption at scale. However, significant portions consist of repetitive restatement of the same metrics and promotional content for the host's course, diluting the insight-per-minute ratio. The healthcare segment introduces genuinely new information but remains somewhat surface-level.
OpenAI internal data is now showing that by June 2026, Codex, so their internal agentic tool, accounts for 99.8% of weekly output tokens at OpenAI
Token has surged from November 2025 to June 2026. On the research side, 56h - 56x. On consumer support, 32x. Engineering, 27x, and legal 13x
The episode synthesizes recent industry announcements (Codex data, Claude Tag, Jalapeño chip, healthcare breakthroughs) but largely reports on what companies have already announced rather than developing original analysis or counterintuitive claims. The host does offer some original framing around demand dynamics and governance challenges, but these sections are brief and underdeveloped. Most content follows industry narratives rather than questioning them.
the agentic shift is not a forecast, it is happening, and it is happening very, very fast
to me, there is a finite amount of demand...I don't see that as increasing the economic value by so much money
This is a solo host episode with no guest appearances. The host (Isar Metis) discusses third-party researchers, executives, and industry figures but does not interview them directly. There is no guest to evaluate on caliber.
This is Isar Metis, your host, and we are going to cover four main topics today
The episode provides abundant specific metrics, company names, dates, and concrete examples: OpenAI's Codex usage figures, Noam Shazeer's $2.7B acqui-hire value, Boston Children's Hospital's 18/367 diagnosis rate, Midjourney's $74M commitment with 50,000 unit deployment target by 2031, Aidoc's 120M scans analyzed. However, some claims lack supporting evidence (e.g., Gemini model performance vs. Chinese open-source models mentioned without specifics), and promotional claims about the course lack data.
OpenAI internal data is now showing that by June 2026, Codex, so their internal agentic tool, accounts for 99.8% of weekly output tokens
in an acqui-hire back in August of 2024, Google brought him and his team back for a sum that was disclosed at $2.7 billion
As a solo monologue episode rather than a conversation, there are no guest interactions to evaluate, no challenging questions, or dynamic follow-ups. The host maintains a structured narrative flow and occasionally poses rhetorical questions to the audience, but without an actual conversational partner, the dimension cannot be fairly assessed. The presentation is clear but lacks the adversarial testing of ideas that defines strong conversational craft in a B2B context.
This is Isar Metis, your host, and we are going to cover four main topics today
Now, why is this a big deal?
Computed from the transcript - who did the talking, and the words that came up most.
Is your business ready for a world where AI doesn't just answer questions - but gets the work done? The age of agentic AI has officially arrived. Companies are rapidly shifting from chat-based AI to autonomous AI agents capable of completing hours of work with minimal human involvement. The organizations that adapt early are likely to gain a significant competitive advantage, while those that hesitate risk falling behind. In this episode, Isar Metis unpacks the biggest AI developments of the week - from OpenAI's dramatic internal transformation and Anthropic's latest enterprise agent, to the growing talent war between leading AI labs and breakthrough healthcare applications. More importantly, he explains what these changes mean for business leaders and how to prepare your organization for the next phase of AI adoption. In this session, you'll discover: Why OpenAI has shifted almost entirely to agentic AI internally. What the end of Custom GPTs means for businesses. How Anthropic's Claude is becoming a true AI teammate inside Slack. The biggest challenges around AI governance, security, and trust. Why AI agents are creating new budgeting and infrastructure challenges.
Transcribed and scored by The B2B Podcast Index.
Hello and welcome to a weekend news episode of the Leveraging AI podcast, a podcast that shares practical, ethical ways to leverage AI to improve efficiency, grow your business and advance your career. This is Isar Metis, your host, and we are going to cover four main topics today. One is breakthroughs in healthcare and AI, which I'm personally very excited about. We're going to talk about how agentic the AI world is becoming very, very quickly right in front of my eyes.
We are going to talk about the talent wars and the exodus of top leading researchers from Gemini to other labs. And we're going to mention a few important releases that come from the major labs this week. Not the big models that we expected. They will probably still come in the next few days, but there's still interesting things to share.
And then we have a few other small things to talk about. So let's get started.. The first topic I want to cover today is the topic of agentic AI and how dramatic the change have been in the past few months, and how mainstream agentic AI is right now compared to what it was, well, just a few months ago, definitely in the beginning of this year. So OpenAI internal data is now showing that by June 2026, Codex, so their internal agentic tool, accounts for 99.
8% of weekly output tokens at OpenAI. Now, that's not a test, that's not a pilot. That is the full organization shift, and this is not just writing code. This includes legal, finance, recruiting, basically every major aspect of the company.
Most of the token production is done on Codex and not on a regular ChatGPT chat, meaning everybody in the company is using agentic AI to deliver outputs, and basically nothing is just using old school ChatGPT chat So let's dive a little deeper into the numbers and connect a few dots beyond just this mind-blowing fact. First of all, the majority of code writing moved to Codex back in December of 2025. But legal, finance, recruiting, and other departments are now using the majority of their usage as of April 2026 on Codex as well versus the regular GPT chat.
Token has surged from November 2025 to June 2026. On the research side, 56h - 56x. On consumer support, 32x. Engineering, 27x, and legal 13x.
That shows you two different things. One is that agentic AI drives significant work because otherwise some of these departments will not do it. Two is that for OpenAI, that is not associated with direct immediate cost. There is compute cost obviously that they're spending, but they're not paying a third party to use that compute.
If you are going to do the same thing, if you're going to go all in on agentic AI in your company, you need to expect a huge surge in token usage. And I can tell you as somebody who's working with several different companies, and some of them are large enterprises, this is the talk of the day, how to save costs and first of all understand cost because it's not easy even on a single platform. Even if you commit just to ChatGPT or to Copilot or to Claude, if you just commit to one, and many companies commit to a few depending on departments or sections or needs, but even if you just commit to one, it is very difficult and really confusing to understand how the tokens are being used.
But I'm sharing this with you so you understand if you are planning to go into the AI universe, this is what you need to expect, a huge surge in token usage. Now, a few additional interesting pieces of information that came out of this ChatGPT data. Long horizon tasks adoption is the next topic. 80.
6% of sampled individual users made at least one Codex request estimated to exceed thirty minutes of human work. 70.2% exceeds one hour, and 25.6% exceeding eight hours.
So again, this is not eight hours of the AI working, this is eight hours of human work that is now done in a single pass with a single task done by the AI in an agentic way. Now, to put things in a bigger perspective, the really big number is that by June 2026, heavy users at the 99th percentile generated more than 60 hours of Codex agent turns per day across multiple parallel agent simultaneously. Now, I work in multiple agents in parallel all the time. I'm pretty sure I don't come even close to 60 hours in a single day.
I'm probably over the 24-hour mark, even though I'm working a lot less than 24 hours. And so this is another thing that you need to pay attention to, the ability to run agents in parallel. So while you are working on one thing, they are doing something else, and you're jumping back and forth between different tools and between different sessions in order to maximize your time because you're the bottleneck in defining and in testing and in checking and in guiding. You can run multiple of these in parallel and make the most out of these incredible capabilities.
Now, the usage of AI in OpenAI internally in what they call organizational non-developer users rose 189x from August of last year. So in less than a year, it is almost 200 times more AI usage, through non-developers compared to what they did just a year ago. Again, this is something you need to take into account, from a budgeting perspective, from a training perspective, from a compute perspective, from a planning perspective, uh, because this will happen in every organization, and And those who do it faster will gain huge competitive advantage over the ones who take longer to do this Another interesting parameter that came out of this paper that, by the way, is called How Agents Are Transforming Work, which uses data from OpenAI themselves.
They said that over one-fourth of work done by Codex by business operations, so not people who write code, but the people who run different operations in the company, involved engineering or coding tasks. Meaning, while you don't understand how to write code and you don't know how it works and you can't read it, allowing the AI to write code can generate immense benefits and business value and faster operations and more precise outcomes, and more or less everybody can do it right now.
And what they're saying is that right now, one-fourth of all tasks performed by non-developers included writing code. I write code. When I say I, I don't write any code. I don't know how to, but tasks that I do involve writing code all the time, and it's now a second nature.
I don't really care about it. I don't even think about it. I know what the outcome that I need. The AI does what it needs to do in order to get me there, and everybody's happy, and the tools are getting better and better, and there are fewer and fewer bugs in the things that I'm creating and deploying, and that saves me the trouble of trying to help the AI troubleshoot the issues.
It just knows how to write code and troubleshoot it itself very well Now, OpenAI announced this week that they're further expanding their Workspace Agent, which they announced in April of 2022. So they have released multiple types of agents that run across multiple aspects of businesses, and they keep on adding more and more to these, and there was another batch of these released just this past week. So even if you do not know how to create your own agents, which by the way, is not rocket science, and if you wanna learn, we have a course that teaches that.
The multi-agent orchestration course that we have been teaching since April of this year to both organizations as private workshops, that are open to the public. We are almost sold out of the August cohort. So we sold out April, May, June, and July, and now we're almost sold out of August. So if you wanna join the next course, you need to hurry.
If not, you'll have to wait till September, which is still fine, but it is something you have to learn because it is complete game changer to more or less every aspect of the business. And if you are in a leading position in your organization and you want to train a large number of people in your company in a private session, please reach out to me on LinkedIn or on my email. There are links to all these things in the show notes. But back to the news, OpenAI is providing more and more agents out of the box that you can use in order to run agentic capabilities across more or less every aspect of the business Now, if it wasn't obvious that that's where the world is going, OpenAI are going to sunset custom GPTs.
So they are - OpenAI are going to shut down all custom GPTs by August of 2026. This is one of the most used tools for people before the agentic era to build different kind of automations, and now you have just two to three months left to convert all your custom GPTs either into projects or into agentic solutions. Now, they may keep it for private users, I'm not sure, but for business users, the usage of custom GPTs is going away, and the idea is to convert everybody to agentic usage.
It will be more effective and will drive significantly higher revenue to OpenAI because agentic AI uses significantly more tokens But it is a clear cut showing you how we're shifting away from the text chat era to the agentic era in a very clear timeline that is in the immediate future From a different company, Anthropic just launched Claude Tag this week. This is now available to enterprise and teams plans. And what it is, that lives inside of Slack, and you can the robot, and it can take action and do a lot of things inside of Slack channels Now, to explain what this is not another Slack bot.
We had Slack bots for a very long time. Some of them are slightly more AI-ish in their nature, but Claude Tag is something completely different. It is an always-on agent that runs inside of Slack that you can call by @Claude inside of a Slack channel where you allow it access, and it can do amazing things It can track every piece of information in the channel. It can pull data from multiple sources and bring it into the channel as needed.
It can monitor everything that's going on. It can do pull request creations, data analysis, go-to-market tasks, solving bugs in the code, and so on, all by simply calling it as you would do other team members with one huge difference. It has access to your entire code base potentially, and it can do things significantly faster than humans can, uh, but it will participate in the chat just like a human. This is something that we're going to see more and more of in Slack and in other communication channels where we start working with humans and agents at the same time in a seamless way, and this is a very strong move by Anthropic in that direction.
Now these are just two examples of what's happening in the industry right now. If you're looking at the competitive landscape, we just informed you last week that Microsoft Copilot CoWork is just became available to everybody. We have Google Agent Space, Salesforce Agentforce, Anthropic Claude Managed Agents, and so on and so forth. Everything is becoming more and more agentic, and the amount of tokens that are being used or consumed in these agentic agents are skyrocketing across the board with almost every single direction you're looking at Now, that also raises a lot of concerns on how far are we pushing the limits of trust with these AI systems, both in means of access to data as well as in means of just risks to humanity if you want, if one of these goes rogue.
And Google DeepMind just released an interesting paper this week. It is a 35-page framework introducing what they call TraitR, which is spelled T-R-A-I-T and R, which stands for Taxonomy of Rogue AI Tactics and Routines. It is basically a threat modeling framework that is modeled specifically to prevent an attack that is done by an actual AI or a rogue AI and how we actually monitor its work and stop it before it generates any serious damage. It includes three identified threat categories.
The first one is loss of control, where they define as an agent that secretly creates unmonitored copies or disables safeguards. Number two is work sabotage, an agent that intentionally introduces flaws into research or code. And number three, direct harm, an agent that steals sensitive data, damages critical systems, or assists cyber attacks. Now, these aren't theoreticals.
They're all things that are happening right now and that could dramatically expand once much more powerful AI is being released into the wild. We just heard what happened with the government stopping Anthropic from releasing or actually pulling back their latest two models because of risks to cybersecurity. But this is right now, you have to start running the trajectory of the current scaling laws and think where we'll be a year from now, two years from now, three years from now.
And so Google DeepMind is starting to think how we can actually keep these models safe from harming us intentionally. They're already causing issues from mistakes. We heard multiple crazy stories about complete databases or email backends that have been deleted or entire sections of code. But all of these were mistakes.
We're now talking about an AI agent going rogue and doing this intentionally. So the full stack of the solution includes dynamic real-time access control, chain of thought monitoring, asynchronous alerts, shutdown infrastructure, and trusted AI supervisor agents that are monitoring the other agents, making sure they are aligned with what they need to do, and they have a kill switch to actually stop the other agent from doing the thing that they don't want it to do Now there's obviously different opinions on whether this approach is effective or not.
First of all, Rohin Shah, who is the lead of the AGI Safety and Alignment team at Google DeepMind, explained the philosophy, and I'm quoting, "The first line of defense is always to align the AI systems, but it is always good to have multiple layers. We borrow a lot from security, of internal employees who might be malicious, and we can apply these to a new setting." Now, part of the problem with dealing with agentic security is that standard role-based access control breaks down once you try to apply to agents, because agents automate workflows across multiple domains and multiple systems and multiple tools, and they blur the lines between different departments and different roles or traditional roles in the organization.
So a single agent may perform tasks associated with multiple old school human roles, and that means that the existing safety and security system is not relevant anymore, and that's something that doesn't have a solution yet. A similar problem, by the way, exists when you come to deploy not problematic agents or connecting your tools into an AI universe. And that means that anybody that now has access to the AI tools has access to the entire data in that tool. So if you connect your ERP to an AI system and give people access to that AI system, they might be able to see information that they shouldn't be seeing, and that's a problem that you have to solve before you actually deploy these kind of solutions.
So there are multiple aspects to this question, not just the specific question of security from an agent going rogue. Now, the flip side of that, a different opinion was proposed or sounded by other people, including Professor Don Song from UC Berkeley, who said that there is a serious flaw in what they're suggesting, which is basically saying that because these models, the supervising models and the actual models they are supervising, are using a similar architecture and a similar set of values that was built by the same company, meaning it may fail to flag misbehavior because it is running on a similar DNA of how it runs.
Now, whether he's right or wrong, I don't know. I assume Google thought about this problem. But what he's saying is that either by mistake or, and I'm quoting, "If the monitor model won't flag failures because it is protecting its peer, the entire oversight architecture breaks." Now, does that mean that this is completely wrong?
No, I'm very excited about this actually, and I'm excited about this because it is the first time I hear any of the big labs actually paying attention and trying to build a mechanism, a process, a model, call it whatever you wanna call it, that will look into this kind of possibility, which I am more or less certain we will hit at a certain point, and trying to think how to prevent it from getting into catastrophic outputs. Now, Catastrophic doesn't have to mean the world comes to an end.
Catastrophic could be for a single company losing everything it has overnight because nobody was paying attention and the AI agents had access to too many things. So it is important to start thinking about these things from a statistical perspective DeepMind has audited approximately one million coding agent tasks while they were testing this system, and what they found is that they successfully caught unintentional data deletions before the damage occurred, and their internal evaluation reported a 94% reduction in jailbreaks.
Now, while this being said, they said specifically that all of these cases, 100% of them, were done by mistake and not by a malicious intent that might be harder to detect. Now DeepMind are not the first company who's launching something like this. Databricks have launched Beyond Identity. Hugging Face have their own agentic security frameworks.
But this is definitely an area that is in just making its first baby steps because we are not sure exactly how AI will evolve. We're not sure exactly what to plan for or how to plan for it. But again, I'm very happy that people are paying attention and actually doing these kind of things If you look at the bigger picture or the broader picture, you know, McKinsey economic projection for agentic AI projected earlier this year that AI agent will unlock 2.9 trillion of value in the US economy alone by 2030.
That is a crazy, insane amount of money. But if you are following the path and you're thinking of how much work every single individual can generate, that tells you how much it can unlock. Now, that still needs consuming power on the other side to actually consume all that value to make it worthwhile, and this is the question I haven't heard anybody talking about, but maybe I don't understand things that other people understand. But to me, there is a finite amount of demand.
Demand for new things, yes. New things will have new demand because they're new, but they will replace some kind of an old demand. And if there is no demand for such a huge spike in economic value, and I'm not even talking about having a 20, 30, 40% unemployment. I'm just saying how much demand there's going to be for new stuff.
I don't see that as increasing the economic value by so much money, uh, but I hope I am wrong. Another aspect of the bigger picture comes from the EU AI Act for high-risk provision, which takes effect on August 2026. So the act has been in place for a while now, but it has a section that talks about high risk of autonomous system requirements, which includes mandatory robust human oversight and control mechanisms, and it is going live in August of 2026. So we have about two months from now for anybody who launches large scale agentic solutions in the EU has to have such mechanisms in place.
Now, both OpenAI and Anthropic have shared multiple use cases from actual users in the industry of their agentic solutions in the papers and the new releases they released this week. So documented enterprise use cases from early adopters include things such as drafting and sending emails, pulling data across connected platforms, creating presentations, researching accounts, summarizing calls, posting structured briefs to team channels. and if you want specific examples, Rippling had a sales agent built by a non-engineer five to six hours of weekly manual work per rep.
This is very significant. I can tell you that I'm working with multiple organizations, and I'm seeing what their developers and what their non-developers are building on weekly basis, things that are dramatically changing the way they work, the way they run their operations, the way they help their customers, the way they collect their data, the way they plan their operations on weekly or biweekly basis. It is complete game changers and moving from things that are in people heads right now, and only Jill or Joe knows how to do effectively to fully automated or mostly automated solutions that take significantly more information into account, making the outcome more effective and more valuable to the company and or its clients, while doing it significantly faster, more or less across every aspect of the organization, from operations to planning, to customer service, to marketing, to finance, and so on and so forth.
Every one of these departments I see on weekly basis people who learn how to build agentic solutions build their own solutions because they are the subject matter experts. They understand exactly what the issues are, what the bumps on the road are, what to avoid, and they can build their own solutions right now to augment what they had to do manually. And there are hundreds of such use cases in every organization. On the Claude Tag side, they're talking about incident response, so debugging latency issues, pulling latency and pull request doing any kind of code work you can imagine, doing data analysis, go-to-market tasks, et cetera, et cetera.
Approximately 65% of Anthropic products team code is generated by the internal version of Claude Tag. So it's not just talking about coders writing code. We are talking about people inside of the regular rest of the organization talking to Claude inside of a Slack channel and then the Slack bot doing the work that they need to do So a summary of where we are in the agentic world. The agentic shift is not a forecast, it is happening, and it is happening very, very fast.
Again, if you just look at OpenAI's internal data, 99.8% of token output is agentic right now. Now, they're obviously the cutting edge of how to use their own tools, yes, but 99.8% is pretty freaking high even for an organization that is using its own tools, but it is showing you where this is going.
The other great signal of where this is going is that custom GPTs are being sunset for enterprise customers by August of this year, so this is not a very long time to prep for that. By August 26th, 2026, there is not going to be a way to use custom GPTs anymore, which will force more and more organizations to build agentic solutions or build projects to do similar things The fact that major companies are investing time in building security frameworks for such agents, and the fact that the regulation, at least in Europe, is going to demand such things in place also tells you how fast this is all evolving To me, out of this entire segment, the two most significant numbers is that the growth of non-developers Codex use has been 137X since the launch of Codex.
And yes, this is since launch, so in the beginning it was very low. But again, this is non-developers, and the other half of it, that's 25%, a little more than that, over a quarter of the usage of non-coders is writing code in the back end because this is what agentic systems can do, shows you how fast this is moving. Codex did not exist a year ago, and now we have 25% of users around the world are generating code with Codex to solve problems that they have in their professional lives What does that mean to you?
First of all, it means you need to learn agentic usage, how to build agents, how to understand how agents work, how to plan for agentic world in your universe. That means both from a budgeting perspective, from a data access perspective, from a security perspective, from a training perspective, et cetera, et cetera. There are multiple aspects of this that every single organization, whether large or small, taking very seriously because it's not something that's coming in the future.
It is something that is available right now that is a complete game changer, and if you're not using it, you're going to be left behind. If you are using it, there are multiple unsolved questions because we're all running into this uncharted territory together. So which system should you connect it to? How much access should it have?
Who should have access to the AI systems that has access to your most secure systems? How many tokens do you allow them to use, through which channels? Who protects it? Who stops it when it's going rogue?
Like, there's so many questions, and again, there are no clear answers, and every organization now figures it out on its own. I fully agree with the call we heard last week from the leading labs to have some international collaboration to help at least answer the bigger questions while we all struggle and try to solve what are very big questions for business people and specifically business leaders on how to approach this as an enterprise or even as a small business. What is the right budget?
What is the right approach? Who should have access, and so on. I will keep on reporting how we are doing this in the different businesses that I'm with as much information that I can share, keeping their data safe. But I will start sharing more and more on different use cases and how we're handling these kind of situations because I think this is going to be a critical aspect of AI adoption moving forward.
Now, since we just talked about this crazy race and the agentic capabilities from Google and Anthropic and OpenAI, let's talk a little bit about the people that are making it possible. And there have been some major shift, tectonic moves of some of the most talented, most capable researchers in the world between these labs in the past few weeks So just a few weeks ago, if you remember, on May 19th, Andrej Karpathy, which is one of the most renowned researchers in the AI space who has held roles, he was one of the co-founders in OpenAI and director of Tesla Autopilot and AI Vision, and then Eureka Labs announced that he's joining Anthropic.
He was basically sitting on the fence for a while, and now he understood that he has to play a role in this next era, and he joined Anthropic after being in OpenAI twice as a co-founder, and then after Tesla, he came there again. But now in the last week, we heard two major transitions, both from Google into the two different companies. So the first one is Noam Shazeer, who left Google DeepMind to OpenAI, just announced on June 19th. Now, in OpenAI- Now, why is this a big deal?
Noam Shazeer was one of the co-authors of the 2017 Attention Is All You Need paper inside of Google that created the transformer architecture that is virtually running everything we see in the AI space right now. He then left Google and started his own company called Character.ai, and then in an acqui-hire back in August of 2024, Google brought him and his team back for a sum that was disclosed at $2.7 billion to get licensing to Character.
ai and get Noam and his team back to Google, and he was working in Google until right now, and now He is leaving and going to OpenAI, and he wrote on X, and I'm quoting, "I'm excited to share that I'll be joining OpenAI and look forward to working with the exceptional team there. It was a difficult decision to move on. I'm incredibly proud of the amazing team at Google and everything we built together. It has been an honor and a pleasure to work with all of you."
Now the reason just understand research, but also understands how to build products in an effective way But he was not the only person leaving Google DeepMind in the past week or so has announced. Now Jumper is a Nobel Prize in chemistry winner together with DeepMind CEO Demis Hassabis he has been sharing the path with Dem, who is the CEO and the founder of deepMind and Google. and he's now departing and joining Anthropic Jumper wrote on X, "GDM," that stands for Google DeepMind, "is a special place, and I still be excited to what amazing thing they discover next."
Now, what did these announcements do? Well, first of all, Google's Alphabet stock has dropped 7.2%. The second announcement one of the largest dips they had this year.
The decline ended being only around 5 to wiping a $264 billion in value. To be fair, the entire tech was doing great that day. The Magnificent Seven lost 2.2% that day, and the overall tech market was more or less flat.
But Google lost at the peak 7.2% and ended the day between 6% and Now, that basically tells you how critical the market thing these two rivers are for the future success of these companies. On the flip side, both companies, OpenAI and Anthropic, where these people went to, so are part of Sam Altman's jump are Anthropic on the verge of an IPO. So two of the most capable researchers in the AI space in the last decade and a half are both joining Anthropic Now, Google for years were able to keep their zero trend identity and were very proud of that they are leading out there that is attracting talent of who want to do pure AI research, and apparently that's not the case anymore.
They're not the first big names that are leaving. Dave Silva was one of them as well, left earlier to start his own company Jonas Adler and Alexander Pritzel, also top researchers in Gemini, left earlier. So it's been not just a few drops, but a serious exodus at this point of top leading researchers who are leaving Google DeepMind and moving into other places, either their own organizations or joining one of the other competing labs. Again, this doesn't look good for Google, especially if you combine it to the fact that their models are very far from the frontier right now.
they're even now trailing behind some of the open source Chinese models. I think we should all expect, and there's been a lot of rumors about this, that the next model from Gemini is going to be a very big step forward. But that being said, they promised that model to be released in June. They did that in their IO conference, and June is almost about to end, and there's no model in sight.
So maybe they will release it in the next few days before the end of the month. But as of right now, there aren't even any rumors, and there's complete radio silence from Google on when 3.5 it is going to be released I can think of, which is not rocket science, of why it hasn't been released yet, is that it's just not good enough, and it is not aligned and definitely not ahead of GPT 5.5.
And now we're learning that there's also 5.6 already in the making and maybe even circulating without an actual tag and being tested by specific individuals, Not to mention Anthropic Fable and Mythos with all the mess that they've created as far as their real advanced capabilities. So Google is not looking great right now, and they are not releasing products that are competing with the leading edge of AI, and they have top-notch, literally the best researchers in the world leaving them, and the decline in the stock price is not surprising in that scenario On the flip side, Sam Altman is obviously very happy and excited, and he posted on X saying, Noam is one of the people I have most wanted to work with since the very beginning of OpenAI.
Only took 10 years. I think it will be worth the wait." So again, on both sides, it's a very big loss on one side and a very big win on the other side, which puts Google in an even worse situation because it's not only he's leaving Google, he's joining the competition. Now a few final thoughts on this.
Thought number one is I don't think these departures are only for money, and if anything, it shows the other way around. It shows that right now these people will jump ship and will move to where they think the research is going to be most significant. Pulling Andrej Karpathy into Anthropic came because his work was aligned with the work that Anthropic is doing right now on self-recurring improvement And I assume similar things are happening with the other depart-departures that we just mentioned.
Again, if you think about the money, Noam Shazeer was paid 2.7 billion. Again, not he personally, the money went other places as well, but I'm sure he had a nice big chunk of change, and I'm sure Google were willing to pay him or continue pay him crazy amounts of money to keep him there, and still he's moving to a different company. That tells you that, again, right now being at the leading edge of research is what matters to these people most.
And I'm personally wondering what does that mean for other researchers. So these things tend to have a momentum, right? When you see top researchers leaving your company, that tells you that there's better stuff, more advanced research done somewhere else. And when it's not one person, but several people in just a few months, the top leaders, it definitely hints where the wind is blowing, and that may lead and departing from Google's DeepMind, which was definitely the leading research lab.
They did an incredible comeback after their very lazy and not very successful start. And now in the past year or so, or definitely six months, they're not doing great on any perspective. It will not be crazy surprising to me to see Demis, who has been leading DeepMind for so long, move to somewhere else. Because if you ever heard Demis, and again, I highly recommend you watch the movie about AlphaFold, if you hear him speak, all he cares about is using his brain power while he still can for the better of humanity.
He literally feels strongly that every minute he doesn't do stuff like that is a waste of his capabilities. And again, if he will feel, I think, again, I know nothing, if he feels that he can do better for humanity somewhere else that is not DeepMind, he may leave as well. I obviously don't have a crystal ball. I don't know what's going on inside of DeepMind.
I don't know if Demis may be doing everything he wants, and that's why it's slowing down the other type of things that DeepMind could have been doing from a product perspective. I don't know any of this, but it is a very delicate situation for Google and very exciting times for AI research as a whole. now since we mentioned AlphaFold, it will be a good segue to talk about AI in healthcare. And this week we got several different stories that all combine to show you that AI in healthcare is gaining momentum, which I feel is maybe one of the most promising aspects of AI from a better future to humanity perspective.
So Boston Children's Hospital published a study that is showing that AI resolved 18 previously unsolved rare disease cases out of 367. So that's doesn't sound like a lot, but it is 4.8% of additional diagnosis that yielded after humans failed to be able to to diagnose these situations Now, for families of these 18 patients, that means a new future for their loved ones that would have died without the usage of AI tools in these scenarios It is even more exciting because these diagnosis that these tools found come from different fields of science.
The first one is neurodevelopmental, the other ones is neuromuscular. Two were unexpected deaths that wasn't identified for the reason, and two of early psychosis. And that tells you that these tools are not good at a specific field of science, or in this case healthcare, but they just know how to connect the dots between multiple data points in ways that humans just cannot do it on unresolved cases, the AI tools were tested against solved cases to verify that they can actually do the work.
And what happened was that the AI was able to identify forty-eight out of the fifty-one rare disease cases that it was presented in one aspect, and then forty-five out of fifty-seven on neuromuscular cases, and all fifteen long read cases were all identified correctly by the AI. So AI is getting the ones that humans were able to identify very well, as well as it was able to crack some of the ones that humans were not able to do On another aspect that is related to healthcare, Midjourney, yes, the image generation company Midjourney, is pivoting very hard into medical hardware.
They are committing over $74 million to build a whole body scanner that will be able to deploy 50,000 units of it by 2031. That's their plan. Now, I highly recommend you follow the show notes and click on the link to the video. It is a spherical tank of water that the person stands on a little platform that lowers him or her slowly into the water and generates something that looks like a really advanced scan, only it does it in a fraction of the time of existing scanning tools.
So if you've seen an ultrasound done, you know, or if you've done an ultrasound, you know it takes a very long time. It is really noisy and uncomfortable. This thing does it in 60 seconds and gives you a complete 3D body map in a fraction of the time of a traditional MRI. And what they're planning to build is a, what they call Midjourney Spa, which the first one they're planning to open in San Francisco in 2027.
That's next year. So they are planning to take what they've learned in AI imaging and apply it to body imaging with a gazillion sensors around a water tank, which again, I find really, really exciting. I don't know what's going to be the cost of this. I don't know what's gonna be the cost per scan.
But assuming it's also gonna be much cheaper than MRI, and it will allow people to do MRIs once a year to find anything they have, this could be completely revolutionary when it comes to preventative care and being able to identify things in the body way ahead of time, because every person in the world will be able to have an MRI on a very high frequency with AI identifying what's actually happening in the images and raising flags as needed to find early signs of anything you can imagine Now, the 60-second number is their future goal.
Right now, it takes 20 minutes to do the scan, which is still shorter than many MRI scans that are done today So far they've tested it only on a dozen individuals. So again, very limited data right now, but very promising based on the results they achieved so far Now, there are many concerns or skepticism from the medical field. Can this work? Ultrasound has limitations, which is how this system works.
Uh, it cannot penetrate bone or air or deep soft tissues. But the reality is they can now generate very serious results, and they have some serious funding behind them, and it will be very interesting to see where this goes, and I really hope they can make this successful because I think this could make a real difference in people's lives the third topic we're going to talk about related to healthcare is the FDA just approved the device called First Read by Aidoc analyzes chest radiographs and generates preliminary radiology report in text It is the first report drafting tool to receive an FDA designation And it is built in the same architecture as Aidoc's other FDA-approved tool that is called Abdominal CT Triage tool that does similar things with abdominal CT.
It is currently already deployed in nearly 2,000 hospitals worldwide, and it has analyzed already more than 120 million scans, and they're now going to do a very similar thing in chest radiograph reports Now, this is exciting from two different reasons. One, it will give accurate and very fast results to people. Two, is that there's a very serious shortage in this skill right now. The Numan Health Policy Institute study found that outpatient imaging interpretation turnaround times more than doubled between 2014 to 2023.
Now, getting the read fast may save lives or suffering from people by getting their diagnosis faster and starting treatment, the right treatment faster. So being able to use this with AI tools, or at least get the first batch with AI tools and then have a doctor look at it, can dramatically increase the speed of delivering the right results to people based on their scans, which will lead to faster treatment. And again, if you combine this with better scanning capabilities, faster, cheaper scanning capabilities like the one that Midjourney are planning to build, you see the perfect storm.
I am very optimistic and bullish about AI in healthcare, and these are just several examples of where this is going. The last component when it comes to healthcare comes from NVIDIA this week. They just announced that they have built what they call BioNeMo Agent Toolkit which is covering structure prediction, molecular generation, docking, sequencing analysis, molecular design, and genomics. So basically, a lot of things in the sub-cell level that can now be simulated or understood in a better way using AI tools Now integrating Bionimo Skills double the agents that are currently using it.
Token efficiency and task completion rates jump from 57% 0.1 before using it to 100% when using Codex CLI with ChatGPT 5.5. This tells you that combining multiple AI tools into a specific goal allows you to achieve significantly higher rates, which means if you're now using just one tool for highly complex things, you maybe need to reconsider.
Definitely in the medical space, I'm anticipating to see a lot of interesting collaborations with companies or tools or gear that is coming from multiple sources in order to solve very specific problems So what does that tell us? It tells us that AI is now good enough. Again, it's passing FDA approval. AI is good enough to help in healthcare processes on both the research side as well as on the scanning almost and then diagnosis of scan side.
And I'm certain we're going to see more and more healthcare use cases that are truly going to change people's lives in the immediate and definitely in the long-term future. So while we talk a lot about the negative implications of AI across the board in different aspects, these are all very promising healthcare support systems and tools and processes that I know will help make our lives better Now this was a lot of stuff that we talked about, so I'm going to mention just one thing in the rapid fire section, and that's the announcement from OpenAI that they have just unveiled Jalapeño, which is their first custom AI chip that they've built together with Broadcom So this will allow OpenAI to compete with companies like Google and Amazon and Microsoft that already have their own homegrown chips.
This has been in development for nine months, which is a very short amount of time to bring this from design to manufacturing Now, OpenAI stated obviously that Jalapeño performs, and I'm quoting, "substantially better than current state of the art," but OpenAI have not released any benchmark or technical specifications, so there's no clear understanding of what does that mean and does it actually perform better and of the AI space. But it is, first of all, exciting to have another option in that field.
Having more options usually lowers the cost of everything But what it does mean, it means that now all the major players have their own chips to use in the AI space. They're all probably gonna rent some from other players to get extra capacity and not just build on their own, but having your own chip is definitely helpful. What OpenAI said, and I'm quoting, is, "We optimize the architecture around the kernels, memory, movement, networking, and serving patterns that matter most for the frontier AI models.
Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware's theoretical limits." There are many additional pieces of news this week that you need to know, including rumors on when we're getting Fable back or what will it take to bring it back, And there are a lot of rumors about new model launches in the immediate future, so maybe next week we'll have stuff to share. But if you want to learn about the things that did happen this week that we did not cover in the podcast, you can sign up for our newsletter.
You can do this in the show notes. There's a link to do that. In the show notes, you will also find a link to learn more about our multi-agent orchestration course that, like I said, is selling like hot cookies. And so if you want to join, the last few seats of the August cohort are there, but that's it, and then you will have to wait through September.
So if you want to learn this course, and you probably do, come and join us in August and sign up right now. That's it for today. We will be back on Tuesday with another how to dive on how to use specific use cases with AI in the business world. And until then, enjoy and have an amazing rest of your weekend.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.