The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/HR/Hidden Layers
Hidden Layers artwork

AI Year in Review - Key Moments, Hot Takes, and 2026 Predictions | EP. 48

Hidden Layers · 2025-12-17 · 41 min

0:00--:--

Key moments - from our scoring

Substance score

57 / 100

Five dimensions, 20 points each

Insight Density11 / 20
Originality10 / 20
Guest Caliber11 / 20
Specificity & Evidence13 / 20
Conversational Craft12 / 20

2025 proved to be a defining year for AI adoption and capability maturity, marked by the rise of multimodal interfaces, aggressive enterprise competition, and a complex regulatory landscape. Emma Perchowski highlights how speech-to-speech conversation, image analysis, and video transcription through tools like ChatGPT, Claude, and NotebookLM have shifted AI from work benefit to workflow transformation. Michael Wharton documents Anthropic's stunning market share ascent - from 12% to 32% enterprise LLM API usage in three years - driven by focused developer tooling (Claude Code, Claude Skills) and transparent safety research, while OpenAI's market share declined from 16% to 9%. Ziv Tsai emphasizes coding agents as a proven killer app requiring daily iteration but delivering 2-3x productivity gains; video generation models (Sora, Google's Vu 3, Alibaba models) and persistent world models (Marble) advanced from GPT-2-to-3 scale jumps; and data center buildout accelerated competition. Reasoning models (DeepSeek, o1, o3) emerged as frontier performance drivers, while Google's turnaround - compelling OpenAI's code red response - shifted the competitive dynamic. Regulatory tension peaked as adoption and fear both hit all-time highs despite minimal federal guardrails, with Amazon's ReInvent announcements positioning AWS as enterprise tooling consolidator via NovaForge's mid-training capability. Predicted failures materialized partially: IBM data on AI-related breaches averaged $10.2 million per US organization, validating concerns about governance and integration risks.

Key takeaways

  • →Multimodal AI interaction has evolved from text-based prompts to seamless speech, image, and video processing that fundamentally changes how professionals work daily.
  • →Anthropic has captured significant enterprise LLM API market share (growing from 12% to 32% over three years) by targeting developer workflows and safety-conscious early adopters while OpenAI's market share declined from 16% to 9%.
  • →Coding agents remain iterative tools that require consistent human oversight, with reliable unsupervised work limited to approximately 10 minutes before verification becomes necessary.
  • →Reasoning models from DeepSeek and OpenAI (O1, O3) prevented a performance plateau and have become daily-use features in ChatGPT and Claude.
  • →Video generation models like Sora, Veo 3, and Genie 3 showed significant 2025 advances but haven't reached production adoption due to control and consistency challenges compared to text generation.

In this episode

  1. 1Welcome and Launch of Working Models Podcast
  2. 2Top 2025 AI Developments: Multimodal Models and Regulation
  3. 3Anthropic's Enterprise Market Success and Competitive Landscape
  4. 4Coding Agents and Reasoning Models as Key Breakthroughs
  5. 52025 Predictions Review: World Models and Video Generation
  6. 6Amazon's Enterprise AI Strategy and Nova Offerings
  7. 7AI Safety, Failures, and Data Breach Costs

Mentioned

Ron GreenEmma PerchowskiMichael WhartonZiv TsaiOpenAIAnthropicGoogle DeepMindChatGPTClaudeNotebookLMAmazonGoogle

Guests

Michael WhartonEmma PerchowskiZiv Tsai

Topics in this episode

ClaudeChatGPTOpenAIAnthropicNotebookLMDeepSeekGoogle DeepMindSORAVeo 3Genie 3

Questions this episode answers

How much has Anthropic's market share grown in enterprise LLM API usage?

Anthropic's enterprise LLM API usage market share grew from 12% to 24% to 32% over three years, while OpenAI's declined from 16% to 9%, driven by Anthropic's focus on developer workflows and targeted early adopter positioning.

What are the leading video generation models in 2025?

Key video generation models in 2025 include Sora from OpenAI, Vu 3 from Google DeepMind, Alibaba's open-source models, and Marble from World Labs, which features persistent 3D representations and navigable world generation capabilities.

Can coding agents now autonomously complete complex tasks without human oversight?

Coding agents can work unsupervised with high-quality output for approximately 10 minutes; once tasks exceed one hour, human verification is rarely avoidable, though this unsupervised capability window is expected to extend over coming years.

What is AWS's NovaForge offering and what is its cost?

NovaForge is AWS's $100,000 per year offering that enables companies to train their own models, including pre-training, mid-training (where companies see the most benefit), and post-training capabilities on AWS infrastructure.

What was the average cost of an AI-related data breach for US organizations in 2025?

IBM's study quantified the average cost of an AI-related data breach for US organizations at $10.2 million, reflecting governance and security risks from enterprise AI integration without sufficient safeguards.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

11 / 20

There are pockets of genuinely useful data (market share shifts, breach costs, human-in-the-loop statistics), but much of the episode restates widely-known 2025 AI trends and personal usage anecdotes that a following operator would already have absorbed.

about half of their API calls are related to human augmentation, like human assistants. And the other half are pure automation
building agents is fun, but getting value from them is hard

Originality

10 / 20

Most takes - agents are overhyped, humans-in-the-loop matters, ROI reckoning is coming - are the consensus narrative circulating everywhere; the GPU-datacenter-in-space and memory-system predictions add some freshness but remain speculative rather than argued from first principles.

I think we're going to have a, uh, GPU data center in space
I would say, by and large, they don't deserve the hype and credit that they're getting right now

Guest Caliber

11 / 20

The panel are practitioners at an AI consulting firm (VP of engineering, distinguished engineer, AI strategist) with hands-on experience, but they are internal colleagues rather than senior operators who have run AI at scale in large enterprises.

Michael Wharton, VP of engineering, and my co founder and distinguished engineer, Ziv Tsai
Emma Perchowski, senior AI strategist

Specificity & Evidence

13 / 20

The strongest dimension - the discussion cites concrete figures like Menlo Ventures market share numbers, IBM breach costs, unemployment rates, and Nova Forge pricing, grounding claims in real data rather than hand-waving.

Anthropic... went from 12% to 24% to 32
it totaled up to about $10.2 million per organization

Conversational Craft

12 / 20

The host asks some sharp clarifying follow-ups (coding-agent acceptance rates, monthly bills, recapping predictions) and there is mild probing, but it remains a friendly internal panel with little genuine disagreement or pushback.

How frequently are you getting results where you feel you can just accept them?
how big is your monthly anthropic bill?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B32%
  • Speaker C28%
  • Speaker E25%
  • Speaker D15%
  • Speaker A1%

Most-used words

models24agents19anthropic14data13world12feel12openai11training11interesting10impact10important9market9coding9human9space8value8

Episode notes

2025 was another defining year for artificial intelligence. In this special AI Year in Review episode of Hidden Layers, Ron Green is joined by Emma Pirchalski, Michael Wharton, and Dr. ZZ Si to break down what actually mattered in AI this year. The team recaps the biggest developments from 2025, revisits their predictions from 2024 to see what held up (and what didn’t), and shares honest, experience-driven predictions for 2026. Topics include multimodal models, agents, enterprise adoption, governance gaps, workforce impact, ROI pressure, and where AI is truly headed next. This episode cuts past hype to focus on what leaders, builders, and decision-makers should actually be watching as AI moves from experimentation to execution. Chapters 00:00:00 Welcome and Introduction to 2025 AI Year in Review 00:00:56 Emma's Working Models Podcast Announcement 00:01:48 Top AI Developments of 2025 00:16:29 Reviewing 2025 Predictions 00:25:08 2026 Predictions 00:36:49 Closing Thoughts

Full transcript

41 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Foreign.

Speaker B: Welcome to Hidden Layers, where we explore the people and technology behind artificial intelligence. I'm your host, Ron Green. This is our 2025 AI year in review episode. We'll cover the most important events that happened this year, including AI agents, multimodal models, and even AI in space. We'll also review our predictions from last year and own up on what we got right and what we got wrong. And finally, we'll share our forecast for 2026 and what we think you should watch closely. I'm joined by three of my amazing colleagues. Emma Perchowski, senior AI strategist, Michael Wharton, VP of engineering, and my co founder and distinguished engineer, Ziv Tsai. All right, I'm really excited about this. We're going to jump in to cover, uh, what we think are the most interesting things from 2025. But before we do, Emma, I think you have some really interesting news to share first.

Speaker C: I do have some interesting news. So in early 2026, I'm launching a podcast called Working Models, which is all about AI in the real world. And if we think about all of the technical advancements that y' all cover so well in hidden layers here, you know, new papers, new models, new capabilities, um, this podcast is really intended to sort of explore the other side of that equation, so how AI is showing up in businesses in the real world. And to do that, I'll be speaking with C Suite executives, heads of strategy, corporate board members, educators, regulators, you know, really the people that. Shaping the practical application of AI. So I am very excited. And yeah, stay tuned for the Working models podcast in 2026.

Speaker B: That is awesome.

Speaker D: Super excited.

Speaker B: I can't wait. I can't wait. Maybe, maybe you can have some of us on occasionally. Just, uh, absolutely.

Speaker C: Open invite.

Speaker B: Okay. All right, so I'm really excited about this because, you know, I don't think it's an exaggeration to say 2025 was, again, an insane year in AI. Right.

Speaker D: Um, lot.

Speaker B: A lot of big advances. And let's kick off. I want to just hear from each of you what you think your top 1, 2, 3 biggest things that happen in the year. And I know that's really tough because I think we could probably have a list of, you know, 20, 30 really, really big items. But let's just kind of win this on. If somebody's listening to this podcast and they want to catch up on 2025, what are the top, top things that happened this year? So, all right, I'm going to kick off, uh, I'm going to kick off with you Emma, uh, what would you go with?

Speaker C: Okay, so the first one I would say is really sort of my personal, I'm most excited about, which has really just been the evolution of multimodal interaction with LLMs. Of course, the capabilities have improved so much in the past year. But I just think my personal use, if I think about when we were sitting here a year ago, it was, you know, mostly text based prompts. I was using it every day. Um, saw a benefit to my work for sure. But now I would say I'm using speech to speech all the time. I'm, you know, uploading images, the ability for it to parse through video and transcripts and take out next steps and notes and just all of these things that I think we were sort of promised from the beginning I feel like are very frictionless now. And I use it more in my personal life. And again, whereas a year ago I would say it benefited my work, I feel like now these tools really change the way that I work.

Speaker B: Yeah, um, what's your go to? What's your go to if you're using a model?

Speaker C: Uh, I mean it's probably chatgpt. Um, but I do use Claude a lot for more like brainstorming and writing I really like. And then um, NotebookLM. Um, you know, use that for a lot of like interviews and video transcripts and stuff.

Speaker B: But I, I, yes, I am definitely using NotebookLM way more than I anticipated. In fact, if I had heard, if I had heard the product briefing for that, I'd have said no go. That just sounds like a crazy idea. Yeah, this idea that you could take a bunch of documents and generate a podcast of it, but it really works amazingly well. What are you using for a speech to speech, by the way?

Speaker C: I'm using ChatGPT and I'm like having conversations with it.

Speaker B: Okay.

Speaker D: Yeah, that one's very good.

Speaker C: Yeah. M. So the next one, this is such a big topic, but if I could characterize it, I would say the focus, thinking about the Frontier Labs and the US government, this focus on pace of innovation and balancing that with accountability. So just as an example, um, to paint the picture, I would say right now adoption's at an all time high investment's in an all time high investment of course has been kind of concentrated and they're circular investments. There was the big headliner study from MIT, about 95% of Gen AI pilots fail. So there's like so much of the economy and of business riding on this investment in AI. But at the same time, um, the Pew Research that Just came out shows that fear is at an all time high. Uncertainty in our ability to manage these tools is at an all time high. Bipartisan, you know, both sides of the political spectrum want some sort of regulation and have very low faith in both government and organizations ability to manage this. So I just think it's like a really interesting and important dynamic to be talking about because if we look at the regulation, you know, in January at the beginning of this year, the Biden executive order was rescinded and replaced with Trump's executive order on AI. In July, I think it was, the Americas AI Action Plan was launched, which was really focused on kind of national leadership in AI innovation, building out technical infrastructure, and did not place much emphasis on like safety or reporting or transparency. And so again, it's kind of this time where everyone's invested in AI, we're all using it. There's very little kind of accountability and regulation and everybody's fearful about it. Um, but you know, people are also excited and just to kind of close that out, the main guardrail that's been in place from a regulatory standpoint has been on the state activity. We talked about that last year.

Speaker B: Yeah.

Speaker C: Where there was a ton of state regulation. And as recent as, I think just a few days ago, there's sort of chirps of, you know, potentially an executive order that would ban state AI regulation for 10 years. The goal there is sort of to have one federal policy which I think would be fantastic. But the issue is that hasn't been proposed to my knowledge. And so the fear is we would just ban state AI regulation. Have a vacuum and have. And have nothing.

Speaker B: Yeah.

Speaker C: Um, so, yeah, I think that dynamic. I'm really curious to see what we'll be saying a year from now.

Speaker B: You know, one of the things you mentioned is like the fear of AI being at the all time high. I saw a chart the other day. It was essentially, um, um, excitement and fear and anxiety around AI across all the different countries in the world. In America, we lead. We're the most fearful, uh, most anxious about AI.

Speaker E: It's funny that we have the least regulation, then what the heck.

Speaker C: And we're using it the most.

Speaker B: Yeah, we use it the most, we regulate it the least and we have the most fear. I don't know, maybe that goes together. Maybe that's part of the reason.

Speaker E: Yeah, yeah.

Speaker B: I don't know. Okay. All right, all right. That's awesome. Um, all right, Michael, why don't you go next?

Speaker C: Okay.

Speaker E: So this was. I struggled preparing for this because there really was a lot that happened this year. Like reasoning models became important. There's, there are more, there's more competition. There are more players that are actually, you know, competing with OpenAI. Like they had a pretty big lead coming into the year and now there's just more frothy competition, which I think is great. But anybody who's heard me talk on this podcast before has heard I'm a huge fan of Anthropic. I think they're a great company.

Speaker B: Yeah.

Speaker E: And it's not necessarily a David and Goliath story by any means, but they're getting a lot of traction right now and I think it's wonderful. Like, I was actually looking at a, uh, it was a Menlo Ventures summary of market share. And over the last three years, Anthropic, at least when you look at enterprise LLM API usage, they went from 12% to 24% to 32. So they, they're just gobbling up market share of enterprise LLM usage.

Speaker D: Yeah. That's wild.

Speaker E: Is crazy. And then if you look at OpenAI, they started at 16 and they're down to like nine now.

Speaker B: Yeah. It is kind of weird. Like Anthropic is going for the enterprise API market and it feels like OpenAI is mostly going for the commercial market. Right?

Speaker E: Yeah.

Speaker B: And it's not a question that you say AI to most people and that they chat GPT.

Speaker E: Yep. For sure. Yeah. I mean they're doing everything everywhere all at once, it almost seems, where they're just casting a really wide net and going for the scale at all costs approach.

Speaker B: I think maybe to their detriment, they might need to focus a little bit.

Speaker E: Well, and that's the thing that's amazing about Anthropic. I mean there, I think you have this thing about AI adoption that kind of reminds me of Tesla where there's this corner of the market that they, they cornered at the beginning of, you know, releasing the vehicle where it's really aggressive. Early adopters, people that just readily want to try new technology. But to get beyond that, it's a really big uphill climb to just get into the broader consumer market.

Speaker B: Yes.

Speaker E: Anthropic has targeted aggressive early adopters like developers like us and just nailed it. Just fit right in the market like that.

Speaker B: Yes.

Speaker E: And things like Claude code, their whole developer workflow experience, like Claude skills, now they're clearly just focusing and doing it well and just gobbling that up. And I think that just the fact that they gobbled up that part of the market so quickly and succeeded is Just really cool to see, especially because they have such a backbone compared to a lot of the other Frontier labs.

Speaker C: Just to react to that. Uh, I'm curious to get your thoughts on. You know, obviously, Anthropic has also sort of positioned themselves as like, the responsible kind of, you know, company that really values safety as well. Um, do you think that that has played a role or maybe as the public sentiment is where it is with AI, people will value that more?

Speaker E: Yeah, I think so. I mean, I think they're, you know, they're really great about transparency and openness and actually dedicating time and funds to. To research that benefits that transparency, like, you know, their work on circuit tracing. To just try to look at how models think with, you know, the primary intention of governing them and actually trying to put guardrails on them to make sure that they don't misbehave.

Speaker C: Yeah, those 60 minutes. Yeah, that came out.

Speaker E: Oh, my God. Yeah. Like the vending machine that they show that. That internal, like, company that they started, that's vending snacks.

Speaker C: I was actually thinking of, um, how they showed the blackmail when they were doing. Oh, yeah, testing. Yeah.

Speaker E: Yep. That's right.

Speaker C: Yeah.

Speaker E: Yeah. I mean, they're just so much. They're casting a wide net, but they're. They seem almost like a tortoise. It's like a tortoise and hare story. They're very intentional about each step, and I just. I like them as a company. I'm a little bit of a fanboy, if you can't tell.

Speaker B: I totally agree. Well, and, um, not to mention, that was a week and a half ago that, um, uh, Sam Altman at OpenAI purportedly issued a code red.

Speaker E: Oh, the Gemini thing.

Speaker B: And that had nothing to do with Anthropic. Right. That was all about Google, which is amazing. Um, we'll come back to that in a bit. Okay, so. All right, so Michael just kind of put a ball on that. You think? Looking back, um, the work that Anthropic's done is the most interesting thing in 2025.

Speaker E: It hits home. It's personal. Not everyone's going to think that, but their rise to prominence and success is just really exciting to see. Yeah.

Speaker B: Yeah. Okay. That's awesome. All right, Z.Z. what do you got?

Speaker D: Yeah, I totally agree with, uh, what you picked as a top. And, uh, it's a really hard choice because there are so much happening. And my top three is coding agents, video generation and world models and data, uh, center build out.

Speaker E: I can't believe we didn't Cover that yet.

Speaker D: So for, uh, coding agents, um, well, it's kind of, uh, surprising, but not so surprising that this becomes kind of a killer app category for Genai. So it's like, um, as Michael mentioned, um, uh, the adoption in Enterprise has been so big, and Anthropic is leading there. Um, and just looking at how much people are using it day by day, and personally, I'm using it every day, every working hour.

Speaker B: Yeah, you're using it insane.

Speaker E: You're melting GPUs.

Speaker B: Yeah,

Speaker D: my, uh, output has literally, I think, doubled or even tripled. Um, and, um, looking at the models, Anthropic is doing great. Um, I think between the competition, uh, of, uh, anthropic, um, and OpenAI, um, at one time, I think Codex from OpenAI took a brief lead, but then when the Opus 4.5, uh, is out, I tried it out. It's just so great. So, coding agents.

Speaker B: All right, let me ask you on coding agents. How frequently are you getting results where you feel you can just accept them? Are you at the point now where most of the time you're happy with the output, or is it still very iterative?

Speaker D: That is a great question. It's still very iterative, and I think it's, you know, um, people say, uh, that the year of 2025 is the year of agents. I agree with that. And then, you know, there's also opinion from folks like Andrew Capaldi, uh, saying it's the decade of agents. Yeah, I feel like it may be still early, but, you know, right now, in my workflow at least, I need to nudge the coding agents consistently. Just like how you have one on one with your team. Kind of.

Speaker B: Yeah.

Speaker E: Okay. So, um, how big is your monthly anthropic bill?

Speaker D: That's a great question. I mean, I use subscription, so that helps a lot.

Speaker E: So you got the max subscription?

Speaker D: Yes, I got a max m. Yeah,

Speaker B: of course you do.

Speaker D: Um, yeah. So, um, I mean, for coding agents, I got a lot more output, but also, um, I don't feel like I am, uh. You know.

Speaker A: Um.

Speaker D: Well, I also have a lot of anxiety because of the. You know, I'm more busy than before. Yeah, definitely.

Speaker B: Because if I understand you're doing multiple. You're essentially giving coding agents multiple tasks simultaneously, meaning you're not focused. Um, you know, if you're writing code, you can kind of focus on one thing at a time. Now you're kicking off agents to go handle multiple tasks concurrently in different workspaces.

Speaker D: Yeah, exactly.

Speaker B: Yeah.

Speaker D: And right now, they can work unsupervisedly uh, with high quality content for like minutes, like 10 minutes. But once you, uh, go over like one hour, it's very rare that you can just blind, like accept the output without verification.

Speaker B: Yeah.

Speaker D: But over the next few years, I'm looking forward to this length of unsupervised, uh, work getting longer.

Speaker B: Yeah, yeah, yeah. Okay. I know I'm not on this list, but I'm going to go ahead and share what I think. I think that the reasoning models that came out this year were really important because outside of those reasoning models, I think we actually would have seen a little bit of, uh, a plateau on performance. Right. Most of the models that are really moving the frontier are reasoning models and they've really only been out about a year. I mean, DeepSeq has not even been out a year. 03's been January, wasn't it? It was January. Right. Uh, 03 just slightly predates that. But I think that we're seeing the use of, uh, uh, chain of thought reasoning on a daily basis as you're interact with ChatGPT and Claude and things like that. So I think that's huge. And my honorable mention, I have got to go with Google. I'm just stunned that they turned things around.

Speaker E: Come back two years ago from that duck thing.

Speaker B: Yeah, I mean, my God, from two years ago, I thought they might have been toast. And they have completely turned around to the point that, like we just said, OpenAI issued a code Red. And that's a big deal because, um, you know, OpenAI has to fund all of its research, uh, and training and inference off subscriptions and through capital raises. And Google can do that through cash flow. So, uh, I think there should be a Code Red. Okay. Are you guys ready to transition? I know this is the tough part. On before we get to predictions, let's talk about your predictions for last year. And I've got them right here. So no cheating. Okay.

Speaker D: Moment of truth.

Speaker B: This moment of truth. All right, so just. Yeah, so let's look at this. All right. ZZ, you said for 2025, a year ago you said you think it will be the year of world models. All right, explain that and then explain whether you think you were right or wrong or m in the middle.

Speaker D: Yeah, I think I'm like half right, maybe. Um, but we do see a lot of development, uh, in video generation models in 2025. Um, uh, in terms of, just to put things in context, GPT went through 2 to 3 and then ChatGPT and then 4. Those are the big jumps. I Feel like year is kind of like from GPT 2 to 3 kind of jump a lot of breakthroughs in models like Vue 3 from Google DeepMind, uh, the Sora from OpenAI and also open source models like one from Alibaba. So much um, higher quality and also much more efficient to trend and generates videos uh, longer and faster. Um, and those um, also kind of goes hand in hand with world models which starts from video generation models like next frame prediction as you navigate the scene to like um, Genie 3 from Google DeepMind again is another great example that you can prompt it to generate a world that you can navigate and it even has a sense of itself. It will generate a room where a person is typing on a computer that is uh, on the computer. There is the Genie three generating things um, and then there's the marble from World Labs. I um, guess it's one of the first world model product out there uh, that uh, not only can predict the next frame but also has the uh, persistent 3D representation that um, is really uh, impressive. So I feel like maybe half.

Speaker B: Right, okay. We're still early so I'm so blown away by uh, these uh, generative video models and things like that and especially the ones where you can kind of move within them like you know, M and alter the world and it has persistency and things like that. But so to what extent do you think you were wrong if we have that now?

Speaker D: Um, it hasn't gotten into production as much as I hoped uh for which kind of highlights text and video are just very different modalities and hopefully in the future people can trim model that can be more easily controlled, can be really helpful for artists, for movie making, for making marketing videos.

Speaker C: I was going to ask what are some places you would love to see it applied?

Speaker D: That's a great question. Yeah, I mean uh, entertainment like you know, and also making educational videos because you know we, we human, we we really love to consume videos like the short videos, long videos.

Speaker C: Right.

Speaker D: Uh, it'll be great to make but producing such such videos is still very tedious job.

Speaker A: Mhm.

Speaker D: And I think it can be made better.

Speaker E: Maybe Yann Lecun will figure it out and that why he's leaving Meta right now to go pursue world models.

Speaker D: Could be.

Speaker E: Yeah, maybe we'll get a lick.

Speaker B: All right, Emma. All right, you ready? So I've got it right here. Amazon uh, is about to show up in a big way in 2025.

Speaker C: Just made it.

Speaker E: Maybe re Invent was what?

Speaker C: Yeah. So you know I think it is still of course it's still playing out. I think it's going to take years.

Speaker B: You have a couple more weeks.

Speaker C: Yeah, a couple more weeks. We'll have to come back. But so re Invent was uh, recently and you know they had just so many announcements that came out of that. The overall sense that I get is they're really kind of banking on being like the one stop shop for enterprise tooling. So do I think they're competing from a model perspective with like Anthropic or Google, uh, or OpenAI? No. But I also don't think that that's really what they're trying to do. And so just looking at some of their announcements, um, you know novaforge, they're expanding the uh, their open weight models on Bedrock. I think they're adding 18 new models to that. And Nova Forge seems really interesting where my understanding is companies can basically pay $100,000 a year and will be able to train their own models. But what I think is really interesting within that is it's pre training, mid training and post training. And I think the mid training is where um, companies presumably are going to see a lot of benefit for that. And I think they're sort of banking on this idea that companies are going to value the ability to choose, kind of choose their models rather than just go with what you know, whoever the top leader is at the time, um, and the trade off cost like to someone that is, you know on their infrastructure is going to continuously go up as they become this one stop shop. So maybe, maybe still's easy like partially. Right. I think we're definitely still going to be talking about this a year from now. Um, but of course this announcement was very recent so. Yeah, we'll see.

Speaker B: The thing I'm most interested in is that that hundred thousand uh, dollar offering where you can, you can do that, you can not just do post uh, training or fine tuning on corporates, you can actually bake some of that knowledge in the model through that sort of um, not pre training but mid training process. And I think that's really interesting. I don't know anybody else really doing that at scale. All right, that's really cool. All right. Michael, how about you?

Speaker E: Okay. So yeah, last year I think I said something along the lines of I

Speaker B: got it right here. Yeah.

Speaker E: All right, well I got it right here. What did I say?

Speaker B: You said we're going to see a huge AI face plant.

Speaker E: Yeah. And that. So to add some context that was all about um, some sort of failure that was at a colossal scale. That had some big financial impact because it really is the wild west out there. People are integrating third party apps without thinking too much. They're putting their data in places M without much authentication. I see this on a day to day basis like the governance is severely lacking. So I was, I would say much like you both, I think it was kind of mixed. There wasn't a single failure that I could find that you go point to. But there is some early data about data breaches that are AI related and uh, you know, all sorts of compliance, uh, related issues that just resulted in failures. So IBM did a study where they quantified the average cost of an AI related failure like a data breach specifically in the US and it totaled up to about $10.2 million per organization.

Speaker B: Okay.

Speaker E: And that's just all the, it could be legal fees, it could be you know, actual like hiring a team to go do the disaster mitigation failure, all that stuff.

Speaker B: Yeah.

Speaker E: And then if you look at the quantity of these sorts of breaches that are happening now, two years ago to, to a year ago, it was about like a 50% jump even in that year. And although the numbers are early, it's looking like a 40% jump in overall quantity. And those are in the hundreds per year. So you just multiply it out. I mean it's easy to justify that it's in the hundreds and millions. But I also think that intuitively we feel this because we've started to see a lot of more, a lot more sophisticated scamming phishing attempts. In our email we're getting robo calls that are like, I'm getting close to tricked now where it's just, it's crazy.

Speaker B: Yeah.

Speaker E: I mean I would even say I got an email a week or two ago from OpenAI saying there was a big leak with Mixpanel, uh, one of their third party vendors. They released a bunch of PII that they're having to deal with. I think it's still going to go up, but I would say that like the financial impact is probably on, but the effects were more distributed than just one big failure.

Speaker B: Yeah, I, I thought you were, I thought the prediction of a big AI face plant was probably a sure bet. Uh, I thought that would be certainly as more and more money and more and more, um, you know, time is spent on this. Um, a lot of companies just aren't ready and I thought that would happen. So I agree. I think you get about half credit on that.

Speaker E: All right.

Speaker B: All right. Uh, I want to transition now. I think I would If I was going to grade us, I'd say, I think zz, you probably got the closest to being on the mark there.

Speaker D: Okay. That means I'm not ambitious enough. I need to be more ambitious and

Speaker E: I'm going to break up.

Speaker B: Exactly. Okay. So let's talk about our predictions for next year. This is really tough. I mean, things are moving incredibly fast. We've seen. We've seen just changes kind of come out of the blue. So I'm really kind of curious to hear what you think if, you know, if you. Again, we're going to review this next year. So what do you think is going to happen this year that's going to make the news, that's going to have an impact, that's going to be meaningful? And let's just go in reverse order. Michael, you go first.

Speaker E: All, uh, right. Pressure's on. Let the record state.

Speaker B: Yes.

Speaker E: Um, so I'm going to frame this a little bit with a kind of a hot take. I love agents, especially when using them for things like development and a couple other, you know, pretty narrow use cases. Right. But I would say, by and large, they don't deserve the hype and credit that they're getting right now.

Speaker B: Agreed.

Speaker E: You know, right now you're getting a lot of people talking about, you know, like, two years ago it was rag. This past year it was agents. Now it's swarms of agents. I don't think the world's not ready for that.

Speaker B: We can barely control one. Let's go to swarm next.

Speaker E: Exactly. Yeah. Like, if you have an agent that's making decisions that are 90% accurate and you put 6 of those in sequence, it's a coin flip. And if it's 80% accurate, it's three. Three decisions in sequence, it's a coin flip.

Speaker B: I know, it makes me crazy.

Speaker E: So we're not ready for that, except for very narrow use cases. And the data are already kind of showing that humans are a lot more important in AI workflows than we originally thought.

Speaker B: Yes.

Speaker E: Pure automation is not. We can't just immediately leap there like most people want to. So Anthropic every year releases an economic index. Actually, I think they update it regularly. But right now, if you look at their usage statistics, about half of their API calls are related to human augmentation, like human assistants. And the other half are pure automation.

Speaker B: And just so I understand, you mean that on that first half, it's that the humans are interacting and changing or like entering the prompt or correcting output.

Speaker E: Yeah. Humans are involved in the workflow in some way that are checking at each step or sequence of steps to supervise. But I think that it's not going to move more toward automation. It's going to move more toward human augmentation next year. Okay, so right now it's. Yeah, it is just about 50, 50. And I would guess you ask most people, they're going to think we're just going to automate more and more. I think it's going to flip the other way.

Speaker B: So just to be clear, you're essentially, you're a little, um, uh, doubtful that agentic AI, it's not just that it won't be the main focus. You think there's actually going to be a little bit of a slowing or that it's just still several years out.

Speaker E: I think people are going to give up on the pure RPA approach. Or I say rpa. It's just an analogy, but like the pure automation approach. M Most M people I think this year have found out that building agents is fun, but getting value from them is hard.

Speaker D: Yes.

Speaker E: Myself included. I don't want to leave myself out of that. Um, but I think that the way that these agents are going to be used to build production systems, they're going to have humans in the workflow rather than pure automation at scale.

Speaker B: I get you. So you're essentially saying agents, people are still going to be embracing agents using agents, but they will, even before they begin the initiative, essentially say, oh, there will be humans in the loop at these different important checkpoints.

Speaker E: Yeah, they're going to augment supercharge human workflows that exist, rather than the classic example of replacing someone in a call center, which Klarna had issues with earlier this year.

Speaker D: Yeah, yeah, yeah, that makes sense. Human AI teaming.

Speaker E: Exactly.

Speaker C: Yeah.

Speaker E: Human machine teaming.

Speaker C: Yeah.

Speaker B: Yeah. Okay. All right, I think I agree with that one. Okay. All right, Emma, what's your prediction for 2026?

Speaker C: So mine's a little similar. It's kind of related to the impact on the workforce that there's been a lot of chatter about this year. You know, there's data that shows layoffs are happening less, you know, less job postings and things like that. And there's a lot. There's studies that prove it. There's studies that debunk it and say that it's mark, you know, broader market forces. It's kind of the downstream aftermath of COVID or things like that. But my prediction for next year is that there's going to be a big pressure on the C suite to start to show roi I feel like this year was kind of our reckoning with the, maybe the bottleneck or the challenge isn't necessarily in the technical capability, but it's the organization's ability to like reinvent their workflows and get value out of it in a meaningful way. And so I think that is sort of a, a realization that's come to light this year. And now the question is going to be, okay, we're investing in AI, where and how do we expect to see roi? And I think that's going to push companies in different directions. One is, you know, potentially the kind of cost cutting, um, direction, maybe showing that from a labor standpoint by doing layoffs, you know, debatably, kind of a more short term way to show this is how we're realizing value because we can do the same work with less people. I think the other companies are going to be more focused on augmentation, on really thinking about how can we um, enhance the work that we do, how can we change it, have new products and services, offer more differentiated competitive value. And I think we're going to see companies go in those two directions. Um, I also think we're going to start to see the impacts to the workforce play out a little bit more. My kind of takeaway right now is that the data is a little bit mixed to very confidently be able to say like the impact that we're seeing on the workforce is because of AI solely. But I do think the data around entry level rules is starting to show that. So from this year, the unemployment rate for 16 to 24 year olds, this was as of like July I think, is 10.8% and compared to the national average is 4.3% I believe. Yeah, so we're starting to see that impact and I think a year from now that's going to continue to play out. One of the reasons why I think this is such an important topic as well is there's this researcher, her name's Molly kinder, and she kind of focuses on economic inequality, AI and the future of work. So a very relevant intersection. And she calls this the um, great mismatch, which is basically if we look at sectors in the economy that are most subject to being impacted by AI, like jobs that are predicted to have the biggest displacement or impact, they have the lowest union density. So in other words, people in those roles only have like 1 to 4% representation in a union. So 95 plus percent of people are in this position where it's projected that AI is going to impact those jobs in some capacity. And they have no bargaining power. And so I think just this is going to be a really important dynamic that plays out next year. The ROI pressure, um, you know, kind of combined with seeing how impactful these tools are in work. And so I think the narrative that we've heard of AI is going to reshape work is kind of here, and we're going to see how that plays out next year.

Speaker B: Okay, so just to recap, you're essentially saying, all right, companies are putting hundreds of millions, billions of dollars into AI, and 2026 will be the year where, uh, you know, there's going to be pressure from shareholders, the board to say, well, show me the return on that. This is great that you embraced AI, but show me the return on that. And you think it's going to be a combination of, uh, maybe some workforce reduction. Right. Coupled with, uh, increased, uh, integration on the human AI teaming.

Speaker C: Yeah.

Speaker B: Is that a fair summary?

Speaker C: Yeah, absolutely.

Speaker B: Okay. All right. I think I would bet pretty highly on that. I think there's going to be a lot of pressure. We see this all the time. Time. You know, we try to really guide our clients towards, uh, using AI in a. In a way that is impacting their business meaningfully. Like, have real roi. Just don't. Don't do AI. Just to say you did AI and check a box. And I think most of the time we're successful at that. And I think that the companies that are just doing it to say they do it, the. The reckoning is coming. I totally agree.

Speaker D: Yeah, that's a good point.

Speaker B: That's great. All right, Zizi, what is your big prediction for 2026?

Speaker D: Um, I'm going to throw out a crazy one. I think we're going to have a, uh, GPU data center in space.

Speaker E: Oh, my God, I love it. Zz.

Speaker B: I did not see that coming at all. Man.

Speaker E: That's amazing.

Speaker D: I just, uh, did some research recently and found out that people have already launched a GPU satellite. It's a trial, it's kind of. They're trying to collect, uh, stats about how GPU operates in space. But, I mean, this past year we have seen so much data center built out. And if you want to look at large GPU data centers being built, look no further than in Texas. Yeah. And for, uh, GPU data center in space and Center PTAI is talking about it. Um, I think, uh, you know, um, Elon M. Musk said it's interesting. And, you know, there. There are a few other big tech companies that are, uh, really, uh, Interested in this space. So my bet is maybe something will happen in this space and we will have a.

Speaker B: Are these solar panels? I mean, solar panel driven?

Speaker D: Yeah, that's. That's what I heard too.

Speaker E: That kind of shocks me because I think. I mean, an average GPU is like a few hundred watts or more or something. If you have a bunch of, um. That's a lot diffuse. I almost would think nuclear reactor on or something.

Speaker B: Yeah. Like. And what about heat dispersion? That's tough in space.

Speaker D: Yeah. But a lot of smart people are working on it, so figure it out.

Speaker B: Yeah.

Speaker D: So let's see. Yeah. On. On a less crazy note, I feel like, you know, the agent coding system, they're. They're good, but they're so, so forgetful.

Speaker B: Yes.

Speaker D: Like every day they, you know, they don't know what they did yesterday and it doesn't know about me, you know, so I'm. And I'm betting on something happening in a memory space. So a good memory system that's much better than RAG and other existing systems that can make this agent really work while you sleep so that next morning you can just see some good work.

Speaker B: Oh, yeah, yeah, yeah. And you're not starting over from scratch every time. Okay. So I'm really surprised nobody has a prediction about some breakthrough in AI from like a modeling or an intelligence perspective. Is that a sign that we think we may be kind of leveling off on the. On the progress?

Speaker D: That's a great question. That's a great question.

Speaker E: It does feel like it's, uh. Okay, so I'm going to agree with Andrej Karpathy on this point. I mean, I think, uh, actually this is maybe like mincing his words a little bit, but in that interview with Dorkesh, uh, he was talking about how the impact from AI is much smoother from. From how people think in terms of economic impact.

Speaker B: Yes.

Speaker E: And I think that's going to continue to be the case. And because we're waking up to this reality that, you know, we need to get value from these things, it makes sense to me that at least the focus is going to shift a little bit more toward execution than pure R and D. There, of course, are breakthroughs on the table that could be made. But. But in terms of just like, emphasis, it feels like extracting value from what we already have is going to be more important this year than.

Speaker B: That's really in line with what you were saying, Emma. Um, I definitely feel that. I think there's this sense that a lot of money's been Spent. Let's see the return on it. The other thing that I really believe this deeply, which is every now and then you'll see posts about all this money's going to AI and where's the return on it? And it's perhaps quote overhyped. I believe that if we saw no improvement in AI from the modeling capability, context, windows, any of that, we have barely begun to be able to absorb and integrate these capabilities into our lives and our businesses in a meaningful way. And I think it just happens so quickly that it's going to take a while for us to learn. It's literally a human thing. We've got to learn how to leverage these technologies and use them effectively. And there was this, I think false illusion, um, that people were using ChatGPT and it was so effective that they could just throw any problem to an LLM and it was going to work. And the reality is that these systems are really jagged. Um, I don't want to say brittle, but they're probabilistic and it's complex and it's really hard to build these things in a way that they are entirely reliable. And so I'm incredibly optimistic. I actually do think we're going to see some breakthroughs this year on the modeling side that will be meaningful. Um, I think we're going to continue to see uh, uh, a lot of investment in the reasoning approach. And I'm still a big believer in sort of uh, reinforcement learning with verifiable rewards. I think if we can push that domain a little bit further and expand that beyond math and coding and I think there are ways we can do that that will uh, uh, see continued improvement on that sort of top intelligence level.

Speaker A: Mhm.

Speaker E: It is interesting to see how the original pre training historically has been the big biggest cost driver for training a big model. And it still is. But the RL component of that cost seems to be just.

Speaker B: That's right, all this, you know, uh, post training is becoming more and more uh, where the money goes.

Speaker C: So there's this quote that I find myself always circling back to and I, it's from a biologist, I can't remember his name right now but it's that humans biggest challenge is that we have godlike technology, medieval institutions and Paleolithic emotions and it just always feels relevant. But like the way that I view it is the frontier labs. Like when we talk about how much investment is going into AI, it's not like all that money is going into enterprise AI, it's going into sort of this race towards AGI or, you know, whatever we want to call it. And then that's trickling down to enterprises where we're saying, okay, these breakthrough capabilities, how can we actually apply them in a meaningful way within our medieval institutions? And so I think there's, you know, is there as, um, companies still relentlessly pursue this godlike technology? I totally think there's going to be development and advancements in 2026, but I also think there's a lot more focus within companies and within the economy now of, like, all right, we need to actually apply this in a meaningful way.

Speaker D: Like, how do we orient ourselves to use this technology better?

Speaker C: Totally.

Speaker E: We still haven't even refactored to incorporate innovation. That happened two years ago, I would say. Yeah, like, we're refactoring. It feels like we're.

Speaker B: Yeah, okay. I couldn't agree more. And what a great note to end on. Thank you all. Um, we'll come back here a year from now. We'll see how your predictions go.

Speaker E: Fingers crossed.

Speaker A: Thank you for listening to Hidden Layers. This series is hosted by Kung Fu AI, a management consulting and engineering firm focused exclusively on artificial intelligence. If you have any questions or thoughts about today's episode, or if you know someone we should feature, please visit us at Kung Fu AI.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Eric Ries on Why Good Companies Go BadPodcast Archives · on Anthropic92 / 100
  • The 18x Midas Lister Betting $3B on AI (and calling most of it fake) | Navin Chaddha, MayfieldThe Peel with Turner Novak · on OpenAI91 / 100
  • 183: Why Trusted Data is the New AI Moat (w/ Rick Kranz @ AI Marketing Automation Lab)Move The Needle · on Anthropic91 / 100
  • Is Your Business Invisible to AI Search? (And How to Fix It) ft. Ray YoungRevenue Science · on Claude85 / 100
  • Welcome to the Software Renaissance. Your Strategy Isn't Ready | Martin ErikssonProductized Podcast · on ChatGPT82 / 100
  • Miles Rowland: Why Every Portfolio Company Needs an AI Engineering TeamAI Pathfinder for Private Equity Podcast · on Claude81 / 100

More from Hidden Layers

All episodes →
  • AI Is Designing the Next Cancer Fighter | EP.5383 / 100
  • Anthropic Code Leak: A Rare Look Inside Frontier AI | EP.5282 / 100
  • The "AI Bubble" Bubble | EP.5174 / 100
  • Did AI Kill Programming? | EP. 5072 / 100
  • Your AI Is Too Big, Too Expensive, and Probably Wrong | EP. 4978 / 100
Explore the best B2B HR podcasts →
All Hidden Layers episodes →