The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/The neXt Curve reThink Podcast
The neXt Curve reThink Podcast artwork

The Ultimate NVIDIA GTC 2026 Recap (with Karl Freund and Jim McGregor)

The neXt Curve reThink Podcast · 2026-03-22 · 45 min

0:00--:--

Key moments - from our scoring

Substance score

64 / 100

Five dimensions, 20 points each

Insight Density13 / 20
Originality12 / 20
Guest Caliber14 / 20
Specificity & Evidence14 / 20
Conversational Craft11 / 20

NVIDIA's GTC 2026 marked a significant strategic shift from GPU-centric acceleration toward a heterogeneous computing platform. Freund and McGregor discuss how the company acquired Groq's LPU technology and leadership to address token generation and agentic AI workloads - a departure from previous focus on large-context prefill problems. The conference revealed a multi-rack architecture combining Vera (CPU), Ruben (GPU), Rock LPU, and SDX storage solutions, representing a fundamental move beyond single-processor optimization toward data-center-as-a-unit-of-compute thinking. Key software announcements included Nemo Claude (NVIDIA's security wrapper around Anthropic's Claude), Lang Chain integration, and over 1,000 CUDA-X libraries for domain-specific applications. The memory hierarchy evolved with SOCAT becoming an industry standard and GPU-attached storage now serving agentic workloads that require persistent access for weeks or months. Speakers emphasize that Open Claude (or Anthropic's Claude itself) represents the killer app for agentic AI, enabling rapid custom agent development on-device or in cloud. The conference also revealed concerns about circular investment in AI and a shift toward financial sector attendees over developers, suggesting potential market saturation questions.

Key takeaways

  • →NVIDIA acquired Groq's LPU technology and shifted strategy to building heterogeneous rack systems (Vera CPU, Ruben GPU, Rock LPU, SDX storage) optimized for different AI workloads rather than relying solely on GPUs for all problems.
  • →Nemo Claude and Anthropic's Claude represent the killer app for agentic AI, enabling custom agent development with Lang Chain while Nemo adds security wrapping to address Claude's tendency to drift from rule constraints.
  • →Memory hierarchy expanded with SOCAT becoming an industry standard and SDX storage providing persistent, high-speed access for multi-week agentic tasks, shifting from GPU-optimized to AI-optimized storage architecture.
  • →Jensen Huang's roadmap philosophy shifted from predictable four-year GPU roadmaps to adaptive, heterogeneous platforms responding to rapidly changing workload diversity, reflecting a move toward application-specific and function-specific optimization.
  • →NVIDIA is architecting data centers as single units of compute with entire functional blocks as racks, enabling simulation of physical systems (genomics, Earth modeling) at scales previously impossible.

Guests

Karl FreundJim McGregor

Topics in this episode

Anthropic ClaudeNVIDIA Groq LPU acquisitionNemo ClaudeVera CPU rackRuben GPURock LPUSDX storage rackSOCAT memoryLang ChainCUDA-X libraries

Questions this episode answers

What is Nemo Claude and why did NVIDIA create it?

Nemo Claude is NVIDIA's security wrapper around Anthropic's Claude, designed to prevent the agent from drifting away from its assigned rules due to memory limitations or other factors - addressing a key vulnerability in using Claude for agentic AI applications.

How is NVIDIA's architecture changing beyond GPUs at GTC 2026?

NVIDIA introduced a heterogeneous platform combining Vera CPU racks, Ruben GPU racks, Rock LPU racks (for token generation), and SDX storage racks, marking a shift from GPU-only solutions to multi-processor architectures optimized for different workload types.

Why does agentic AI require different storage architecture than chat-based AI?

Agentic AI agents must run continuously in the background for days, weeks, or months to complete tasks, requiring persistent, high-speed storage near accelerators; chat-based interactions only need storage for milliseconds, so SDX storage was designed specifically for this persistent access pattern.

What happened with SOCAT memory at GTC 2026?

SOCAT, the proprietary memory architecture developed by Micron and NVIDIA to address KV cache issues, became an industry standard within a year and is now being adopted by other AI accelerator makers and even HPC applications seeking high-bandwidth configurations.

Why does Jim McGregor call Claude the killer app for agentic AI?

Claude enables rapid custom agent development on-device or in cloud, with Lang Chain tooling making it easy to build agents that orchestrate multiple LLMs and tools - addressing a five-year search for a scalable AI application monetization model beyond chat interfaces.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

13 / 20

The episode packs substantive technical content about NVIDIA's strategic pivot toward agent AI, the LPU acquisition, memory hierarchy evolution, and heterogeneous infrastructure design. However, significant portions are consumed by congratulations, housekeeping, and conversational padding that dilutes insight density. The core insights (agent AI as killer app, GPU+CPU+LPU rack strategy, finops unpredictability) are valuable but not uniformly dense throughout.

Open Claw, I believe is the killer app of agentic AI
you're gonna have a Vera, you're gonna have Rubens, you're gonna have STXs. You're gonna have LP Xs. Yeah. So you're gonna have, and eventually you're gonna have cpx.

Originality

12 / 20

The guests offer some genuinely fresh observations - particularly the framing of NVIDIA's shift as addressing agent-specific bottlenecks rather than abandoning GPUs, the emphasis on finops/token budget unpredictability as a blocking issue, and the critique that telcos don't have urgency. However, much of the framing recycles well-known NVIDIA talking points (heterogeneous compute, Cambrian explosion, speed-of-light philosophy). The 'killer app' claim about Open Claude is intuitive rather than contrarian.

the business model for the telcos has to change, and we've been telling 'em this for over a decade now
Nvidia is dramatically changing. There's obviously a huge reaction to the diversity in the com, the complexity of the, inference opportunity

Guest Caliber

14 / 20

Karl Freund (Cambrian AI Research) and Jim McGregor (Curious Research) are credible, long-standing semiconductor and AI infrastructure analysts with demonstrated industry access. They attended GTC in analyst capacity and bring technical depth, particularly on memory architecture and agent economics. However, neither are operators who have built or scaled AI infrastructure companies themselves - they are expert commentators rather than practitioners with P&L responsibility for major AI initiatives.

I'm joined by. Paul Fre, who is the center of the Cambrian AI event that's happening right now. And he has a company called, or a firm called, Cambrian AI Research.
Jim McGregor. Of the fame curious research.

Specificity & Evidence

14 / 20

The episode provides concrete technical specifics: Raybolt 3 and Raybolt 3 LPU, NVL 72 racks, Vera Rack, CPX, Spectrum X switches, SK Hynix Cambium memory, SDX storage solution, and agent token budget allocation. However, it lacks hard numerical data on performance deltas, cost comparisons, customer deployments, or timeline commitments. Claims about 10x-100x efficiency gains and 1000x improvements per year are asserted but not substantiated with data. Finops economics remain vague despite being identified as critical.

you're gonna have a Vera, you're gonna have Rubens, you're gonna have STXs. You're gonna have LP Xs
we're seeing about a 10 x improvement in efficiency. Just in terms of the model sizes. So the model sizes are coming down 10 x At the same time, the performance of these systems, or performance efficiency of these systems is going up. 10 or a hundred x.

Conversational Craft

11 / 20

The host Leonard asks directional questions and allows guests to elaborate, and there are moments of genuine disagreement (depreciation vs. opex on older hardware, obsolescence definitions). However, follow-ups are often soft; claims like 'Open Claude is dangerous' and 'nemo-claw proves enterprise grade' are stated without sharp pushback on evidence. The host lets interesting threads drop (e.g., telecom urgency, CPO pull dynamics) without drilling deeper. The tone is collegial but lacks the intellectual friction that would force precision.

Well, you know, there's a lot of big numbers being put out there, but, you know, I, I know that you guys. You guys are very bullish about all this stuff. There's no butt. Yeah. No, there is a big butt.
I don't think they truly see the opportunity and they need to, I really think that the business model for the telcos has to change

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

nvidia17storage17fact16open15claw14last14seeing14solution13network13problem12three12cost12racks11question11industry10whole10

Episode notes

Send us Fan Mail Leonard, Karl, and Jim attended the AI event of the year, NVIDIA’s GTC 2026. This year, the agentic AI theme went into overdrive with its DeepSeek moment, OpenClaw. Or was it a ChatGPT moment? The trio unpacks a deluge of announcements from NVIDIA’s technology stack and across the domain-specific industry and computing solutions that are meeting the agentic AI moment. In this episode, Leonard, Karl and Jim talk about the top headlines and share their analysis from NVIDIA GTC 2026. ️ Staying ahead of the fast-rising tide of AI ️ Agentic AI's ChatGPT moment with OpenClaw and NVIDIA's NemoClaw ️ NVIDIA pivots its AI data center inference roadmap with Groq 3 LPX for agentic AI ️ NVIDIA is innovating at and competing with the speed of light ️ NVIDIA is shifting away from GPU-centricity with Groq 3 LPX & Vera CPU systems ️ NVIDIA is becoming more than a GPU company. It is heterogeneous! ️ NVIDIA claims that it will become a $trillion company by end of 2027! Will it? ️ Is NVIDIA GTC becoming more of a VC/Wall Street event than one for developers?

Full transcript

45 min

Transcribed and scored by The B2B Podcast Index.

Next curve. Everybody. Welcome to, uh, this episode of Next Curve's Rethink podcast, where we break down the latest tech industry vets and happenings in the world. Semiconductors and Carl's favorite topic.

Ai. AI and gyms, right? Yeah, absolutely. and we break all this stuff down into the insights that matter.

I'm Leonard, the executive analyst at Next Curve, and I'm joined by. Paul Fre, who is the center of the Cambrian AI event that's happening right now. And he has a company called, or a firm called, Cambrian AI Research. And, we also have the dynamically ated, Jim McGregor.

Of the fame curious research. How do you like that one, Jim? I learned that at GTC 2026. Oh, geez.

I, I have no idea what it really means, but it is a thing. It's called dynamic ation. I'll have to look that up. Yeah.

You go to Jim should thank you, or Hitchy. Yeah, it's really interesting com collaboration between Apple and Nvidia in the whole spatial computing space. So, really exciting stuff, but, I always associate Jim with really exciting stuff. So gentlemen, welcome.

Welcome. And so you're probably wondering what we're gonna be talking about. Well, I already gave you a hint, GTC 2026. We're gonna do a recap and I'm gonna have the pleasure of listening to these two gentlemen who were on the analyst program I was not this year.

and, I'm really curious, what your takes were since you were there, front and center with the program. And I was out in the exhibition hall running around. Trying to find a free lunch. But before we get started, remember, please, like, share, react, and comment on episode.

Also subscribe here on YouTube and Buzz Sprout. Or listen to us on your favorite podcast platform. Opinions and statements by guests are their own and do not reflect those of next curve, or myself. And, we're doing this for informational purposes only to provide a open forum for discussion, debate on all things, AI and, Silicon Semiconductor and supercomputing.

So, by the way, hey, look what happened. One of these hundred thousand subscribers and I wanted to, thank all of our subscribers, but also, Jim and Carl for their wonderful collaboration on the Silicon Features series. gentlemen, I attribute a good portion of this to some of the great. Insights and collaborations that we've done over the years on this program.

So thank you so much and congratulations. 'cause part of this. What are the bad chain? Congratulations.

Congratulations to you guys as well. That's pretty cool. That's pretty cool. Well, you notice he didn't send us one?

No. Just kidding. No, I didn't. I'll go check my mailbox.

Yeah, I'll you guys one, but, no. What I want everyone to do though is, recognize, you and, also reach out to, Jim and Carl for their insights. It's ab, they're absolutely valuable voices and, eyes and ears on the industry and what's happening in AI and the semiconductor industry broadly or even beyond. So yeah, definitely reach out for to them and, stay ahead of this, this, you.

Very quickly rising tide of ai, that everything, it's not possible. Everything. It's not possible. You can't stay ahead of it.

It's impossible. No, no Open claw. Just prove that. Tell yourself short.

I'm trying to make you look good, Jim. Open Claw, what's your problem? I was telling Jim, I've been working here on the new house in Tucson and I totally miss the importance of open claw and, I got to GTC and I'm like, holy cow, the world has changed While I was. Changing light bulbs here.

And he's looking at me. He is like, Jim, why didn't you tell me about Apple? Why did you call me and tell me? I thought you knew.

So, Jim, takeoff for a few weeks in the world changes. Yeah, yeah, yeah. Let's get, I mean, you brought up, let's go fucking Claude. So, I mean, what stood out this year?

That that was, that was probably one big event. The biggest things that stood out was open claw or Nemo claw, I should say, which is, NVIDIA's wrapper, kind of security wrapper along with their, open shell, the entire. Nemo Claw solution is designed to put a wrapper around open claw. And Open Claw is, very innovative in the fact that it's basically an agent to call other, LLMs or to call other tools and everything else.

It's the easiest way we've had so far to build custom agents personally and actually on device. So it's a major innovation. Unfortunately, and we've played with it, it has its limitations. it can go rogue, because you don't really out there, the code when you give it rules.

And those rules can be forgotten over time, especially if you're running into memory limitations. So it has some serious security limitations. Well, yeah. Uh, Nvidia recognized that it came out with Nemo Claw to put the security wrapper around it.

Very innovative, along with tools like Lang Chain, the fact that you can build these agents so quickly, so effectively, in a very short period of time to do pretty much anything you want. And they can operate on device or in the cloud. Yeah. And maybe a hyperbole of it.

It may be an overstatement, probably is, but to me, this is the killer app. Yeah. We've been waiting for a killer app for AI for five, six years, and we thought it was chat GPT, but it's hard for a lot of people to make money off chat. GPT, open Claw, I believe is the kill the killer app of Agen ai and it's interesting, I can't imagine what it must have been like to work at NVIDIA for the last three months because three months ago, Nvidia did the unthinkable.

They decide that GPUs weren't the right solution for something, and that something is called Gentech ai. They wanted to get on top of that. They previously were working on the front end of inference processing the large context, prefill problem. and they came out with a product, announced a product for that called, CPX Ruben.

Yeah. And then as Ian Buck said, along the way, we found something better. Something different. It's not solving the large context prefill problem it's solving at the other end of ai, which is token generation at massive incredible speed and very low latency, and importantly, deterministic latency offered by GR language processing units, LPs in the last three months, not only have they integrated these new engineers and leadership into their team.

They've come, they've taken their next generation chip called Rock three three and Rock three LPU three, and so they built an entire. NVL 72 kind of rack for them and them as they now. So, they have now shifted strategy to focus on agen AI as opposed to what they were focused on and still are, which is the large context problem that with, let's say, using movies video as a input token. Okay.

Or body of code, a million lines of code, using that as a near input, input, sequence. They're now focused on genetic ai and they're, they're gonna have probably anywhere from one to four racks of Brock Lp Xs. alongside of Ruben, right? So you're gonna have a Vera Ruben Rack, and you're gonna have like anywhere from zero to four LP X racks for each Ruben Rack.

And it's like, wow, this is a major shift. This company has been so focused on GPUs. You got a problem, we'll solve it with GPUs. Right?

Yeah. no, this company was amazingly, Agile to be able to say and fully believe, no, this is not the right problem. GPU is not the right solution for this problem. Rock has the right solution for this problem.

It would take too long to acquire the company, so we'll just acquire its assets and leadership team and engineer. Yeah. so well, it is pretty damn amazing. On top of that, they were coming out with a storage rack solution.

The SDX? Yeah. Yeah. And even re-architecting future generations of racks with Chiver to where they're actually gonna have blades so they can double the density of GPUs from 72 to 144 and have, and get away from using cables and actually using back plane solutions.

So, I mean, the level of innovation is just phenomenal that we see going on at Nvidia. This, Jensen likes to say that Nvidia. Innovates at the speed of light, and it definitely looks that way and that speed of light. I haven't heard them talk about speed of light much lately.

When it was first discussed, it was more than just speed. You're absolutely right, Jim. They're moving an amazing speed. But it is the idea that your competition is not another company.

Your competition is fundamental laws of physics, like speed of light. Yes. So speed of light is perfection. How close can you get to perfection given the technology you have today, the resources you have today?

how good can you get, forget the competition. And video guys don't think of competition. They think of the speed of light. Yeah.

Well that's, yeah, that's, a huge shift. And I know that during the keynote. he was mentioning Cuda quite a bit. In fact, he, well, it's, it's 20 years, kind of preface the whole thing.

20 years, right? 20 years of Cuda. Right. But like, what's really interesting about the comment that you just made, or the assessment that you've made, Carl is this huge shift away from, the GPU.

yeah, I think that was really noticeable. But also, it wasn't just the LPU, they introduced the Vera Rack, so now, yeah, They're making a big play into CPU, and so all of a sudden we're seeing this big change in the quote unquote accelerated computing story where I think two, three years ago, the impression was everyone's gonna move everything toward these quote unquote GPU infrastructures right now, it looks like Cuda might not have that kind of scope. I, I don't think they're moving away from GPUs.

I think they're adding things beyond the GPU to the GPU based solution. So if you look at infra from Nvidia a year from now mm-hmm. You'll have racks of GPUs, all with all their fancy NV link and everything. You're gonna have racks of CPUs, which is Vera, and you're gonna have racks of LP to solve The ve, the agent AI problem of lots and lots and lots of small models.

Mm-hmm. Uh, all calling each other. And so it's a much more heterogeneous environment. I don't think it's more away from GPU, I think it's move beyond GPU.

And I think if you go a step That even that's a huge statement right there. 'cause in a lot of people's minds, Nvidia, a GP company. Yeah. And I put it quotes because not anymore.

We're not just talking about the. And I think if you take it even a step further, they continue to work, look at the workloads, which are changing rapidly and very, very quickly, and they're trying to identify the bottle X. So, I mean, last year they, along with Micron, introduced silk cam memory just to address, the KV cash issue and everything else. Yeah, they introduced a storage solution this year to also address some of those, the storage specific solution or needs that you need for ai, not long term storage, but AI type of storage that, short term or midterm type of solution.

And they keep looking at that and it goes beyond just even the hard. I mean, the fact that, uh, they continue to look at, enhancing foundational or frontier models for specific types of applications, especially for the physical world and the fact that they now have over a thousand Cuda X libraries for modeling pretty much everything from, DNA genomics all the way to the planet Earth. I mean, basically really looking at is it hardware? Is it software?

Is it interconnect? What is it? What are the bottlenecks and how do we get over that? And we're even seeing that with next generation platforms, with Kyber going to those blades and everything else, it's, it is amazing that at one company has this breath to look at what has to be done to address these needs.

For me, the sparkling moment of GTC when Jensen went home to, to, to, and and told his wife, Hey, today I announced we're gonna do a trillion dollars in revenue and. No. So what? No big deal.

Take out the trash. Now, Jensen, I've been telling you for days, you gotta take out the trash. That was his wife's response supposedly. And if you look at what the stock market's done, it's done the same thing.

Trillion dollar, eh? So what? Yeah. Well, you know, there's a lot of big numbers being put out there, but, you know, I, I know that you guys.

You guys are very bullish about all this stuff. There's no butt. Yeah. No, there is a big butt.

There's there's a big butt. there's a big, it's money with two's either, actually. There might be, anyway, it might be one of those too. I, I'm definitely one of the overhanging, clouds at this conference is, circular investment.

I, yeah, I've. Have a lot of discussions about that, so we can't discount that. Even though from a technology front, there's a lot of excitement. But from an ecosystem standpoint, one of the things that I don't think really, was well received as a comment that Wall Street is their, there's more Wall Street people or investment or finance sector people there than developers.

the conference itself has changed a lot. We saw that with, the event last year. Uh, I don't know if they had a Q day this year. No, they did not have a quantum day this year.

Quantum. Okay. Thank God. I thought that was terrible.

And then CES also, less focused on technology, more on the markets, but, I, I think the. Important thing here though, is when we look at the transformation of that, the company as we're alluding to through the comments that have been made so far, Nvidia is dramatically changing. There's obviously a huge reaction to the diversity in the com, the complexity of the, inference opportunity, and this is stuff that we've actually been talking about. Quite extensively in prior episodes, whether it's related to the memory architecture and or, the role of application specific, silicon and, presenting a different envelope of optimization for some of these in, these models in runtime.

You guys mentioned adaptable. The identity of a company is evolving really rapidly. Last year it was all about, hey, we're a AI infrastructure company, right? This year, AI Factory seemed to be that identity that, Nvidia wanted to portray or at least push out there?

I don't know. What do you guys think? I don't think it really changed. I think they still mention, they may not have mentioned AI infrastructure as many times as they did last year, but I still think that's the focus, they want, and more than anything else, Carl can speak to this, but we're seeing more of a message around not just hardware, but obviously software and tools and everything else.

Yeah. And support for the entire ecosystem. Like the S stx. Yeah.

They're not coming out with the STX, that's going to be up to partners to actually develop that rack. and the fact that all these models and all these libraries are now open source. Yeah. Yeah.

And speaking of the SDX, they're pretty stringent on the requirements and their adherence to the reference designs. But, one of the things that I think is interesting is last year it was about GPU optimized storage, right? Mm-hmm. It seems like, Jim, what you're saying is this year's, let's say.

The evolution of that idea is AI optimized, right? Yes. Because when we're looking at AgTech, the workloads are more expansive. They're not singularly focused on just the AI element.

that the key problem they were solving is that an agent's working on the job you gave, it has got to work continually in the background. Yeah. For days, weeks, or months. Yeah.

And so you need all of its storage that agent's. Cadre of, of other agents. They all need very fast storage to be able to continue to work on that problem until completion. Yeah.

when we talk about ai, we think in terms of asking chat, GPT something, right? That's a one shot deal. And once it's done. So the storage required for that only exists for, a few milliseconds.

Yeah. Whereas an age, it may need to have its storage close by for weeks or months. Yeah. And there wasn't a good solution for that until NVIDIA came out with the STX.

So that's, it's a really different animal. Again, everything's being driven right now from magenta ai. I can't say that often enough. Well, we, last year we made that observation around tiered memory, and then the very bottom of that, that tiering was the storage bit, right?

Yeah. The, G-G-G-P-U optimized storage for long, long memory, right? Mm-hmm. And so, yeah, I mean, Jim, since you're like the super technical guy, what did you see changing there in terms of the thinking around the memory tiering because.

That was such a big deal for the course of last year. Well, we came out, they came out with so camm and it was a proprietary, basically architecture last year developed between Micron and Nvidia. And within a year it became an industry standard. And we're seeing not only interest in it from anybody else that's doing AI accelerated platforms, but even from like HPC applications, they don't want necessarily the height.

Density, but they want the, they want the high bandwidth. So they're looking at it from, for lower density configurations. So just the fact that, that's become a new, new, level of, hierarchy within the memory hierarchy. To support KV Cash and now, we were already talking last year about GPU attached storage.

Well, now essentially what we have with, STX is accelerator attached storage, and I hate to say GPU because now we have the LPU and everything else, but you have that storage solutions. It's really dedicated to those AI accelerators to be able to handle and. Perform those solutions in addition to that. Yeah, that changing memory stack that we're developing.

Yeah. And at the risk of self-promotion of the idea of Cambrian explosion, we're seeing a camian explosion of the heterogeneity. Yes. Yeah.

Of AI solutions ev. Not one solution, one size will not fit all. No. And I think Jensen realized that, and really that's what drove him to expand his portfolio to CPU Racks and LPLP racks, along with the GPU Rack and storage racks and network.

it's a remarkable shift that I don't think the market fully appreciates right now. Most people don't appreciate that right now. And then Leonard, kudos to you and your platform for helping other people. More people understand what's really going on here.

It's a major ground shift. It's you guys as well. And the fact that we had these debates instead of just agreeing with each other. I think that surfaces those salient points that, hopefully, shapes a better dialogue around all this stuff, but that helps people, maybe think,, in more lockstep with what's going on.

I mean, so thank both of you guys, and again, congratulations. You guys are all part of this. Achievement here, with the YouTube channel. but no, you know what?

I, here's the thing though. I don't, I really don't think, Jensen knew what was. coming down the pike. I know a lot of people love to think that, but here's the thing.

Let's go back to the roadmaps that he's laid out over the years. Mm-hmm. This stuff was not part of the roadmaps and he was very vocal about saying that. We give you assurance that, over a four year period, you guys can depend on us.

To have that, reliable roadmap that can guide your investments. What we're looking at now it's turning into a much more complicated, outline even for AI supercomputing, and I think, You guys made the comment about adaptability. Yes. They're adapting really quickly.

And one of the things I'm noticing, even with some of the hyperscaler plays, whether it's Maya 300 or, Ironwood or Google's whole line of TPU systems or it's, train. Everyone is trying to. Adapt and we're seeing roadmaps as well as platforms that are transitional. Right?

And the way I looked at, the Grop three LPX system, it is a weird add-on to a Vera Reen and you mentioned Carl, you're gonna have a Vera Ruben, and then maybe one to four. LP X Racks Actually, we should specify. You're gonna have a Vera, you're gonna have Rubens, you're gonna have STXs. You're gonna have Lp Xs.

Yeah. So you're gonna have, and eventually you're gonna have cpx. So yes, and CPX will be a part of that as well for some customer. This is a lot.

This is a lot for, Customers and the industry to digest. I mean, it, it we're going away from what looked like for three years ago, a very simple looking and extensible roadmap to something that, to your point, Carl has become very heterogeneous and mm-hmm. Really a diverse stack to meet, a wide range of. Domain specific requirements or even function specific requirements.

And that was something that, that really resonated with me in Jensen's talk is that he actually gave to application specific stuff. Right? Yeah. at least that was my.

Take, I don't know how, what, how you guys, well, I, I think it fits in with that whole view that, the data, the new unit of compute is the data center. Yeah. And the fact that, we would've thought about this as a single server. How do we architect a server?

Yeah. 15, 20 years ago, how much memory are we gonna put into what kind of storage? What, what type of processor do we put an accelerator in? Now we're at the point where.

these functional blocks are entire racks, and we're thinking about the data center as one unit of compute, and really how do we build this system? What pieces do we need to put in there? And let's face it, if we went back 30 or 40 years, it was, what components do we put together? And then it went into an SSC.

Yeah. And then it was, so, we continue, building this, I just think that it's gotten to a point where we're, yeah. And you have to remember that we're now addressing problems. When you think about simulating the human genome, when you think about simulating the earth, when you think about simulating anything physical.

Yeah. We're now reaching into areas of scientific research and analysis that we only dreamed of, previously. So, we're building systems. Bigger and better than we ever have.

Just look at what we're doing with, AI and multimedia and the fact that we now really have realistic AI generation capabilities for entertainment and gaming and stuff like this. It is amazing what we can now do. the roadmap just keeps taking us to where we're gonna get. We're gonna be able to do more, and we're gonna be able to do more with less.

Because, from what we've seen, every year we're seeing about a 10 x improvement in efficiency. Just in terms of the model sizes. So the model sizes are coming down 10 x At the same time, the performance of these systems, or performance efficiency of these systems is going up. 10 or a hundred x.

So we're seeing a thousand X improvements in capabilities and efficiency almost every year. almost every year. It's mind boggling. I had a, yeah, I had an interesting conversation actually on my other podcast study, iot Coffee.

Talk about that performance versus scaling. And this is where things like times X whatever, a hundred times or billion times there's some charts that gentle was showing. I think one of the things that, that oftentimes gets conflated is the difference between, the scaling versus performance with performance. It's really about.

Do you have more efficient compute and are you able to then, have that more efficient compute in a system that can. maximize the density or push the envelope with density, right? Whether it's, innovations on the thermal front or the power or compute, right? And then you have scaling, which is the scale out.

It has really been pushed forward with scale out and now scale across, right? And so some of these times a million or larger. Multiples that we're seeing are typically on the scaling of the size of the clusters that are now spanning a data center campuses. Right?

And so I think it's really important to have that, proper lens as we're ingesting these large multiples that people are throwing out there. I think of it in the same form factor. And I'll be honest with you, I do see, anywhere from a hundred to a thousand x improvement per in the form factor every year. I was at Embedded World in MWC.

And I'm talking to the network guys and saying, listen, we need to make the network part of the platform, part of the workflow, because in many cases, like physical AI slash robotics, it has to be the intelligence. It has to be the backbone of the solution because A robot has a form factor, it has a power budget, it has all these things and a performance limitation. You're gonna want a lot of that intelligence in the network to do the sensor fusion Between all these different hundreds or thousands of sensors between robots and other sensors, you're gonna want it to be able to do the orchestration between all these different, units.

You're gonna want it to do LLMs that you can't do on device, and you're gonna have to be able to do that in sub one millisecond. So the network, whether it's a router, whether it's a ran, whether it's whatever, has to be able to do that. And it has to be able to handle a thousand x increase in data. yeah, some of 'em looked at me and said, oh no, we're gonna do all that in the cloud.

I'm like, no, we, you're not. And some of 'em looked at me and said, well, we'll get, there we're six G. I'm like, no, you need it. Tomorrow.

I don't think they get the urgency and the need and the fact that, we really are, you talk about, yeah, we are seeing some of these astronomical numbers thrown out there for a million X and everything else with these data centers. But quite honestly, when you look at the amount of data that's being processed, it is literally increasing a hundred to a thousand X every year. And it has to be done in the same form factor, and so that's the volumes. But if there's a suggestion that generation over generation, we're seeing these large multiples in terms of performance.

The other topic that was, debated pretty heavily, at least out, I was doing my own thing, so the, I wasn't part of the analyst program. I don't know how much this was talked about, but the whole obsolescence topic, right. when you look at something like that in terms of a generational improvement. It really makes the whole, Hey, you can extend a useful life of any investment, 10 years really like untenable.

Because at the end of the day, it's not about how do you repurpose, you can repurpose things. People do that all the time. It doesn't mean that it, let's say that you get a hundred times improvement in on the opex and with the energy efficiency and producing a token. That alone makes the prior generation obsolete because it's a hundred times more.

But that's not what's happening. The price, the cloud price of AMP peer. Mm-hmm. Remember AMP peer?

Yeah. Okay. The cloud price of AMP here is increasing, not decreasing. So that tell me, demand is demand.

Demand is outpacing it. Thank you, Jen. Right, right. Exactly right.

Demand is not pacing. There's the price. But the question with everything related to monetization is it profitable? And the question is it profitable to run that workload on the Amp P versus, hey, FOYA, Blackwell.

you fully depreciated the asset. Of course it's profitable. You fully depre appreciated it. The only cost, the only cost really is power.

Yeah, right. But so, so six, seven years old. Yeah, So the cost, you have to factor into opex as well. But then if it cost.

10 times more to run that workload from the opex perspective on Ampire versus Blackwell. But if I can't get Blackwell, if I have more work that I have capacity, yeah, of course I'm gonna use that ampi. It's fully depreciated asset, so I I have to disagree. Yeah.

Which is what I love about these conversations. That's fine. Yeah. Well, that continues to be what I call, Jenssen's Paradox.

And I think it's gonna continue to be a question 'cause I don't think people factor in the opex end of it. And then also the idea of obsolescence is, is of course, in the opex. I mean, these guys know how to use spreadsheet. I think you're bringing up the key issue, and this is, the fact that.

Obviously there's an entire ecosystem, there's an entire value chain. Yeah. And everyone needs to have a positive return on investment at some point. Yeah.

In that value chain. Yeah. And that's brought up the whole question, especially by the financial community of, is there an AI bubble. Well, in terms of demand, there is no AI bubble.

we, as long as I can see or that we forecasted, we have wave after wave after wave after wave coming. Yeah. That is pushing just the demands that we have that are as, up astronomically. but, does that mean that we have the business models in place for all parts of that value chain to make money effectively?

That's a good question. And I don't know. And I think you do a good job of bringing that up, Leonard. I don't know that all parts of that value chain are really set up to effectively do that.

And that is still the vaccine question for the industry. Yes. Right. And then it's not just about monetization as profitable monetization and in terms of.

One of the things, especially for, here's another topic that came up. especially as it relates to AgTech thin ops, the economics of, AgTech is not well known. One of the things that I'm hearing quite a lot these days from the developer community is that the unpredictability of the cost of, work, when you assign it to an an agent, it's inconsistent. Mm-hmm.

Because you never know, like, like say let's for the same unit of work, the cost can vary dramatically and we have to still remember these things can still confuse themselves and get into these infinite loops. Well, and we have agents using agents, so you never know how many agents it's calling or how many models it's using or how much data it's used. Matter fact, you can ask. The same question twice of an agent saying, well, I want more detail.

Right? Or are you sure about this? Just to make it double check itself. And that changes, everything changes, right?

And Jensen had an interesting concept that I hadn't heard of before. I hadn't really thought about. when you hire an engineer, you need to give that engineer a token budget, okay? And you make sure he's, if he doesn't spend his token budget, he's outta here.

Right, because that means he's not being as productive as he could be. but to Jim's point, he could also rapidly outspend his token budget and become a cost problem. So I think this uncertainty is something we'll have to live with for a while. Yeah, it will be solved.

There's the efficacy of how tokens are spent. It's good that he brought it up, but, there's a big. Question about what is the value of a token? not from a cost perspective.

The cost perspective each, the cost of each token, depending on the system is somewhat predictable. But from at the application level, it's completely. Unpredictable. It's very difficult.

Whereas, let's say for a deterministic system you have for transaction costs and pricing, that can be calculated. But how do you do that when you're using agent systems? it can be variable and yes, very dramatically different. So the variance can be different.

How do you cost for it? So the whole finops. Topic, I think has largely been avoided because of exactly this. and then also introduces risk to the value conversation around AG agentic ai.

But anyways, I, I think this is something, we'll, probably it's great topic next year. Yeah, it's next year. It's this year. Denial.

I, I, I agree. I it's a great topic and it's great that you keep bringing it up and that, you've made it a key focus lender. It really is. Yeah.

I think remember this year is the year Gentech AI became. Real impossible. Yeah. But we, we'll worry about the economics down the road, right?

yeah. Right now people are saying, oh my God, what's I, I like the way Jensen put it. He says, every company needs to have a strategy for open claw. And I more thought about that thought, wow, he's right.

This is that transformational. What is your open call strategy? Everybody from IBM to Pillsbury. we're working on it.

Matter of fact, we've been working on it for a month. Yeah. And the first thing we found out was open claw was dangerous. Yeah, yeah.

Oh my God. Exactly. I love that. Yeah.

I like the way that Nvidia says when you set up your first set up, your open claw, it has zero, access. You have to grant it access, it can't inherit your ACL L right? It can't in inherit your capabilities. You've gotta give very specific, very fine grained access and very specific files and data.

Although, or Jim's rice's gonna run amuck, it's gonna put you out of business. Although I will tell you, I, I discovered some really interesting things about through. Readiness of being able to sandbox, even put together a sandbox environment. So it's not stuff that's publicized, it's stuff that's in my research and the discovery that I did at gt, GTC.

So, clients, if you wanna find out, gimme a ring. Yeah. I want to give a shout out to all the people who have been doing AgTech on device. the likes of Apple, Google, Microsoft.

Lenovo Honor all these guys that have been looking at on device and trying to figure out how to do personal ai. Jim, you alluded to this earlier, that this is really about on device and doing something more localized than, your own thing. Right. this stuff is really hard and one of the things that I really struggled with when Jensen got up on stage and said, this is enterprise grade, prove it.

Because these other companies have been doing, been struggling with this for two years, right? And they're being labeled as being behind and all this other stuff. now you have this very dangerous thing called open clause. You mentioned Jim.

And then in two months, probably at best, now you're gonna come and say, claim that this is enterprise grade chart looks great. I'll tell you right now. Most people don't even understand this stuff. That's true.

I much less know how to secure it and I would, what I want to see in 2026 is nemo. Claw to prove its claims. I'm not believing this stuff face value because we're working on, we're working on that now. So gi give us a month and we'll see.

Yeah, see right here. And you're gonna right here. Call. We're doing it now.

We're doing it live. exactly, we got three software developers on staff And to the audience if you want to know what's up with all this stuff. 'cause this is probably the most important question coming outta GTC. Jim and his frigging amazing kids and the developers and the team there.

Definitely. you have to talk to Jim, but, oh, one thing before we call it, wraps. Jim, I wanted to get your take on AI grid. did you have any.

Points of view on that because that, for the telco industry, it's a big thing. There's a lot headlines about it. There's a lot of attempts to associate with some sort of. Game related to AI ran, what was your like impression?

I'll be honest with you. I was still rather disappointed. obviously there were a lot, MWC was all about AI ran and how we're adding AI into the network. especially, with six G being, basically an AI defined.

solution. however, Nvidia had a big announcement with Nokia and several, telcos over there in Europe, talking about AI ran, seeing the, I think that the telcos and the equipment companies start to see more opportunity integrate intelligence into the network, but they're still looking at it from a network management perspective. In other words, how do I more efficiently run my network? How do I do this?

How do I do that? I don't think that they truly see the opportunity and they need to, I really think that the business model for the telcos has to change, and we've been telling 'em this for over a decade now. Yeah. But the fact is that six G may be the last truly defined network because they're gonna have to be upgrading their networks.

Yeah. Whether it's hardware, software, or both, every two to three years going forward just to keep up with this demand. Once again, like my robotics example, the network becomes a critical part. The workflow of the platform that you're looking at.

In a lot of cases, you have to do the intelligence in the network because you can't afford to go to the cloud for deterministic applications. For applications that just require low latency, a 10 millisecond delay in a robot can be very hazardous. You can't afford that. You have to have sub one millisecond.

Yeah. So I think that, and. I stood up in front of a whole crowd in front of Noia event and said, I don't think you guys have a sense of urgency. and I really don't.

I think that, I think that they need to get urgency. So I still think that the network and it'll be interesting 'cause I think this is an area of innovation and if the traditional equipment guys and telcos don't rise up to the task, somebody else will. Yeah. And I love the fact that you were the last, last person to ask a question.

I was like going, oh my God. So there he is. I haven't seen him this whole time. And you were telling, what did you tell the, The mic handler when I was up there asking Oh, I, you were right before me.

And, he was trying to hand me the mic. I says, oh, don't worry, he's gonna ramble for a while. I love that. But my question was actually pretty short.

But you know what? You keep bringing up networking. and I guess, we'll wrap up with this. I'll just make a comment.

I didn't hear too much, which was really weird given, I think we all recognize how important networking is and interconnect is for the continual scaling, especially when we look at all this heterogeneous stuff. Mm-hmm. The networking has to look really different. It's not gonna be GPU or the old roadmap centric, it's gotta diversify.

But the CPOs, the CPO stuff, I, I was surprised to see the semi committal, push. For the CPO, stuff, and they introduced new, rack, or a switch, right? The spectrum X. Oh, I'll be honest with you, I'm seeing more interest, I think, more pull on the CPO stuff than I am push.

So, and then instead of seeing the push, I think maybe from NVIDIA and other technology providers, I think the end market is going to pull that in even quicker. So I, I would agree. I was surprised I didn't see more from it, maybe from Nvidia, but I still think that opportunity and that demand is there because I think the hyperscalers. The neo clouds and all these guys building data centers out, realize how much that's what the value of that is from a Yeah.

complexity perspective, a cost perspective, just an overall TCO. So I think we're gonna see a lot more pull on the networking side than we ever have before. And I think in a really weird level, I'm hearing a lot of that pool that you're talking about coming from, like, the, The, Sienna, menta, that kind of level where you're talking about, this really high capacity, low latency, networking for, scale across is, I don't know, that's where I'm hearing a lot of it, but where this is real, and the requirements are being driven at that level.

Not so much at, let's say the rock level, you know what I'm saying? No, I get it. Yeah. That's just what I'm noticing.

yes. Awesome. And gentlemen, I think this is one of the best, episodes we've ever put together. we, we did some really, tasty disagreements, which I love.

Isn't this why we started this Carl, right? Mm-hmm. Yeah. It all started with our fundamental disagreement around the economics of ai, right?

Yes. People don't realize that we didn't do this to come and we're still talking about it, like, oh, hey, I want you to just, repeat what I just said, Ah, you guys are awesome. Love doing this with you guys, thank you so much. thanks to the audience for listening in, and please reach out to Jim McGregor.

he and his team are working on one of the biggest problems, today, which is AgTech economics. taking a look at. What are the finops aspects of all of this? How do you, mm-hmm.

How do you stabilize our, the ROI argument and the value propositions around, gentech applications? And of course, curious research, www.curiousresearch.com call.

The mind expanding AI Cambrian expert, reach out to him to understand what is happening in the landscape of AI and the silicon, as well as software that's, changing industries and changing the semiconductor industry and pushing the AI industry forward. And also, please subscribe to our podcast. Yes, we got one of these. We're featured on YouTube, obviously because of the YouTube placard.

And, check out the audio version on Bud Sprouts and, your favorite podcast platform. Also subscribe to the next research portal, at www.next-curve.com, as well as our substack.

You can find some substack and of course, all this for the tech and industry insights that matter with curious research and Cambrian. AI research, LLC. Gentlemen, thank you so much and we'll see you at toward the end of the month. Right.

We gotta do our March recap, so, yep. Have a great weekend. Cheers. Take care.

Bye.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Rebooting Enterprise AI with MCP and KubernetesPractical AI · on Anthropic Claude88 / 100
  • AI Is Having Its Dropbox MomentAI Proving Ground Podcast · on Anthropic Claude85 / 100
  • Generative AI Meets Accessibility: Benchmarks, Breakthroughs, and Blind Spots with Joe DevonAI Engineering Podcast · on Anthropic Claude85 / 100
  • Open-Source AI Battle, Google Throttles Meta, Micron Margins Moon | Edward Coristine & Tai Groot, Chad Rigetti, Pim de Witte, Yadin Soffer, Jack Morris, Neil Movva, Jakob Diepenbrock, Chris AltchekTBPN · on Anthropic Claude80 / 100
  • CAD, BIM, and the AI Leap: Qonic & Raven!AI Across The Product Lifecycle Podcast · on Anthropic Claude80 / 100
  • No Code, No Problem - How a Speech Pathologist Built an AI StartupDesigning Successful Startups · on Anthropic Claude79 / 100

More from The neXt Curve reThink Podcast

All episodes →
  • Silicon Futures for May 2026 - Qualcomm hyperscale AI, Cerebras IPO, Huawei 1.4 nm68 / 100
  • Edge Continuum: Where AI belongs from sensors to cloud (with Azita Arvani)86 / 100
  • Edge AI Foundation at Sensors Converge 2026 (with Pete Bernard)55 / 100
  • Silicon Futures for April 2026 - The AI CPU craze, Qualcomm's custom AI, Google TPU 8 explained80 / 100
  • The Blueprint for Agentic Security (with Raj Chopra)75 / 100
Explore the best B2B AI & Data podcasts →
All The neXt Curve reThink Podcast episodes →