TechSurge: Deep Tech Podcast · 2026-09-16 · 1h 12m
Key moments - from our scoring
Substance score
70 / 100
Five dimensions, 20 points each
The inference era is fundamentally reshaping how AI systems are bought and sold. Where cloud operators once demanded commoditized components, they now purchase integrated systems from vendors like Nvidia and AMD because frontier LLM models (2+ trillion parameters) require terabytes of memory distributed across multiple GPUs, CPUs, networking, power, and cooling - a complexity that startups struggle to replicate. Austin Lyons, a semiconductor analyst at Creative Strategies and author of Chip Strat, explains that inference spending has surpassed training spending for the first time, creating new opportunities and constraints. The conversation explores system selling dynamics, why Nvidia's early integration of rack-scale solutions created defensibility, and the emerging pre-fill/decode split, where inference workloads are being disaggregated based on different compute and memory requirements. Pre-fill (processing the entire prompt in parallel) is compute-intensive but memory-light, while decode (generating tokens sequentially) is memory-bound but compute-light. This workload split has created opportunities for specialized AI ASICs like Grok and Cerebras to compete in the decode space, though consolidation may be inevitable as only three or four vendors typically dominate semiconductor segments long-term.
Frontier LLMs require 2+ terabytes of model weights in memory, but individual GPUs only have ~288GB of HBM, forcing models to be split across 4-8+ GPUs. This creates a complex integrated system with GPUs, CPUs, networking, power, and cooling that Nvidia solved first at rack scale (e.g., Grace Blackwell NVL 72 with 72 GPUs), and customers now prefer turnkey solutions from single vendors over assembling components from multiple vendors.
Pre-fill is the initial parallel processing of your entire prompt to establish context (compute-heavy, memory-light), while decode is the sequential generation of output tokens one at a time based on previous tokens (memory-bound, compute-light). This split means different silicon architectures optimize for each phase - pre-fill benefits from high compute but low memory, while decode needs high memory bandwidth and lower compute.
Yes, but it requires hundreds of millions in capital and very careful trade-offs. Startups like Cerebras took a risky approach by inventing wafer-scale engines and all associated cooling/power/communication systems from scratch. A smarter approach might focus on one piece of the system (like decode-optimized silicon) and partner with CPU vendors like Intel, as Samanova did with x86 integration.
Decode is memory-bound - GPUs are general-purpose and designed for balanced compute/memory, leaving them underutilized during decode. Specialized decode ASICs like Grok and Cerebras use SRAM-based memory architectures optimized for high bandwidth, enabling 1000+ tokens per second versus GPU decode rates, which matters for latency-sensitive applications worth paying a premium.
Unlikely in the near term because buyers prioritize speed-to-market and simplicity over cost optimization. However, as market matures and standardization increases, the CPU/GPU precedent suggests 3-4 dominant vendors will emerge, each selling complete systems, while some competition may come from neo-cloud companies that buy diverse silicon and sell inference tokens as a managed service.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode packs substantial technical and market-structure insights, particularly around the shift from training to inference spending, the pre-fill/decode split, and why system-selling emerged. However, there is notable repetition and throat-clearing (e.g., multiple restarts explaining the same concepts, lengthy throat-clearing in Speaker A's longer monologues) that dilutes density. Most operators would learn concrete ideas around 2-3 per 10 minutes rather than the 4+ that would merit a 18+.
Almost 2% of US GDP will be spent on AI infrastructure this year, nearly double the amount from 2025. Yet these huge numbers hide a quieter chain. In the last year, the sum spent on deploying models in production, known as Inference, was for the first time larger than the amount spent on training.
pre fill, I've got all these GPUs that are just doing all of this, uh, parallel computation and it actually doesn't require a huge amount of memory. And so I've paid for this, uh, high bandwidth memory that's very expensive by the way. And in pre fill, that HBM is actually just sitting there being underutilized.
The episode covers some genuinely fresh structural insights (e.g., the inference > training spending inversion, pre-fill/decode silicon disaggregation, neo-cloud financing dynamics via hyperscaler backing rather than traditional VC). However, the core framework - Nvidia's dominance through system-selling, startups' need for anchors and scale, the eventual consolidation to 3-4 winners - is well-worn in tech discourse. The guest articulates existing ideas clearly but rarely arrives at truly contrarian or first-principles arguments that would surprise a thoughtful operator.
we've got uh, a bunch of GPUs that are being heavily utilized for compute and their memory is being underutilized. And then we have a bunch of, during the pre fill phase, then we've got a bunch of GPUs in the decode phase that are um, basically under utilizing their compute and just totally utilizing all their memory.
there always seems to be like three or four vendors in a certain thing. Whether you look at like wafer, fab equipment, foundries...We are in an era, of course, as happens whenever there's like, drastic innovation where a ton of competitors have popped up.
Austin Lyons is a semiconductor analyst with relevant credibility (Creative Strategies, Semi Doped podcast, Chip Strat Substack) and demonstrates genuine expertise in hardware markets and supply-chain dynamics. However, he is a pure analyst/writer rather than an operator who has actually built or scaled a chip company, managed a data center, or made capital allocation decisions at a hyperscaler or neo-cloud. His insights are informed but filtered through secondary research and customer interviews, not direct execution experience.
Austin Lyons is a semiconductor analyst at Creative Strategies, co host of the Semi Doped podcast, and the author of the Chip Strat Substack.
I definitely was the type of person where right away I was like, okay, this is different, this feels funny. I need to dig in and understand it and try to understand both sides.
The episode includes concrete examples (Grace Blackwell 72-GPU rack, 2 TB+ model memory, Cerebras wafer-scale, Grok/TensorDyne, Rivian autonomous driving workload) and references specific metrics (800 tokens/second, ~$200M+ chip startup funding, HBM costs rising, 100 megawatt constraints). However, many claims lack hard numbers: exact inference > training spend split not quantified, neo-cloud margins unspecified, OpenAI's in-house chip performance vs. GPU baselines not detailed. Several important assertions rest on anecdote (son's 70K-line game) rather than verifiable data.
It takes hundreds of millions of dollars now for a chip startup when maybe back in the day you used to do several rounds of just like a couple million dollars to prove out your little proof of concept.
even a GPU when it was running decode, just the way that GPUs are more general purpose and designed and their memory hierarchy decisions, maybe they can only run it like [paused, restarts] 1000 tokens a second
Host David Goldman asks several sharp clarifying questions (e.g., 'Why wouldn't a cloud buyer just piece together components instead of buying full systems?', 'Is pre-fill/decode split always necessary or use-case dependent?', 'How much does cost play into the equation?'). He also pushes back thoughtfully on the guest's neo-cloud investment thesis. However, many follow-ups are surface-level (e.g., restating the guest's point rather than probing deeper), and the host rarely disagrees or challenge-test claims that warrant it (e.g., the claim that regulatory hurdles won't prevent further consolidation, or that on-prem diffusion will match cloud scaling). The dialogue feels more like co-exploration than adversarial interrogation.
But if I'm a cloud buyer, Nvidia has famously high margins and they charge that on all of the different parts of the system, not just on the gpu. And if you go out in the valley, there's all sorts of companies who are going and offering one piece of this puzzle...what kind of value do you get from getting it all at once?
So is tokens per second per user that interactivity KPI still the right one to think about for startups? Or are there changing needs because of power constraints, cost constraints, new workloads?
Computed from the transcript - who did the talking, and the words that came up most.
Almost 2% of U.S. GDP will be spent on AI infrastructure this year, nearly double 2025's figure. But beneath those headline numbers, the composition of that spending has quietly flipped: for the first time, dollars spent on running models in production now outweigh dollars spent training them. In this episode of TechSurge, host David Goldman speaks with Austin Lyons, a semiconductor analyst at Creative Strategies, co-host of the Semi Doped podcast, and author of the Chipstrat newsletter. Lyons previously worked as a hardware engineer at Intel and as a product manager on John Deere's autonomous tractor and Blue River Technology teams before turning to full-time chip industry analysis. The conversation opens with why AI buyers have moved from assembling commoditized parts to buying entire pre-integrated systems, tracing how Nvidia's rack-scale approach, exemplified by its 72-GPU Grace Blackwell racks, made turnkey deployment the default, and why that raises the bar for any chip startup trying to compete.
Transcribed and scored by The B2B Podcast Index.
Speaker A: No one yet has really brought a chip to market that was designed specifically for LLMs.
Speaker B: Well, that's capitalism. They're not going to do it if they don't benefit. Almost 2% of US GDP will be spent on AI infrastructure this year, nearly double the amount from 2025. Yet these huge numbers hide a quieter chain. In the last year, the sum spent on deploying models in production, known as Inference, was for the first time larger than the amount spent on training. This shift is creating monumental ripple effects through the buyers, designers and suppliers of AI infrastructure. Austin Lyons Austin Ly Austin Lyons Austin Lyons is a semiconductor analyst at Creative Strategies, co host of the Semi Doped podcast, and the author of the Chip Strat Substack.
Speaker A: It takes hundreds of millions of dollars now for a chip startup when maybe back in the day you used to do several rounds of just like a couple million dollars to prove out your little proof of concept. And then if they get access to 100 megawatts, they're gonna have to ask themselves, how can I get as much revenue as possible out of this hundred megawatts?
Speaker B: So you wrote an article about the conditions are for the next trillion dollar chip company and you in this conversation we talk about why AI now gets bought as entire systems, how the inference workload split in two, how a new type of cloud company grew to over $100 billion, and where the next trillion dollar chip company will emerge. Hi everyone, this is the Tech Surge Deep Tech podcast presented by Celesta Capital. Each episode we spotlight issues and voices at the intersection of emerging technologies, company building and venture investment. David I'm David Goldman, partner at Celesta Capital. If you enjoy our discussions, now is a great time to hit the subscribe button and you can leave us a review on your favorite podcast platform. Visit techsurgepodcast.com to sign up for our Substack newsletter and check out all of our past episodes. Austin, welcome to the show.
Speaker A: Thank you for having me.
Speaker B: AI infrastructure is in the middle of one of the biggest capital buildouts in history, but I think under the surface, even though the numbers are going up, there's been a pretty big change as the dollar spend on Inference has gotten bigger than the spending on training. And this is a trend that you've covered a lot in your writing and something that you, I think have a very unique perspective on. So I want to spend a little bit of time unpacking the implications of this shift. So maybe let's start at the top and how these systems get sold and what system selling is so Historically, when cloud buyers went out to the semiconductor industry, they used to be trying to commoditize server hardware as much as possible. So they would work with industry groups, they would try and get standard designs done so that they could drive down pricing and negotiate with each part of that design. But increasingly and over the last few years, they've been buying whole systems. So maybe let's start at the top and just say what goes in a system? And why is this trend happening?
Speaker A: Yes, yes. I like the idea of starting with systems when we're talking about AI, because that's really what it's all about these days. It's rack scale all the way out to of course, like full data center clusters. Um, so, you know, we're talking about the compute at the heart of these Systems is ultimately GPUs, and that's kind of what gets the airtime or AI accelerators. But it's actually a full system. So if so, stepping back, the question is, you know, how did we get here? And why are people buying a full system from one vendor and not commoditizing it? So I think if you start and you look, obviously the workload that we all care about today is LLM inference. Um, and so if you're running at the frontier and you're wanting to use the Latest and greatest OpenAI model or anthropic model, these are models that are like 2 trillion or more parameters. And what that means is even, um, as you quantize these and try to run them, uh, without using as much memory as possible, you might still need something like a terabyte or two terabytes or several terabytes of memories just for the model. Because with these models, as we saw with the scaling laws over late 2022, 2023, 2024 is the bigger the models, um, and the more compute you have, the better answers you get. And so way back in the day we might have taken, uh, an AI model and put it just on one gpu. Um, but if you start to think about like today's frontier models, you need maybe terabytes of memory. And yet a single GPU, um, only can have, for example, like 288 gigabytes of HBM on it to store these model weights. And so quickly you say, well, wait a minute, one model can't fit on a single GPU's memory. So what do you have to do? Oh, you need several GPUs, so you might say, okay, take that model and split it across, you know, four GPUs or eight GPUs. And that's how we started to work our way into needing like, oh, this isn't just a single chip anymore, this is a whole system. And that's how we got our way to, for example, like the Grace Blackwell, uh, NVL 72 rack. You know, it's a rack with 72 GPUs in it. We've gotten to a place that we are talking about AI systems, even just for inference, even if you were only to have one rack. You've got, you've got GPUs, CPUs, networking, uh, power and cooling and um, so back to the question of um, you know, is this, why isn't the industry commoditizing it? You know, I think a big piece of where we are at today is Nvidia, uh, did a really good job of getting to this rack scale first where they said, hey, we can help design the whole system and make sure that it works across 72 GPUs and 36 CPUs and that they can all talk to each other. And oh, by the way, it also takes software and compilers so that you know, you model developer can write your model and actually run it across all this. Even though the industry always wants um, commoditization and competition in multiple suppliers, Nvidia did a good job of frankly getting there first and making this complicated system turnkey so that the software developer, it's as easy as writing their model, deploying it, and it quote, unquote, just works.
Speaker B: So you need a very big complex system to make these AI systems work, to make the LLMs work in the cloud. But if I'm a cloud buyer, Nvidia has famously high margins and they charge that on all of the different parts of the system, not just on the gpu. And if you go out in the valley, there's all sorts of companies who are going and offering one piece of this puzzle. And if I'm a cloud buyer, why would I be willing to give so much margin to Nvidia instead of trying to piece each piece together and buy my GPU from Nvidia maybe, and my networking from someone else and my software from someone else and my cool from someone else? What kind of value do you get from getting it all at once? Is it around speed? Is it around simplicity? What are the kind of trade offs that people think about?
Speaker A: Yeah, I think so. We are not yet in the era where people are trying to squeeze down their costs as much. We've still been in the era of I need to IT'S speed to market. We saw that with Elon Musk and Xai, um, where he's been able to stand up data centers very quickly, um, using Nvidia systems. And I don't think Elon's goal was to say how can we do this as cheaply as possible, but more of just how can we stand up this, compute and get productive tokens out of it as quickly as possible? So even though a company can go out and potentially pull the components from many different vendors, um, there's just a desire to buy it and stand up quickly. And then to your other point of, there is definitely ah, simplicity there. And we see even AMD with their Helios rack that they're bringing to market, same kind of thing, they see that pull from customers that say just, I want to be able to buy the full rack. I want to be able to get it and stand it up as quickly as possible and get software running as quickly as possible.
Speaker B: So each of Nvidia and AMD have had decades to put together this puzzle and tons of financial firepower to do M and a Nvidia acquired Mellanox, AMD has done a bunch of acquisitions to build out these system selling strategies. Is it possible for a startup to compete now that you have to sell systems?
Speaker A: Uh, yeah. That is a tough and very interesting question because to your point now, if you're a startup, you can't necessarily just come in and say, oh yeah, I'm going to build a better GPU or better AI asic. People are going to say, okay, great, now what do I do with that? Um, are you going to make your end customers have to piece together the whole system? We just said that they want to stand it up as quickly as possible and deploy it as quickly as possible. Or are you as the startup now going to have to take on building the rest of the system? Um, I think there's probably different approaches here. And by the way, this is why we see that it takes hundreds of millions of dollars now, um, for a chip startup when maybe back in the day you used to do several rounds of just like a couple million dollars to sort of prove out your little proof of concept. And so yes. So can a, uh, startup compete here? Yes, but they have to be very smart. I think it could also be death by a thousand paper cuts if you're trying to invent a new AI accelerator and you're trying to invent your own proprietary networking and maybe your own proprietary cooling. And actually for example, um, Cerebras, they took a Very interesting approach where they're uh, an AI accelerator startup, they've gone public, um, they make these wafer scale engines is what they called them. And not only uh, which by the way, for those listeners who aren't as familiar, instead of just making a chip, um, instead of taking a whole wafer of chips and dicing them up and making these individual chips, like two GPUs and a CPU that you package, they said, hey, why don't we just leave it at the wafer scale and have all these chips that are on there and actually make the whole wafer our compute engine. But they had to invent now how do you make them all communicate, how do you cool all of that, how do you deliver power to all of that? And they had to basically invent all of that. And that's going to take a lot of time. So I think any other AI ASIC who followed could have an example like Cerebris to look at and say, wow, they did very innovative technical things here, but they had to invent everything. And that's time consuming and expensive. How should we maybe do this differently?
Speaker B: Yeah, I mean, I think you see some parallels with the build out of the Internet infrastructure in the 90s, early 2000s. You start with proprietary, then as open standards proliferate, there's more capability for more people to offer things and you can get a little bit more mix and match. I know you've seen Samanova, for example, has partnered with intel to offer x86 CPUs in their system alongside. So there seem to be some M trends, even though incumbents clearly have an advantage in the system selling. One of those trends that I think you've touched on a lot in your writing is as we move into this inference era, you're starting to see the inference workload itself get split and it gets split into what's called pre fill and decode. So can you tell us what are those two things and why might you want to split that workload into different types of silicon?
Speaker A: Sure. So at the end of the day, if you just think of like a simple talking to an AI sort of chatbot, you write a paragraph of um, explanation. Hey, I want you to go research this and give me an answer. And the first thing that the AI system under the hood, the model needs to do is read through the whole prompt that you gave it. And how this works is this is the pre fill stage. And what's happening is in parallel, um, sort of one of the key innovations with the transformer model is this Idea of attention where you work through everything in the long, you know, let's just use English, for example, the long English paragraph, and you ask, you know, um, which of these words are sort of connected to other words in the sentence. And that way you can sort of like piece together the context of the whole paragraph and what each word is referring to. And so you can do all of this in parallel. You can go through and if there's 100 words, you can look at all 100 words in parallel and compare them to all one other hundred words and do some of these calculations to figure out like are they related or not. Um, that's pre fill. You're ultimately doing a bunch of, um, linear algebra, a bunch of matrix multiplication. Now when it's time to give you an answer, you know, especially if you remember back to like when ChatGPT first launches a lot slower, you know, you essentially would start to see this stuff kind of come out word by word. And that's because, you know, every word that I am about to say depends on the words that I just said. And then when I say that word and I say the next word, it also depends on those other words. So basically this is decode and this is where you're predicting the next tokens or the next words. If we're talking about English, and that is sequential, you can't do that in parallel. And people who are running this inference at scale started to realize, hey, wait a minute, pre fill, I've got all these GPUs that are just doing all of this, uh, parallel computation and it actually doesn't require a huge amount of memory. And so I've paid for this, uh, high bandwidth memory that's very expensive by the way. I mean, getting more expensive by the day. And in pre fill, that HBM is actually just sitting there being underutilized. Hmm, interesting. Okay, now I've got these other GPUs that are running D code and they're actually not fully utilizing all of their flops, all of their units of compute. Um, because actually decode is what's called memory bound, which is it's a lot of, um, hey, I made a prediction and I need to get all the weights and whatever I need from the KB cache and then I need to do a little calculation and then I need to go back and forth and back and forth one at a time. And you're not doing things in parallel and you're really just waiting on memory. And so, uh, you know, at the highest level, even Nvidia led the way with this saying, wait a minute, um, we've got uh, a bunch of GPUs that are being heavily utilized for compute and their memory is being underutilized. And then we have a bunch of, during the pre fill phase, then we've got a bunch of GPUs in the decode phase that are um, basically under utilizing their compute and just totally utilizing all their memory. And so naturally any engineer is going to look at that and say like huh, huh. Well maybe we should disaggregate these and maybe the pre fill should run on systems that have lots of compute but don't necessarily need all that memory. And maybe on the decode side we should really emphasize memory bandwidth and how quickly information could get shuttled around and maybe it doesn't even need quite as much compute. Or we should essentially over index on the memory part. And so naturally right there, um, we started to move into a world where Nvidia shipped um, for their systems, a software layer called Dynamo which helps orchestrate this and would let you take Nvidia GPUs and it would orchestrate across um, and that actually gave rise to the Grox and the cerebras. These um, AI A6 startups who actually were started even before uh, transformer based LLMs was sort of the defining workload of our area. But they had made architectural choices where they used a lot of this really fast on chip memory called sram. And it's actually where you use transistors to store the memory instead of um, dram, which is capacitors and transistors, which is what HBM is made of. We won't go way down into those memory details. But basically these early startups had made a bet on having really high memory bandwidth. And once the workload got separated into pre fill and decode, they could actually sort of raise their hand and go oh uh, wait, we're actually really good at decode. And in fact we can go even faster than GPUs. And this kind of took us from the inference era where everything was on GPUs to saying, oh wait a minute, what if you could slot in one of these AI ASICs that are heavily built on SRAM that can maybe unlock 1000 tokens a second, whereas maybe a GPU when it was running decode, just the way that GPUs are more general purpose and designed and their memory hierarchy decisions, maybe they can only run it
Speaker B: like is that something that you should always do? Do we always need to split up pre fill and decode or Are there just certain applications where it's really good to have speed? I'm willing to pay a premium for speed, therefore I'm willing to go through the hassle of splitting these things up and having orchestration software and having different types of silicon and all of the things that are associated with that versus just running it in the GPUs or the system that I've already bought.
Speaker A: Yeah, yeah, yeah. There's like so many different nuances here. So if we take a look at the hyperscalers when they're deploying, uh, you know, tens of billions of dollars of GPUs and as they're starting, which of course early on, GPUs very general purpose, very flexible, uh, you know, gives you the freedom to change your workloads. But let's say you're an OpenAI and you're like, no, these are our specific models that we know that we are iterating on like the biggest, the frontier one and the medium sized one and the small one. You start to say, hey, uh, we should really sort of cater to this workload's needs. And therefore, therefore it would make sense to deal with the complexities that you pointed out of having to split up, pre fill, decode, orchestrate that um, it might be worth it, ah, for. And you know, maybe that unlocks for example being able to sell an ultra premium tier where you get really, really fast inference and maybe there's going to be a small subset of users who would pay, you know, 10x more to get tokens that are 5x faster. So I think there are certain uh, model labs and hyperscalers that are at scale that are saying the complexity is totally worth it. We're willing to deal with different SKUs, if you will, like different chips. Um, now on the other hand, I think there's going to be tons of enterprises and arguably the long tail of consumers who of course they want inference and they're thinking about cost and they're thinking about speed, but they aren't going to want to manage all of that complexity necessarily. Um, so for them, I don't think where we're at today, they're not going to necessarily need that complexity. Now of course, everyone always wants, ah, faster inference. Um, so if there's a way for them to get that outcome without having to deal with all the complexity, of course they're going to want it. For one, when I'm trying to think about this space and ask like, oh, is it just one size GPU fits all or is it going to be many, um, SKUs, SKU, you know, many different chips. Where's this going to go in the end? I do look to CPUs. Um, like when you look at like uh, any CPU vendor or any um, cloud, you know, like Google cloud, they don't just deploy one CPU even for customers who want to rent them from their cloud. They still have sort of a portfolio that say like, hey, this one has a lot of money memory in case you're running a database. Oh, this one's actually pretty vanilla and it's nice and cheap if you're just running like an API server and you'll see there's still like sort of shapes of. We got this family of chip, this family of chip. I do think there's a world where we could get to having different uh, you know, a couple different shapes of AI accelerators. Um, now where we are, what we led into is we said okay, we had this training era and then we had this inference era, those are all on GPUs. And then we've got this next era where it was like, hey, let's have a GPU plus this specific um, decode chip that's you know, made a good memory hierarchy trade off so we could have high interactivity is what they call it, which is like really fast tokens. So where we are today though is those are multi vendor. So it was like, you know, um, Nvidia plus Grok. Of course Nvidia acquired Grox. They're trying to bring that on their house, um, you know, someone else plus Cerebras, AMD plus Cerebras or Trainium plus Cerebras or whatever. And now uh, my thought is that we will ultimately go to those SKUs living inside the same silicon vendor. Um, because it starts to get complicated when your route to market is to depend on another company, right? Like oh, I sell a GPU and they sell a decode specific thing. Um, it does feel like it is working right now. There's definitely demand for it.
Speaker B: So I mean, yeah, it feels like there's this inherent tension in the market because you want to buy systems. That's what customers are saying. They want someone to do the work of putting all of these things together. But then at the same time they also want the right silicon for the right job and they want specialized things for decode and they want to have the right proportion of CPUs to the workload. So it's not like a one size fits all system.
Speaker A: Right.
Speaker B: Ultimately, do you think that you get more and more fragmentation here? Or are you going to have customers saying, Nvidia, please solve this problem for me. Buy Grok, buy the next company, buy the next company. Keep selling me systems, you know, amd, buy Thales, buy, you know, work with Cerebras. Figure it out for me. Um, because I mean, to give you like the counter side, there are separate companies for CPUs and GPUs. Like, we've decided that those are separate enough markets that they can have standalone companies. We can figure out the system with CPUs and GPUs from different vendors. Like, that's been a solved problem for a long time.
Speaker A: Yeah.
Speaker B: So maybe that could be an end state where in fact you have decode silicon. That is a completely different market and we give it a catchy name like DPU or something, even though that's a one, not that one, because it's already been used.
Speaker A: Right, right, totally. I do think, you know, zooming way out, when you look at the semiconductor industry in the long run, there always seems to be like three or four vendors in a certain thing. Whether you look at like wafer, fab equipment, foundries. It's complicated, but we're getting back to maybe having two or three, um, CPU vendors, GPU vendors. We are in an era, of course, as happens whenever there's like, drastic innovation where a ton of competitors have popped up. So I think it would not be crazy to zoom out and say, like, okay, in the grand arc, maybe only three or four people will shake out and therefore there'll be some sort of consolidation. Um, I personally think this could look, this isn't as simple as just like, oh, Nvidia buys them, AMD buys them. Even though we're seeing some of that. Um, of course there's regulatory things to talk about there, but I actually think if you look up and down the stack, there's actually a lot of competition. For example, um, at the NEO cloud layer. And, um, they might be. A NEO cloud might be incentivized to say, like, hey, I can tell that the customer, they want the right silicon for the right job, but they also don't want to deal with the complexity. Um, I could buy a bunch of different silicon and I could deal with all the complexity and I could just try to give them, I could sell them tokens as a service and try to give them the speed or the cost or whatever that they're looking for. And I could take that onus on and differentiate from other NEO clouds. So I do think there's still routes to market where startups um, can come in today and they could just say like, hey, I am the best decode solution. You should try me. Um, or maybe I'm a pre fill solution, um, and yet be able to figure out a way to get to customers and to let, uh, the end, sort of the end developer, let's say to not have to deal with all that complexity. Um, therefore you might just keep pulling on that and play it forward and say like, oh, wow, could a AI ASIC startup ever merged with the NEO cloud? I don't know. Maybe, you know, like, maybe it won't just be Nvidia or AMD buying all these.
Speaker B: You touched on an interesting thing here because I think neoclads are really under discussed when people talk about AI infrastructure. We, we tend to focus a lot on semiconductor companies and systems companies and not as much at the cloud layer and sort of have a pet theory that it's because most of the traditional VCs missed out on those as investments. And so we don't like to spend too much time giving credit where we don't get to claim any.
Speaker A: Yeah, yeah.
Speaker B: Um, but the reality is that if you look at the handful of NEO clouds that came up in this first wave of NEO clouds, they've created a lot more equity value. I mean these companies are worth. The top three are worth something like $125 billion in public markets, which is a lot more than what we've been talking about with some of these chip startups. Why do you think that investors maybe had a lot of trouble understanding the first wave of NEO clouds and missed out on those as investments?
Speaker A: Yeah, yeah. Uh, for listeners, by NEO cloud, we mean a cloud company that started by just renting GPUs. I'll just talk through my hesitation when I first saw the idea of NEO clouds, and maybe I am a fair proxy for some investors. Um, maybe they thought this way, which was, you know, ultimately I saw. Okay, well, let's back up. Um, so who's renting from these NEO clouds? Well, it turns out a lot of it is hyperscalers, which is ultimately driven by demand from the biggest model labs. And so the question is like, well, wait a minute, you're saying that, you know, uh, OpenAI is running on GPUs that Microsoft is renting from some NEO cloud. Well, doesn't Microsoft just have their own data centers? And the answer is, well, of course they do and they're trying to build more. But, um, at the end of the day, we are limited by power. Access to power. Um, we are limited there's like, uh, financial reasons to go to rent versus buy, um, maybe even go off balance sheet as Capex has, you know, increased year over year over year. And so there's very legitimate reasons why even someone as sophisticated as Microsoft Azure might say, like, actually I want to rent GPUs from someone and I'm also building other data centers. But it's like maybe it's a stopgap. So I, looking at that, I said, okay, so you're going to have a company that's a Neo cloud, that maybe they have access to power and maybe they have um, a good way to raise capital against assets like this. For example, like Bitcoin miners historically have um, had this experience. Um, and maybe they're shifting into GPUs. And so I looked at it and thought, huh, okay, so they were doing bitcoin miners. Now GPUs are hot, so they're going to do GPUs, but they're just going to rent them to Azure. But Azure is also just standing up their own data centers. So like, how is this sustainable? They've got one, it's like huge customer concentration. Literally maybe one. Um, uh, but to their credit, that's how you get the financing, as you say, like, yeah, I've got Microsoft, who's gonna rent these from me. And so people will lend against that and sort of believe in that. Um, and so that's part of why I missed it. As these companies were popping up, there's a legitimate need in the marketplace for people who have access to, to power, who can get financing, who can manage this and who can move quickly to stand up GPUs to run models on them, to rent it out of bare metal or maybe at a higher level of abstraction. And even though they have serious customer concentration risk, so does everyone else in the semiconductor industry right now. Right. Like who's Nvidia's end customers? Like uh, even Nvidia, the biggest, baddest, they're selling a lot of GPUs to a certain set of customers.
Speaker B: Yeah, I mean it's not uncommon to see someone go public with 90% customer concentration these days.
Speaker A: Totally. That is the name of the game. If you look at any component supplier, like um, even in the, in the interconnect space or the switching space, like you look at like a credo, who made these active electrical cables, really awesome invention. And same thing they're selling it to, like even when they went public, they're selling it to a handful of customers. And so the name of the game too, for this era is like watching how this unfolds and seeing, like, how do some of those people who win one big customer ultimately win a couple more? Or how did they build out more of a portfolio and sort of reduce a little bit of that risk? But it just is the way that the industry, uh, works right now.
Speaker B: So if you wind the clock back five years before these Neo clouds got big, I mean, you have AWS is an incredible business. Amazon, Microsoft, these companies have fortress balance sheets, great relationships with Nvidia, all of the semiconductor companies, expertise in setting up data centers, software, customer relationships. It seems like to me that they could have done this, and certainly most people thought they would have, which is why so many of them missed out on these investments. Was it a strategic decision by them where they thought, hey, maybe this isn't going to be a big enough market. I'm not sure I want to spend all of the money and take the risk. Or was there some special sauce in what these NEO clouds were able to do in setting up quicker, in being a little bit more creative in financing, in converting old Bitcoin centers? Like, were they doing something different from what Microsoft might have done? Or was Microsoft ceding market share to them or some other factor? Yeah, not to pick on Microsoft anyhow. Correct.
Speaker A: Correct. Insert anyone here?
Speaker B: Yes.
Speaker A: Uh, maybe a little bit of everything. I mean, at the end of the day, this is risky business because the investments might be $40 billion this year and $80 billion next year and $100 billion a year after that. So obviously for these clouds, I think a lot of the existing players weren't necessarily. It's not that they didn't believe that the future that we're in now would manifest. I think it's all about timing. And I do think that they're probably not incentivized necessarily to just sprint out and stand up all these data centers. Obviously, it's a huge investment. Um, I think that the NEO clouds were able to, uh, move faster, take on more risk. They can kind of go for broke here. It's like, all right, cool, let's convert, let's buy a bunch of GPUs and let's get moving. And so, yeah, I just don't think that kind of like the innovators dilemma. It's like they're already big three clouds. They already have all the customers. I think they're probably looking at each other. The other big CSPs, probably like the AWS and Google and Azure all looking at each other and just saying, like, well, are we standing up? Are we all believing that this future is coming and investing in GPUs sort of at the same rate as each other? Um, but that could still not actually be enough supply to meet demand. Now, I do think that demand actually went higher, faster than everyone had expected, especially once we got reasoning models. And then now, of course, in the agentic age, um, and so therefore, I do think that even if they all looked around and said like, yeah, we think if we grow supply, like, this is a little risky. Ah, and it feels like these numbers are really big, but we think the demand will be there. But I do think demand has skyrocketed and gives an opportunity for neoclouds to come in and say, like, yeah, we'll fill that demand.
Speaker B: I think there was also maybe like a difference in interest level. You heard some of the hyperscalers make comments about how these bare metal GPU instances are low margin and sort of commodity and therefore they don't want to support them as much. Whereas the neoclouds read them as revenue and therefore good.
Speaker A: Yeah, yeah, yeah, yeah, you make a fair point. I mean, uh, anyone whose business was in renting CPUs or selling services on top of CPUs, the cost structure is much better than GPUs. And so, yeah, you could see CFOs saying, like, wait a minute, we're going to spend a ton of money and our margins are going to go down. Even if our margin dollars go up, uh, there's still conversations to be had that might make you slow down or hesitate a little bit.
Speaker B: So then, is this a durable state of affairs? Obviously this is a very quick moving market, but first mover more willing to take a little bit of risk, willing to take on lower margin. Those are not necessarily durable advantages that will last for a decade plus. What do you see playing out with these neoclouds? Do they get acquired by hyperscalers? Do they consolidate into a NEO hyperscaler or something?
Speaker A: Yeah, yeah, yeah, yeah, yeah. It's very interesting. I mean, I'm not sure there will always be need for 100 Neo clouds, but I do think that it's real demand and it will always be. I don't think that the big three clouds will always meet the needs of people, you know, indefinitely. Um, not only that, so like, okay, we've got, uh, particular workloads, you know, now where it's like, yeah, I want really high interactivity or medium actor interactivity. Um, but there's always going to be innovation, um, like the world Labs Company with Fei Fei Li, they're coming out with world models. Who's going to make the bet there? Like what if those don't run to the exact shape that all these big clouds have invested in? And I do think that there will always be, you know, new workloads or new demands or whatever that are popping to existence where uh, uh, uh, NEO cloud is going to pop up quickly and say like, I can meet your need and I can innovate there. And I think that that will always exist. Again, it may not be enough sort of cutting edge frontier demand for 100 Neo clouds to hop on it and try to offer something different. But I definitely think there'll always be a need for these more nimble, smaller essentially like gpu, AI, ASIC rental companies that can innovate a lot closer to where the frontier of like NeoLab and AI enabled software companies are innovating.
Speaker B: I feel like we can't talk about NEO clouds without at least discussing circular financing risk. Uh, this is probably the thing I hear most from people who are skeptical and they have questions about whether or not there is durability in AI infrastructure. And so the counterargument that people pose is the NEO clouds get investment in equity from Nvidia. In many cases they use that equity investment to buy GPUs. They then use those GPUs as collateral to take on debt. And then a lot of that debt is also sometimes backstopped by either their hyperscaler or by Nvidia itself. And so it all sort of perpetuates. And then the revenue from the NEO cloud buying the GPUs goes back to Nvidia. They invested in more NEO clouds and people have this idea that there is a circular finance financing issue like you saw with some of the vendor financing that happened and revealed in the dot com bubble to be an uh, inflator of that bubble. Now, on their last earnings call, Nvidia addressed this directly and said they see it differently. Do you agree with them or do you think that there are some concerns here?
Speaker A: Yeah, you know, I definitely was the type of person where right away I was like, okay, this is different, this feels funny. I need to dig in and understand it and try to understand both sides. I could definitely see why it looks like circular financing and you could even call it as much like call it what you will. Um, I think going back to thinking through, uh, demand and supply and the cost of capital here and just knowing that it is a fact that there's just insatiable demand, especially Again now with agentic AI, where literally the cost, the barrier to entry for software development has gone as close to zero as possible. I mean, I've got a son who made like a 70,000 line video game this summer and he didn't write those lines by himself, right? He used Codex to do it. And like, that's amazing and it's unreal. And I think, dude, just wait till he's in high school and in college and beyond. Like, he's gonna use AI so much more intelligently than me. He's gonna use way more tokens than me. Like, I can't even believe what, what it's going to be like in the future. And so, you know, you can just look in every industry and see people doing like, lots of people writing software, doing interesting things that they couldn't do before. So, like, the demand totally real, the supply very fixed. Ultimately, at the end of the day, you might look at like someone like a tsmc and there's just only so many wafers that come out, only so much COAS capacity. Um, but even as GPUs get built, so, you know, even if we can increase the amount of GPUs that get built, the question is, you know, who has the capital to buy them? Because today we might be talking, you know, five, $10 million a rack or more. Like, who has that kind of money laying around? And so there is this cost of capital, this financing thing that comes into play where it's like, okay, um, you know, if people want, if customers are just like, I just want inference, I want it as fast as possible, I want as soon as possible. Well, please make it happen. Like, who in the supply chain has the money to invest in standing up all these data centers and running them and getting the inference? And then, you know, oh, it could be neoclouds. Okay, well, can they, do they have access to the capital that they need to, you know, make whatever, um, tens of billions of dollars investment or even a few billion dollars investment? And you know, the answer is a lot of these might be early companies or they were bitcoin miners or whatever, so they might have some access to capital. But, but if you're in video and you're staying on the sidelines and you have all this money and the world's biggest company, you're saying demand is incredible, uh, and supply is what it is, but we've got some supply in the market. But it's now, it's not just building it, it's like getting it powered up, finance stood up of Course it makes sense from their perspective to say, if we can help get this stood up more if we can. The neoclouds, like banks aren't so sure they want to lend to them, but if Microsoft or Google or AWS says I'll be the off taker, so, you know, remember that I'm the, you know, it's not, don't think about the Neo, uh, cloud, think about me when you're lending against it. And if Nvidia could also come in and say like, hey, we want to help make this happen, like, can we put our brand behind it, backstop it, whatever, can we just make this happen? I can see why Nvidia would want to do that. Now, does that mean also that they benefit from it? Of course, totally. It's customers and maybe even if there's some sort of, uh, revenue sharing, it's a new source of revenue for Nvidia.
Speaker B: Well, that's capitalism. They're not going to do it if they don't benefit from it.
Speaker A: Yeah, exactly, exactly. But I think as sort of a techno optimist, if people are like, hey, it's going to feel funny, but there's ways to get more compute stood up faster so that more people around the world, world can do the uh, awesome things that they're trying to do, then I'd say, all right, I can get behind that.
Speaker B: Yeah. I think it ultimately boils down to differences of opinion on the durability of the cash flow that comes from these assets. If you went out and you said, I'm going to build a toll road, you can get a lot of financing for that, you don't have to put a lot of equity and you can get a lot of debt because people know, okay, this road's going to have X number of cars. We know what the traffic patterns are, are you'll collect this amount of money, it's very safe. You can raise lots of money in debt at very low rates for projects like that, even if you're a new company. Yes, obviously there's a big difference between a toll road and a GPU based AI factory, as some people are calling them. But from Nvidia's perspective, and from some of the hyperscaler perspectives, they think that these are fairly safe assets, that in two or three years they'll be completely paid back. Then there's going to be a stream of cash flows coming out of them where even if demand goes down a little or doesn't grow at the same rate, you'll still be able to get value out of them and they're not going to depreciate super quickly. And so I think that just Nvidia has one view and they're willing to put their balance sheet behind it and the hyperscalers feel the same way and not everyone else has the same view. And that's really where the rubber is going to meet the road. Yeah. Well, the flip side though is that you have the existing Neo cloud set who are now very embroiled with Nvidia and they're in lockstep. They rely on them for financing to the extent that customers are going to demand more and different silicon. That creates an opportunity for new Neo clouds who pop up, who can handle and figure out how to make it all work together, who can choose the right silicon for what customers want. Do you think that that's effective counter positioning and we'll see another wave of Neo clouds like we did the first time? Or will the existing set, the core, weave the nebius, the uh, uh, you know, companies like that figure this out and just start using Nvidia plus GROK or AMD plus.
Speaker A: Right, right, yeah, I mean, I think, I think both will exist. But you make a good point, which is the trade off, it's like kind of like golden handcuffs. Like the trade off for a Neo cloud is, hey, they're backstopped by Nvidia, uh, maybe they got some financing, they're obviously getting allocation for GPUs. And so those particular Neo clouds might feel like, hey, if there's different silicon out there that is very competitive, even if it's for like a subset of workloads, we may not feel like we can go buy it and offer it, because what if we don't get as much allocation in the future? Or essentially, you know, you might say like, don't buy the hand that feeds you. Um, so I do think there'll be opportunities for other Neo clouds to come in and say, hey, we've got a bunch of different silicon and maybe we can abstract it and we can run your workloads across it so you don't need to worry about it. So I do think there will be opportunities for someone to come in and counter position.
Speaker B: Yeah, I mean in the last couple of years really there's been two big outcomes we've touched on both in the semiconductor space. Grok, which sold to Nvidia or sold.
Speaker A: Yeah, right.
Speaker B: Uh, and Cerebras, which went public, but neither one really won a hyperscaler before they were able to do this. And in the case of Grok, they Sold to Nvidia. And so you wrote an article about what you think the conditions are for the next trillion dollar chip company and you had four conditions here, um, capable of running trillion plus parameter models, RAC scale chips, which we've touched on, beating an incumbent on a KPI and landing a frontier anchor. So points 1 in 4 like running trillion parameter models and landing a frontier anchor are correlated around the idea that the frontier matters the most. Why do you think that that's the case and is a prerequisite to be the next big breakout company in silicon.
Speaker A: Yeah, uh, very good, very interesting question. I mean ultimately today at the Frontier it's the best models and especially obviously if you can run them at fast enough speeds. My belief is that's where the outsize value will accrue today. Um, yes, there are lots of use cases where you can use older models, smaller models, you don't have to run them as fast. I think that pie will continue to always expand. I just don't think that people will pay a premium for it. So I think if you're a AI accelerator company and you're trying to put as much muscle behind a few arrows as possible, um, you would want to really compete at the frontier. Um, you know where I think again there's going to be today it's software developers that are saying like dude, yes, I will pay not $200 a month, we'll pay tokens, we'll pay thousands dollars a month, tens of thousands of dollars a month if we can get it fast. And if we can get, you know, Claude Fable for example, or the latest OpenAI and of course the frontier will always keep getting better. And so yeah, it just feels like if you're an accelerator startup, um, that's also where m, maybe there'll be the least competition because GPUs for example, can't get there today. And we know that. So if you're aiming at uh, 70 billion llama 3 and you're uh, trying to go after all those workloads that are valuable but they don't necessarily need the highest intelligence, I also think there's going be to a lot of competition there and it could be literally um, old hoppers or old amperes from Nvidia, you know, but, but when you go like on that uh, Pareto frontier curve that we always see, and I'll describe it for people who are just listening, where the X axis is interactivity, which is just how fast are those tokens per second per user and then on the Y axis is throughput like the slower you go, the more tokens you can generate, uh, concurrently. But the faster you go, like really way out there on the far right. Even if you can't serve as many users, that's where the value is accruing today. And it's hard to see a world where that changes.
Speaker B: So that graph gets shown a lot, particularly when Jensen or people from Nvidia talk and they talk about what they're going to be able to do with GROK plus Nvidia. But is tokens per second per user that interactivity KPI still the right one to think about for startups? Or are there changing needs because of power constraints, cost constraints, new workloads, like agentic coding? Like, is it still all about only speed?
Speaker A: It's not all about only speed and it's definitely in my opinion, I like to think about it and compare people, um, at a fixed interactivity, um, for a given unit of power. So we are power constrained. So ultimately if someone, a NEO cloud, if they get access to 100 megawatts, they're going to have to ask themselves, how can I get as much revenue as possible out of this hundred megawatts? Um, they might say, I want to bet on allocating some of my megawatts to really fast tokens because I think we can charge more for it. Um, so if you're going way on the right of the interactivity curve, let's say they're aiming for 800 tokens per second, they're going to want to know, okay, I want fixed interactivity. Let's say I want 800 tokens per second or higher, um, because I feel like I can charge a premium for that and I've only got so many megawatts. So then they're going to ask how many concurrent users can I serve? What is my token throughput? So I do think it's about like token throughput at a fixed interactivity for a normalized by power. Um, but I don't think again, not everyone work.
Speaker B: I didn't hear you say the word cost. And you know, one of the ways you get better interactivity is by using more expensive memory, using sram. Yes, was more expensive, may not always be.
Speaker A: Yeah, yeah.
Speaker B: Uh, so how much does cost play into that equation?
Speaker A: I mean cost. I think that cost plays into it. Uh, obviously if you're, you're in that use case where you're a NEO cloud, you've got 100 megawatts, you're trying to generate as much many tokens at a fixed interactivity that you Can, I mean cost is one way you can get more tokens, which is like, oh, you also have a fixed budget to spend on compute. And so if you buy Nvidia Vera Rubin Rack, the latest and greatest plus nine accompanying Grok LPU racks, you might get uh, really high on that interactivity and it might be pretty good power normalized, but you might have spent half your budget or all of your budget just right there. So I do think cost comes into play that when I'm thinking about the user experience, I'm thinking about interactivity, how many people can be served. But ultimately, if you're a, uh, NEO cloud or any buyer of compute, you're definitely thinking about cost. And again, cost could be like, oh, I could buy, I can get the same performance out of two racks from this person versus 11 racks from that person and therefore for the same fixed cost, what if I could get 11 racks from this new competitor and therefore, uh, five, ten times more the tokens at that interactivity.
Speaker B: Yeah, this kind of buyer thinking is very emblematic of a cloud or hyperscaler who's got a huge instance that they're trying to spread over lots of users. I personally always kind of struggle with holding two ideas in my head at the same time. So on the one hand, all of the sort of initial demand and value has been going to the frontier labs and on the infrastructure side has been being served by a combination of hyperscalers and neo clouds. And then on the other hand, we and many other people believe that AI is going to be something as big as the Internet. It's going to diffuse into businesses all over the world. Every company is going to have some AI element. It's not just going to, you know, in the same way that every company has a website now.
Speaker A: Yes, yes. Oh yes.
Speaker B: There's no more any.com versions of companies. Everyone's got a website, everyone's got an app. Soon everyone will have some AI element in their business. And so if you end up in that end state world today in the Internet world, more workloads exist on prem than in the cloud. If I am a, uh, coffee shop, I might have a server in my coffee shop. I'm not going to have an AWS account, most likely, yes. And so ultimately do we end up where enterprise is actually the big market here instead of cloud?
Speaker A: So yeah, that is a very good question. I think they're both going to be massive. And let's talk enterprise because I don't think people appreciate that enough. And you can even look even Nvidia is trying to get ahead of it where they change their reporting and their business units to essentially like for data center, it's hyperscaler and non hyperscaler. Like it's like ACIE or something. Something, you know, it's like enterpriser and AI clouds, which, it gets a little fuzzy there. But, um, and they're saying that, oh, by that way, that non hyperscaler one's growing faster than the hyperscaler one. And right now the revenue is pretty close on both. And again, it's a little fuzzy because they put NEO clouds under there. But um, so the question is like, well, what workloads are going to go to the uh, enterprise and why? And I think that you're totally right, which is there is going to be the diffusion of generative AI across every industry. And I definitely don't think we're there yet. Um, and when you think about the implications of that and the incentives for enterprises, I don't think that they are going to say to the point that you made, um, every company, you should be a whatever company in its own industry, logistics, manufacturing, healthcare, whatever. And then Marc Andreessen said like 15 years ago, no software, uh, is in the world. Every company's gonna be a software company. And to some extent that's right, because even if it's just internal tools now, all these logistics and manufacturing, healthcare, they're all using software. And I think we're going to a world, you know, where uh, agentic AI is going to eat the world and every company is going to be an agentic AI company. And like, like I said, you know, pointing at my son, like, just imagine when, uh, you know, fast forward 15 years, like, yes, they're all going to be agentic AI. Okay, so if agentic AI is core to how businesses run, are they going to all have, you know, hundreds of millions of dollars that they spend in tokens every year? Totally not. I think there's going to be all sorts of reasons. One cost to, you know, uh, owning your own data, figuring out how you even differentiate in a world like that. There's going to be all these incentives for companies to want to deploy workloads on premises. And it's not going to be today. It's a lot of like, oh, we'll go to the cloud for the frontier workloads and let's do as much as we can of the older, smaller models on premise. Because by the way, um, you could also look at Nvidia's proxy for hyperscalers and non hyperscalers as Frontier closed models and open source models. Because if you're running enterprise AI today, it's really meaning you're. If you're running it locally, on premises or whatever in a server or on your desktop, um, it's got to be an open source model. And so I think that's a little bit why we're in the world we are today, which is like, oh, if you want the best model, you have to go to the cloud. And if you can do anything else, you should. Do you want to pay for tokens or token generators? I think a lot of people, if they can afford it, would rather have token generators. Um, so. Okay, but if we fast forward a little bit, what happens is if front. What happens if open source frontier models can keep up or be as good or good enough? I do think they'll continue to be a rise in the amount of workloads that you do, um, on premises. And yeah, there's all sorts of reasons. Of course you can look at like regulated industries and say, like, well, they're going to have to run that on premises. But maybe even if I'm not regulated, like, maybe in the future I think that like, hey, the data that, like the labeled data that my humans generate, where it's like we have all this agentic stuff and then let's um, say I'm an insurance company, we have all this agentic stuff that's doing the claims processing and then my humans are going in and correcting it. That was good. That was good. That was wrong. Let's keep that data internally and let's fine tune our own model so that it gets it right in the future. So maybe I want to run that locally because I'm in charge of the model, I'm in charge of the data, I keep it. It's all now my intellectual property. And maybe I. See, that's how I differentiate in the future is like, I've got better agents than my other insurance competitor. And uh, you know. Yes. Can you do this, all this stuff in the cloud and feel like it's secure? You totally can, but you do lose. Um, like I think at the end of the day when we're talking about diffusion, we want every engineer at every company to be able to tinker and touch it and play with it and, you know, use it themselves. And sometimes when stuff's in the cloud, you get the convenience, but you lose the ability to maybe like get in the hood, whether it's a closed model or even an open model in the cloud. And so, uh, it's Funny you say that.
Speaker B: I almost feel the opposite way about enterprise, which is that to me it seems like it's gated a little bit by software. Where so many of these companies have not moved workloads even to the cloud for data sovereignty, regulatory, uh, privacy, IP protection reasons. They're very concerned about stuff leaving their corporate premises, their IT premises, and they would gladly do more things in AI, but they don't have engineers in house who know how to post, train a model or know how to tinker with this stuff. And those people are expensive and they don't want to hire them. And so they want this thing and they want to do it on prem, um, in enterprise. But no one's quite figured out how to help them do that thing yet. And that to me feels like the missing piece that would unlock a lot of enterprise hardware sales as well.
Speaker A: Yeah, yeah. And I definitely agree with you. That's where we are today. Um, where now. So, you know, back up eight years, everyone wanted to be a software company, but they didn't have software engineers. You know, I live in Iowa, and so it's like if you're a software engineer there and you're willing to work in insurance or ag or retail, like you were a rock star because you could walk in and they're like, yes, thank you, we need you. We didn't have this capability before. Fast forward now and anyone can vibe code, which is actually pretty awesome because now these domain experts who were the person in insurance that knows insurance really well, they can actually build the solution they want. Now I think we're where we were like eight years ago, where everyone's like, okay, but now we don't know how to fine tune a. Like, I'm an insurance expert and I can vibe code a thing, and I've got some software people here, but none of us know how to fine tune yet. And I do think that's like the education piece, you know, if I have to tell my children, like, hey, what if they had to go to college like today and they had to pick a major? I'd just be like, pick anything and pick machine learning and learn how to fine tune stuff. Because you can go in and you can understand domain and you can also understand how to like fine tune AI and so essentially actually apply AI. And so I think that there'll probably be a rise of AI engineers, if you will. Um, maybe agents will do this for you and bring that down to zero faster than agents brought software down to zero, which maybe took 40 years. But I do think, um, it is a pain point today that companies don't have the AI generative AI familiarity yet. I don't think that pain point will be there forever. Maybe it's five years, maybe it's more, I don't know. But I don't think that will always be a blocker. I think eventually if every company became a software company and they have software literate people on staff, uh, I think eventually everyone will have like, you know, fine tuning LLM literate people on staff too.
Speaker B: So when that day comes, is this just a huge unlock for Nvidia and they get that much more revenue or do you think anyone else has a chance at that market?
Speaker A: That's a great question. Like what happened to IBM? Um, you know, like I think that ultimately there are giants and they're first and they ride a huge wave and then to all the points we've talked about in the past of like zooming out and there's, you know, three or four winners and people want competition. I think the more people can tinker, the more that this diffuses, just the more opportunity, like one company cannot meet everyone's needs at the right price point, at the right speed or whatever, um, they can meet lots of people needs but there's always going to be people who are trying to do some interesting bespoke thing and they're going to say the off the shelf stuff uses too much power. I know I've got this crazy setup, but I can only, I can't do 130 kilowatts, I can only do 50 kilowatts or something. Right. And there's always, I think, going to be workloads that are emerging where the, the stuff off the shelf just doesn't meet it. And the tough part is like when you're Nvidia and you've got these huge hyperscalers, like you're not necessarily incentivized to go find those little people. Like you're not invest, you're not interested in picking up pennies, you're interested in picking up a billion dollar bills, you know, um, so I don't think that it's always, I think Nvidia will be totally fine and they have great solutions and they're always going to have customers who are coming to them to get the latest and the greatest and to deploy it quickly. But I do think there will continue to be new opportunities that people pop up, especially as this diffuses where people can compete.
Speaker B: Circling back to this idea of the next trillion dollar company, there's an explosion of opportunities in different workloads and then also we're hearing different markets. Does that mean that there's an opportunity for a trillion dollar company or maybe are we going to get $500 billion? Companies like AMD is still not even a trillion dollar company and they've been around for a really long time. They've got a lot of pieces of this puzzle, right?
Speaker A: Yeah, I mean I just think the size of the market, you know, especially if you could if it. So here's the deal and part of why I sort of came to that conclusion, GPUs, uh, uh, obviously they have a history in graphics and being able to do things in parallel. And that has been changed. Shifting toward AI centric Nvidia saying yes, these data center GPUs, you're not going to run doom on them. We're going to, yes, they used to be able to support FP64 and they still do, but we're going to spend all of our transistors as we do a node shrink on like um, FP4 FPA. So like these, this lower precision that um, AI models really want. Um, but at the same time it was still a general purpose cpu. And that's why I said we went from training with GPUs to inference with GPUs to, then we went to this next era where it was like, hey, LLMs are the workload. We haven't actually had silicon designed specifically for LLMs, we had GPUs and they kind of morphed from their early roots to fit the shape of what we're doing. So let's take a GPU and slap on this SRAM thing. But no one yet has really brought a, um, chip to market that was the designed specifically for LLMs. So then the question is, okay, if you can be the first one that can stand up a uh, gigawatt's worth, which not simple, you'd have to, you know.
Speaker B: Well, so to some degree the Nvidia GPUs of the last couple cycles are chips that were designed specifically for LLMs. It's not like you can take uh, Grace Blackwell and you know, play a video game on it easily. It's highly specialized for this type of workload. And in particular when you get into these combination gpu, LPU or the AMD helios, these are really designed specifically for LLM workloads.
Speaker A: They are morphed specifically for LLM workloads. So they have legacy ways of doing the networking, legacy ways of thinking about the memory hierarchy. I definitely agree that they are iterating toward what is best for the inference workloads that they serve. But the question is, what if you started with the blank sheet? And we actually even saw this from OpenAI with their jalapeno chip that they um, launched at Hot Chips recently where they said, yeah, we started with a blank design and we are thinking very differently about it. We're making very different architectural decisions. Um, they said, hey, instead of shipping all this KV cash around and having all this shared memory and all this contention, like we're like, guys, the data is never in the right place when we want it and our computer is always sitting around. What if every accelerator had its own little HBM slice? And what if we map the workload and try to say like, how could we rethink about the workload such that we don't have all this contention and shipping data around? And I use that example and I think tensordyne is another example where they said, hey, should it be matrix multiplication or could we do log math and would that turn multiplies into ads which are really fast in silicon. So I actually do think there's architectural knobs that when you're taking a uh, GPU that has support like Nvidia's next GPU still has to have backward support for all the software. It doesn't exactly. But to some extent like they want to support all the workloads. If it ran on Hopper, they want it to mostly run on Blackwell. But what if you could start to the clean sheet, make different architectural decisions, whether it's about the way you do the compute, the memory hierarchy, the way you network, um, etched said, hey, let's do low voltage inference. Like what would happen if we ran this instead of at these, you know, really, um, high power and we just put a lot of oomph so it can go really fast. Like what are the benefits that we get if we run at a lower voltage? And um, if you said Austin, that power is fixed, like maybe there's benefits even if there's trade offs, like it doesn't run as fast, but what if it's like a lot significantly lower power? And so I do think that there's opportunity to make clean uh, sheet designs, make different architectural decisions and therefore unlock that KPI like I talked about, like what if you could get 10x more tokens for that fixed interactivity for that particular model out of your hundred megawatts than you could buying something from Nvidia or AMD or something, something off the
Speaker B: shelf you hit on another. I, uh, think Interesting tension. So on the one hand you have someone like an OpenAI who is a large customer of Nvidia, but also now is designing their own chip. And increasingly if you hear them talk about how they did that design, it was very, or at least they claim it was very AI optimized and that they were able to do it much quicker because of the acceleration of AI. And then you also have companies like a tensordyne who's got an entirely new idea paradigm, like a whole new way of approaching this problem that probably a customer would not have thought of on their own because they're a little bit more focused on their specific workload, not new ideas from the beginning. So how do you think about the trade off of how much is going to go to custom silicon companies using AI to create something specifically for what they need and then maybe working with a Broadcom or Marvell to help them finish and take it out and get it into production versus these companies that have entirely new approaches like a Tensor Dyn.
Speaker A: Yeah, yeah. Uh, you know it's so interesting to think about OpenAI and the fact that they're buying compute from, you know, the big vendors, mostly GPUs, um, and then they're also building their own silicon which they have the advantage of their chip designers working hand in hand with their software team and co designing for specific workloads. To your point, yes, um, like off the shelf silicon vendors, they are not inventing in a vacuum. They are talking with their biggest customers and they're saying like, hey, where do you see your roadmap going? How can we make sure that our silicon meets your needs? But that is different than OpenEye internally having their ML team and their um, chip team working very closely together in co designing. Um, but to your point, obviously they, they are going to land, they are going to make a particular set of trade offs. Like everything in engineering is all about trade offs. Do you want more HBM or more sram? Well it's going to cost you something either way. You know, do you want some die size or do you want to uh, you stack it even higher and you know, maybe there's thermal trade offs, whatever, everything. And so to think that uh, you know, the merchant vendors made particular sets of trade offs which of course they need to sell their chips to as many customers as possible, even if it's maybe only a handful these days, like they have to make a set of trade offs, then these internal teams, they're going to make a set of trade offs given what they know about the workloads that they're running. But to think that just those two different sets of trade offs will be all that you need. Right. I do think there's opportunity for someone like a Tensor dynamic or others to say, hey, what if there's these like totally crazy trade offs and we actually are taking the risk in doing the R and D on that trade off. So this log math stuff like yeah, surely people at the merchants looking vendors or the internal XPU teams, that's crossed their mind, they've seen a paper but they may not be incentivized to take the risk on that.
Speaker B: Right.
Speaker A: Like if you're OpenAI's XPU team, you're making your first chip, are you gonna um, play around with log math? Are you gonna say no, no, no, let's pull, let's make some of these very interesting other decisions like HBM slices that have like um, maybe been used in industry elsewhere. Yeah.
Speaker B: People forget that there's like people involved in these decisions who have career risks.
Speaker A: Exactly. Totally.
Speaker B: If you can make a decision that's very high probability and still works.
Speaker A: Yes.
Speaker B: But Maybe isn't the 10x that's probably better for you. If you're working at a big company and particularly you're trying to deliver your first version of something and you don't want to screw up.
Speaker A: Yeah. And it's expensive too. Like you're going to tape it out and you might stand up and you might be all in a billion dollars or something. Like you don't want to get that wrong or have it get canceled before you can stand it up. Um, so yes, there's totally this whole human side, there's these incentives. And so again that continues to be opportunities for startups to say like we're going to take that risk. Or actually we've been, in tensordyne's case, we've been taking that risk and we were trying the log math in a different market and now we're ready to bring it to this market, you know, and so could uh, an OpenAI say that's super interesting. Like we'll take a couple racks of those too. Absolutely. Because I do know that these are very sophisticated buyers and they're always seeing what else is out there because they completely understand that when, uh, designs were made at a particular point in time, it may or may not be fitting into exactly what they need A couple
Speaker B: years from now are more companies looking at these custom chips. OpenAI's got a lot of money, they've got a lot of really talented engineers. Even without AI they probably could have designed their own chip. But as AI makes designing a chip easier, it's not at the level of your son being able to create a video game using Codex. But you talk to teams and they are seeing lots of improvements. There's some sort of floor to how cheap it can get because you do ultimately need to tape it out and do all of these things in the physical world. But do you think that you're going to see a big proliferation of companies that never would have tried to do custom silicon, give it a shot?
Speaker A: I definitely think so. And you can already look at certain examples and see where it's happening. So the question is, why would you design your own silicon? Well, if you have uh, if you know your workload really well, probably what you bought, just like we talked about, probably what you bought off the shelf, design decisions were made and it might not map to your workload perfectly. And that might be okay at first. But like take Rivian for example. Um, they used to use, uh, they actually went through a couple different vendors, but they used some merchants, token vendors and they would map their workload to it and it was fine enough. But if you're running and electric vehicle and you're also trying to do autonomous driving, um, you have very specific needs. You want to use as little power as possible because otherwise you're taking battery away from the customer. Being able to drive another couple miles, whatever. On the other hand, you need to run as fast as possible because if it takes you too long to make a decision, there's another 20 meters that you were thinking before you started braking. Um, but at the same time now all of a sudden they've got LIDAR and they've got cameras and you have all this data flowing around that you need to do inference on as fast as possible. And by the way, it used to be convolutional neural networks that they did the inference on. Now they're doing end to end, um, LLM based, uh, vision language action models is what they call it. And so the workload has been changing and now in this world of like we need to do real time stuff, we need to do these heavy big models, um, but we also need to take power into control. They're saying, well, uh, and by the way, we need to ship a ton of data around so we need to have a particular interconnect bandwidth, um, and memory capacity. Memory bandwidth. They're saying, you're like ah man, this stuff off the shelf, this doesn't really fit our needs and our cost profile and so then the answer could be, well, design your own chip. Well, there's a cost to that. You need engineers, uh, who are familiar with front end design, back end, maybe you can partner testing. But then at the end of the day there is a cost to, you know, um, that might take you three years of development. So you're gonna have to pay for these engineers for many years and you're gonna have a whole roadmap and it's a very big investment. And then of course there's the cost to tape it out and actually get it built. Well, okay, fine, but what if with the help of AI, um, maybe it still takes you 100 people, but instead of taking three years, it takes one year. And so maybe your cost is cut, uh, down by a third. Where when you were running that calculation, maybe even though the performance and the headroom it would give you to do interesting things was just so good, but you're just like your CFO is like, we just can't add another $2,000 to the bill of materials. We just can't do that yet. Well maybe if you can come back and say like it's only gonna be $700 to the bill of materials, maybe they'd say okay, that's really interesting. We think we could hack it, right? So I think that adding AI to chip design, speeding up the time to market, being able to do more with the same amount of people, I think will ultimately be net good for companies and maybe like reduce that cost barrier to entry. Or again, it could be a talent thing where you're like, how do I go find 100 people? Maybe you only need 50 or something. Like I just think it will reduce barriers to entry and we will see all sorts of use cases where people maybe had never thought about making their own chip and now they'll say like, oh, it's actually something we could do.
Speaker B: It's an interesting dynamic because the same speed up that would be available to these companies is also then available to the merchant silicon teams who are probably even better served to use these tools because an experienced engineer who understands trade offs is going to use the tool tool a little bit better than a person who's approaching it for the first time or this is their first chip design and so maybe they can start proliferating their number of SKUs and serve more customers with more custom, semi custom things within what is somewhat merchant. Um, so I just think it's really interesting. I don't know how it's going to work.
Speaker A: Yeah, I would expect, I would hope that um, merchant silicon companies. Because the thing is, when you're a merchant silicon company and you're going to make a particular product, there has to be to be a big enough market for you to capture enough customers where it was worth your time and investment. Right. And there's probably a point where the ROI didn't make sense because the market's maybe only this big, and maybe it's going to cost you this much to develop the sku and you're like, ah, that's not worth it. But now if the cost can come down by half, maybe it is worth it and it clears your internal rate of return that you needed. So I do. I would expect that, um, big companies would be able to create more innovations as well.
Speaker B: Awesome. Well, I don't think anyone really knows how it's going to play out. This is such an exciting time for us. Austin, thank you so much. It's been an incredible conversation. I've really enjoyed it.
Speaker A: Yes, thank you. This was fun. Uh, let's do it again.
Speaker B: Absolutely. What sticks with me after talking with Austin is how much is still unsettled and also how little that appears to matter to the pace of spending. Will the market coalesce around a chip designed specifically for large language models? Can a cloud company build a durable business selling to the same hyperscalers they compete with? For better or worse, no one can afford to wait to find out the answers before investing billions just to stay in the race. Thank you for tuning in to the Tech Surge Podcast podcast from Celeste Capital. If you enjoyed this episode, please feel free to share it, subscribe or leave us a review on your favorite podcast platform. We'll be back every two weeks with more discussions of all things deep tech. Bye for now.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.