
Training Data · 2026-06-30 · 1h 10m
Key moments - from our scoring
Substance score
75 / 100
Five dimensions, 20 points each
Dylan Patel, founder of SemiAnalysis, describes how he built the premier semiconductor research firm from anonymous forum posting to a 90-person operation generating over $100M in revenue. Starting as a teenage hardware enthusiast moderating forums on Intel, Nvidia, and AMD, Patel combines deep technical knowledge of chip design, manufacturing, and supply chain economics with institutional-grade research. He discusses how post-COVID isolation and personal tragedy pushed him to formalize his expertise into SemiAnalysis blogs, which attracted consulting clients and led to conference attendance across the semiconductor ecosystem - attending 40+ events yearly from SPIE lithography conferences to NeurIPS. The conversation covers his most significant contribution: InferenceX, a living benchmark platform tracking real-time inference performance across hardware (Nvidia, AMD, Google TPUs, AWS Trainium), software stacks (SG-Lang, vLLM), and models (OpenAI, Anthropic, Chinese labs) on a Pareto optimal curve showing the latency-versus-throughput tradeoff. Patel argues this throughput-latency curve is foundational to all AI infrastructure decisions, with inference becoming a multi-percentage-point portion of global GDP.
InferenceX is a living, automated benchmarking platform that continuously tests AI inference performance across different hardware (Nvidia, AMD, TPUs, Trainium), software stacks (SG-Lang, vLLM), and models, updated daily. Patel built it because traditional point-in-time benchmarks become outdated within weeks as models, drivers, and optimization libraries update multiple times per week, making it impossible to maintain an accurate performance picture without continuous benchmarking.
The throughput-latency curve maps the tradeoff between response speed (latency) and number of concurrent users served (batch size). Patel believes this curve is the most important because it determines whether an application needs instant responses (low batch size, high cost per user) or can batch process without latency concerns (high throughput, lower per-unit cost), and all hardware, model, and application layer decisions flow downstream from this tradeoff.
After being doxxed on anonymous accounts in mid-2020 during COVID lockdowns, Patel started a blog called SemiAnalysis on his 24th birthday, publishing two high-effort posts instead of anonymous shitposting. The public posts under his real name generated significant traction and consulting business; he then traveled to 40+ industry conferences yearly while living nomadically (in national parks, cheap motels) to deepen his supply chain expertise before scaling the operation.
SemiAnalysis is a 90-person research firm combining engineers and technologists across the semiconductor supply chain with former hedge fund analysts. They produce institutional research subscriptions and high-impact public analyses (like GPU benchmarking and InferenceX) that educate the market on chip design, manufacturing, economics, and AI infrastructure - balancing technical depth with financial rigor.
Patel attended niche supply-chain conferences because they reveal ground-truth information about manufacturing bottlenecks, chemical suppliers, pricing, and logistics that aren't published anywhere. At these events, he learned critical facts like how a factory fire in the 1980s doubled memory prices, or that only three companies globally produce certain chemicals, giving him unique insight into semiconductor vulnerabilities.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode is genuinely dense with non-obvious technical and economic claims - the 100x co-design thesis, the throughput-interactivity curve as the master curve, power density breakthroughs beyond 1W/mm², and Jensen's multipolar world strategy are all substantive. The long origin story and some filler in the opening drag the average down.
you take what could have been a 2x here, 2x here and instead of being multiplicative to 8x it's actually 100x because you've optimized it across all three layers
model costs drop for equivalent quality by like 60x a year
Several genuinely fresh frameworks: the reframing of the CUDA moat as an ecosystem co-optimization problem rather than a developer tooling moat, the local-vs-global minima framing for ASIC bets, and the Jensen multipolar world thesis are all counterintuitive and non-recycled takes. These are not the standard semiconductor talking points.
what people call the CUDA moat is not actually anything to do with Cuda, but it's like the fact that Deepseek, Kimi and Zippuai and Alibaba and Tencent...their models are a co design for GPUs and therefore if I want to run them on GPUs actually in some cases they don't run really well on tpus
A world where open, anthropic and Google models are the only models, is one in which he's screwed
Dylan Patel is the genuine article - he built SemiAnalysis to 90 people and reportedly $100M revenue by attending 40+ conferences a year, cultivating primary supply-chain sources, and producing original research that moves industry. He has proprietary data, named contacts, and is clearly plugged into confidential deal structures. This is an actual practitioner-researcher, not a thought leader.
we have 90 people and like a big chunk of them are technologists, engineers across the whole supply chain
I go to 40 plus conferences a year
Extremely specific throughout: named rental rates per gigawatt for Trainium vs GPU vs SpaceX deals, Anthropic's Q2 profitability status and per-token margins, data center pricing trajectories in $/kW/month, and chip-level architectural details like NV Link connecting 72 GPUs vs Google ICI connecting 8,000. The level of named data points is rare for a podcast.
Trainium uh sells at sub $10 billion per gigawatt rental rate uh to anthropic and to OpenAI GPUs at least before the craziness of the last 6 months usually went around 12 to $13 billion per gigawatt
Anthropic in Q2 is profitable, their net income profitable, um, excluding stock based compensation...their per token margin is so high
Sean's best moment is a genuinely setup-for-disagreement question on where efficiency gains originate, which Dylan immediately contests and uses to deliver the co-design thesis - that's skilled hosting. The oil/Saudi Arabia analogy to probe data center quality differentiation is also creative. However, the origin story runs very long without redirection, and several questions are framed as agree/disagree softballs.
Sean, I completely disagree with you by the way
To me it seems like in the last three years most of the games have come from hardware level and some from the model level. Like do you think that that uh, is what, do you agree with that?
Computed from the transcript - who did the talking, and the words that came up most.
Dylan Patel, founder of SemiAnalysis, argues the biggest gains in AI don't come from faster chips, they come from software-hardware co-design. Optimizing the model, the kernels, and the silicon together turns a 2x here and a 2x there into 100x. He explains why DeepSeek's experts were shaped for Nvidia's Hopper (and why TPUs struggle to run it), why OpenAI's sparser models and Anthropic's denser ones pull them toward different hardware, and why the so-called CUDA moat was never really about CUDA. Dylan breaks down InferenceX, his living benchmark that runs the latest models on over $50M of donated hardware daily, tracking a roughly 60x annual drop in cost per unit of quality. He makes the case that inference will be a bigger market than oil, that the compute crunch persists because models expand the value of useful work faster than compute grows, and why Jensen Huang is bankrolling neoclouds to engineer a multipolar world. Hosted by Shaun Maguire and Sonya Huang, Sequoia Capital
Transcribed and scored by The B2B Podcast Index.
Speaker A: I think it's really fun inside of semiannalysis because we have 90 people and like a big chunk of them are technologists, engineers across the whole supply chain. Um, and then a big chunk is people who are formerly at hedge funds. And you see these arguments, like, people are like, oh, well, that doesn't matter. And it's like, then someone's like, well, but cost. And then someone, the engineer is like, no, no, no. But this technology is the coolest. And you see this, you see this organically, like, fight it out. Um, and we're pretty informal. And you know, given the fact that I was a forum moderator is. You can imagine what, you're enjoying it. You don't wrestle with the. Because the pig enjoys it. Right, Exactly.
Speaker B: We're here in the semi, uh, analysis office with Dylan Patel. You know, I'm Sean from Sequoia, my partner, Sonia Huang. Pretty insane what you've done. Semi. Semis 5 years ago were not very sexy in the West. They were sex in the East. But people, uh, here in the west had kind of forgotten about them. You did not forget about them, though. You went very long. You created probably the premier research company in the space that's been educating the world and state of the art from very technical details to supply chain to the bigger picture. Um, there's rumors that Semi analysis recently passed 100 million of revenue. I don't know how accurate those are. Whatever the numbers are, you can. Guys are crushing.
Speaker A: It's. It's as accurate as the information is. You know, you know, you never, you never know.
Speaker B: Uh, there's also rumors that you might start a venture fund. Like, you know, I, I hear all the time in the ecosystem people wanting, you know, affiliation with Semi Analysis. You, you've built this trusted brand and so whatever you do, it's working this clearly like just the beginning of the journey for you. Congratulations. All of that. But how did this happen? Like, how did you. First question is like, what is the background? How did you kind of get to where you are now?
Speaker A: Well, when I was a young boy and you know, coming out of the womb.
Speaker C: So.
Speaker A: So, okay, so I grew up in like a small business. My parents had a motel. We lived in the motel. We. At our gas station. So, you know, uh, I was selling. You know, I joke a lot of times. The first neural network I trained was, uh, racially and visually profiling people based on when they enter the gas station, which cigarette to, uh, pick.
Speaker B: Oh my.
Speaker A: Basically, you know, the cigarettes were all extruded across the top and I was too Short to actually, like, you know, reach them. And technically it wasn't legal to sell cigarettes at that age, but whatever. I had to move the step stool over to the right area.
Speaker B: I started working. My first job is before it was
Speaker A: legal, too, but it's good experience. Well, I didn't get paid.
Speaker B: Right.
Speaker A: It's a family business. Same, yeah. We had a motel, and then across the street was our gas station. So, you know, sometimes, you know, you know, someone would walk in and so, like, if an old white lady with curly hair walked in, I'd move the ladder or the step stool over to where the camels are. And if, you know, different, different age, demographic, uh, profession, you know, race, etc. I move the step stool over and I joke. This is the first neural network I trained because if I waited for them to tell me, I'd have to move it over and then I'd step up versus just being ready. Um, so menthols versus 100 slims and all these things. I joke. That's the first neural network I trained. But I grew up in family businesses, lived, um, in a motel. And, um, it all really goes back to when I was my eighth birthday. Um, my birthday's in May M. And it was April when the Xbox 360 was announced. Um, for my birthday, I didn't ask for the Xbox or I didn't ask for a birthday gift. My parents asked what I wanted. I asked for it for Christmas. We, uh, celebrated Christmas, but there was no way, at least at the time I thought there was no way they would ask, would give me the Xbox 360 for Christmas. And so I got it for. I asked for my birthday for tab for Christmas. Anyways, Christmas comes around. I get it. Um, you know, fast forward a couple months. My cousin, who lives in Alabama, they also lived in a motel, was going to come over for spring break, um, for his spring break. And we were going to hang out at my house. And he's in between me and my older brother in age. Brother's a bit more jockey. Um, so he didn't really care too much about the Xbox. He played sometimes, so didn't really care. Um, but my cousin, I wanted him to think I was cool, right? So I bragged many times on the phone. I was like, yeah, I got an Xbox. And then the Xbox broke. There was something. There's a hardware defect called the red ring of death. Um, but long story short, I had to open it up and short the temperature sensor and it fixed it. Um, but there was many other tricks I tried first and none of them worked. Um, and so that's sort of how I got into hardware is like open Pandora's box. By the time I was 12, I was on these forums a lot, reading, posting, uh, a lot. And this is around the time when Reddit ate all other forms. And so I became a moderator of Android and Apple and Google as well as hardware and was watching, looking at Intel, Nvidia and AMD and all these other forms. So build a PC. All these forms I was watching, reading, posting a lot, but some of them I was moderating a lot. Um, and so smartphones, watching smartphones develop from very simple to speed racing to being technologically more advanced than PCs, um, in many ways architecturally and same with all the in gpus. Like just tracking and watching that, reading every comment. Always, um, having the economic tinge because I grew up in a small business, so I was always looking at the economics. Right. There was a time where all the, like, I'd say neck beards on the Internet loved AMD GPUs. And like, I personally had bought an AMD GPU too because price performance. But then when it came down to like, what's technically better, I'd always be like, no, no, no, Nvidia is better because they use a smaller chip to get, you know, better performance at better power efficiencies and their margins better. And so like, I would always like talk about how Nvidia's margins were better than AMD's in the GPU landscape. And so it was like very fun.
Speaker B: And you were 12 at the time?
Speaker A: I started moderating when I was 12, but this is all through my teenage tween age and high school years. Right.
Speaker C: Do, um, you have any other weird hobbies or was it just semis?
Speaker A: I played a ton of Starcraft at one point. I was grandmaster on the North American ladder. StarCraft 2.
Speaker B: Serious.
Speaker C: So you've gotten just obsessively good at multiple things?
Speaker A: Yeah, I mean it's, it's obsession is.
Speaker B: How were your grades?
Speaker A: Um, they were decent. Um, I would say like I had mostly A's, but they're classes that I like were thought were really boring or, you know, I just didn't enjoy like Spanish. I got like, not the greatest grades, um, you know, but. But it was like I speak fluent Spanish by the way, so it's really dumb. Um, but like, it's just sort of.
Speaker B: Maybe that's why you didn't get a good grade.
Speaker A: No, I didn't. I didn't learn later to be fair, but yeah, that. So it was sort of. My grades are fine. Right? Like, I mean, they were fine enough for Asian parents. I was better than most of school. But you know, it wasn't like, you know, try hard maxing for like, you know, all A's.
Speaker C: Okay, so you're very much a student of the Internet then. This is how you, how you develop this expertise. Uh, at what point do you decide to start semi analysis? And what's been the biggest surprise since starting the company?
Speaker A: Yeah, so I went to school, I got a few degrees in stuff that wasn't related to semiconductors. Um, was a quant for two years at a small quant risk firm. Um, and then basically there was a culmination of events that happened. One was that my, um, sort of like I got screwed out of a bonus. I'd made my company many millions of revenue of risk free revenue because I exploited like a risk thing in the market. Um, I think well over 10 million. And then someone else took credit for my work and all this sort of stuff. But eventually I did get right size, but I lost a social contract with the company I was working with. Um, add some. My grandparents grew up in my house with us or in the motel with us. They, uh, lived with us and so we were very close with them. And my grandmother got dementia and she forgot who I was and she, she fell down some stairs and had like a tragic accident and passed away. So all of that happened in early 2020. Um, additionally, there were some like, you know, girl things. And so, you know, there's a few things that happened that made me like kind of very sad. Um, and so all of those things sort of culminated. Then Covid happened and my brother's like, dude, just come stay with me. He lived in Nashville, so I came and stayed with him in Nashville. We were like, oh, lockdowns will be a few weeks. You can stay with me while they happen and then you can go back home and you know, whatever. Famous last words. Lockdowns lasted much longer. But, you know, living with my brother for a few months, you know, it was like sort of like, okay, didn't know what I was doing. I was now at my brother's home. Um, everything was his rules, you know, sort of like, you know, him and him and his fiancee at the time, now wife, you know, were like there. And so like, I basically had to tiptoe around, but I didn't care about my job. And so I was like posting even more than normal. I'd always been posting a lot on the Internet. I'D always been trading stocks a lot, but like, I made a lot of money shorting Covid and long and Covid and like all this stuff. Semiconductor shortages happened around then too. And anyways, I was like very much obsessed with posting and things like that. And eventually, um, around that time, someone again for an argument, someone on the Internet. And they doxed me, right? They publicly revealed my identity for my anonymous account. And at the time I was like, oh, no, I was scared. I stopped posting for like three weeks. And I was like, what am I doing? Why do I care? So then I just started posting under. I had like blogs and stuff as well. I made a real blog, semiannalysis. And on my 24th birthday I posted, you know, two blogs. And then from there it just like, it was not a newsletter. But I got so much traction because now instead of posting on anonymous name as a real name. And I put a lot more effort into those two posts than I usually did. Instead of like shit posting on the Internet, it was like real effort into the blog. Um, you can actually go back and read those if you want. They're not that great, but you know, they were, they were good for the time. They were the best stuff you could find on the Internet about semis. Um, and I just kept posting, posting, posting. I started getting a lot of consulting business, you know, 2020. I also sort of. I was again crashing out. Didn't know what I wanted to do. So I packed, uh, everything up or sort of. I took my truck, I bought a tent that fits on the back of the tent truck, um, bought an air mattress, whatever, and would like drove around all these national parks all around America. And so like two or three or four days of the week, I'd stay in a random motel where I negotiated the price to be like $30 a night for a room. And I was working on something else's stuff. And then the weekends I'd read books and oftentimes read textbooks, um, while in some random national park or hiking and listen to audiobooks, um, about semiconductors, about AI, about all the things that I cared a lot about and got way more educated over these six months where I'm just like going to every national park. Um, and the whole time I was alone, that I was alone the whole time I was posting blogs, everyone was like, dylan, what the are you doing?
Speaker B: Pre Starlink or the very early days of Starlink.
Speaker A: Pre Starlink. Pre Starlink. Um, yeah, so it was like very much like, what are you doing? Um, I travel around Latam again, like for For a year, initially with my friend and then with my ex, you know, for, you know, about a year. And then I go then 22, 23, 24, end of 21, 22, 23 and 24. I'm completely, I'm still completely homeless since mid-2020. Right. Um, but I'm traveling around to every conference in the world. I go to 40 plus conferences a year. No matter where in the supply chain it is. I'm like, oh, that looks interesting. I guess I'll go to that. And I'm like, I went to one conference like, wow, this is amazing. You get to talk to the experts and they just like, they, they, they, they're going to talk to you because. And then you're so excited. And in the case of semiconductors, everyone's a boomer, so it's like, it's great to like, you know, they're like, they don't see young people who are like excited about it, so they're really happy to tell stuff.
Speaker B: And so you just didn't have to ask on this. Was there like a part of the supply chain or one of these conferences that you know, particularly change your view of the semi world or that you felt then or feel now is particularly underrated?
Speaker C: I think.
Speaker A: I think the trade shows like in conferences range really widely. Um, obviously some of the, you know, the ones I have the most fun at, you know, include Neurips. Why?
Speaker B: Why is that?
Speaker A: Because it's 20,000 AI researchers and they're generally in my distribution of age range. So it's like a lot of fun. But they're also like leading AI researchers and it's a lot of fun and you learn a lot. Um, there's also a lot of parties and then it ranges all the way. Like, you know, there's Random Chemical conference in Japan where it's 300 Japanese dudes, it's like 20 guys from ASML, 20 guys from TSMC, 20 guys from Intel. And those are the only people who speak English. Everyone else speaks only Japanese. And you're like, uh, I guess there's still pretty interesting and fun, I think. I think like one thing that I have like a skill set of is like I'm able to bond with anyone regardless of their background and like who they are. I'm able to talk to them, find something interesting to talk about. Oftentimes it's the tech stuff, but you know, it's, it's. And so I think like the most interesting conferences are oftentimes like, you know, the really big ones because that's where the big stuff is happening. But I think the niches that are really, really exciting is like, you know, spie. Um, so there's ieee, which is International Electrical Engineering something, um, and there's spi, which is another ecosystem. SPI conferences are super, super deep in details. Every single one that I went to, especially like SPI Advanced Lithography or SPI Photo Mask, I went to them. The first time I didn't even understand 90% what I heard. And then I read, read, read, read. I made some context, of course. And the next time I went, I understood like half of what I went to. Third time I went, I understood like 75% of what I went to. Even now I went and I was like, I still don't understand everything that's going on. Whereas, like, you go to like neurips, you know, a couple times you can understand, okay, what's neural symbolic reasoning? Okay, what's this? What's that? Like, you can, you can kind of get a mapping of what everything is pretty quickly. But some parts of the supply chain are so arcane and so deep and so technical. It takes a lot of times for you to even understand what's happening on everything. Um, for every research paper doesn't necessarily mean you did. You go to a conference for a few reasons. You understand the research, you understand, but it's all the research that's being published. But what you really care about is understanding how does that research intersect with technology. Also how does that research differ from what's there today? And none of these research papers tell you what's happening today. But then you just ask people and you build contacts and you learn, and then you learn about the supply chain and oh, this company supplies this company, even though it's not publicly stated anywhere. Like, you know, you learn that the, the. This chemical is like, cost about this much and a tool uses about this much. And you legitimately, you hear the horror
Speaker B: stories of like, this chemical had a shortage and it totally threw off this part of the supply chain. And then it turns out there's only three companies in the world that make that chemical. And it's like.
Speaker A: My favorite one is I learned, uh, a Japanese guy at that specific Japanese conference, uh, that I went to where no, almost no one spoke English in very broken English. He told me about how, uh, his father worked in this, in this industry in the 1980s, that the only factory in the world that built this chemical, uh, burned down and that caused memory prices to like double or triple. And I was like, wow, not too
Speaker B: different from today, not, not at all crazy.
Speaker C: Um, inference going to be the biggest market on earth, the biggest market beyond earth. Agree or disagree?
Speaker A: Um, I mean, obviously use of tokens is going to be the biggest market, um, and the value that's created from tokens is going to be the biggest market. But I think tokenomics, sort of the use of tokens, adoption of AI, sort of the most important thing that's happening. And inference, whether it's open models or closed models, will be one of the biggest markets in the world. Much bigger than oil, I think, much bigger than many other parts. Inference of AI will be many percentage points of the gdp.
Speaker C: What you've done with inference X, I think is industry standard. Maybe say a word on why you started it, what it does, and what do people misunderstand about, uh, performance benchmarking and inference?
Speaker A: Yeah. So to zoom back, right, like semianalysis, uh, we do a lot of stuff that's like, you know, a lot of it is like research for institutional clients and our subscription first products. But a lot of it is also like, hey, this would just be cool to figure out, let's figure out how to figure it out and just post it publicly. And that begets more and more scale. And so we've done this with a lot of GPU benchmarking and testing and training performance and inference performance. But ultimately we saw inference benchmarking was like point in time. You know, you test it and you take some time and you release it. And it's like slow and arcane and outdated because models change all the time. Every, I feel like every week there's a new model, whether it's a Chinese model or, you know, today, mythos 5 fable dropped and new models are coming out all the time. Um, on the software layer, uh, Pytorch, vlm, SG Lang, new, um, drivers, new, something drops. You know, in fact, the update cycle for most of these libraries is twice a week. So you basically have the software updating all the time and therefore performance changing. Um, new inference optimizations are coming out and those get updated. And so I feel like it's a relentless breakthrough after breakthrough after breakthrough that keeps driving efficiency and cost down, which is why we've seen model costs drop for equivalent quality by like 60x a year. It's incredible. Um, but to stay on top of that, you can't have point in time benchmarking. You need to have benchmarks be living and breathing that is constantly running on the latest hardware, on the latest models. And so we embarked on a project and we got a lot of buy in from the ecosystem. This was only possible because we had enough aura with some of the ecosystem where we were able to get Core Weave and Crusoe and Nebias and Oracle and Microsoft and Amazon and Google and OpenAI to contribute to us, um, compute. And then we were able to work with Sglang and VLM and now Radix, ARC and Infra act, uh, which are the private companies who are sort of leading those efforts, um, the open source efforts, um, to collaborate with us. We were able to get Nvidia and AMD and Google and Amazon. Now because we're adding TPUs and Trainium, uh, to collaborate now we've got all these people collaborating. We've got over $50 million of hardware, uh, donated to us. Um, once we launch TP's and Trainium, it would actually be over $100 million of hardware, um, maybe about like 15 different chip types all running these benchmarks every single day on all the latest model, right? The best model for Moonshot, the best model from Alibaba, the best model from um, there's about five different Chinese models, the best open source models, the best Chinese labs there. We run benchmarks on their models every day and then also the best US open source models, GPTLSS, Nemotron, etc. So we're running these benchmarks every day, um, in an automated fashion. And they run on these, these servers that are dedicated to us for inference benchmarking. And we sweep across so many different configurations and optimization types and then what it creates is, and all the results are public and all the configurations are public. So now we have the Pareto optimal curve because a lot of, you know, times when people are comparing inference performance, they're like taking a suboptimal curve or point for someone else and comparing it to their optimal one. It's like, well yeah, I can make, I can stick. If I drove a Porsche versus some race car driver, obviously I'd drive it slower. This is the same thing with inference benchmarking. And so what we did is we created open source, uh, basically containers for the optimal points across every uh, point on the interactivity. That is how fast is it responding to me versus batch size, that is how many users am I simultaneously serving.
Speaker B: Curve.
Speaker A: And so now anyone who wants the optimal point can just go to inference X, download it and run that as the optimal point and they can check every day if they want, or they can even auto download the most optimal point for that model and their inference performance will be near peak.
Speaker C: Is that curve like the most Important curve in your opinion, the throughput interactivity curve? 1.
Speaker A: Yeah, I think, I think um, most things in hardware infrastructure, model, application layer, everything is downstream of that curve. Right. Is it something that needs to be super, super fast, super low latency? Um, and I don't really care about the cost, so I make batch size very low and I use techniques like speculative decoding or multi token prediction heavily. And there's so many, you know, possible techniques there. Or is it something where actually I'm batch processing a ton of documents and I don't really care about all these things. I don't use these techniques that actually are worse on cost efficiency, but help you with speed for an individual user. Because I just want to pack a bunch of users. I don't care if the document takes all night to process. Right. Um, and right now the way we treat AI infrastructure is it's like one size fits all. But over time we're going to get to the point where there's stuff where you have batch workloads or you need instant response and there's the whole curve that's going to matter for uh, users. And so we see this with anthropic, right Claude code, fast mode cost way more than regular mode. Um, and same with OpenAI's priority queue thing.
Speaker C: Um, sorry, dumb question. How does cost factor into this chart?
Speaker A: So if I, let's say imaginary example, I uh, have a batch size of 100 and I can do 10 tokens per second per user. So in total I'm doing 1,000 tokens per second, uh, off of that one piece of compute, that's one side of the curve, super slow, 10 tokens per second. Um, you know, other side is I have uh, five, uh, hundred tokens per second, but I only have one user. And so maybe 250 tokens per second, one user. And then there's points on the middle that are more Pareto optimal. Right. The average person actually wants like 50 or 100 tokens a second. And maybe you know this the, the number of users I can batch together. So the curve is okay, a thousand tokens total, uh, per second or 250 tokens total per second depending on how many users I batch. And there's a curve in the middle. So, and so ultimately some workloads will actually want the forex cost decrease because the same unit of hardware can do a thousand versus two hundred fifty. And some users I'll pay 4x more because I don't care about the price, I care about time because the person using the tokens Is expensive or the feedback loop that I have here is. Is expensive.
Speaker B: If you had to guess, you choose the time frame, 10 years or 15 years. What percent of inference compete do you think will happen in space? Can be 0%, 50% Sean coming 99%.
Speaker A: Like this is a tough one.
Speaker B: Uh, like whatever time frame. And you're so.
Speaker A: I think, I think the non consensus or at least against SpaceX thing, you know I love SpaceX by the way and I totally would buy the IPO if I could buy stocks. Um, not investment advice. Thank you. Thank you. Not investment eyes. Um, from Sequay either. Um, I don't think that space data centers will really matter in the next, um, you know, three to five years. Um, with that said, I think in 20 years I think the vast majority of compute will be going in space. Um, and so the real factor there sort of what's the cost, time frame? It's the time frame and it's the cost of building power on terrestrial land. And how much power are you going to be able to do on terrestrial land? And I think obviously my views of where inference, how many gigawatts or terawatts are devoted to inference is a crazy curve for me personally.
Speaker C: What's your forecast? How many gigawatts?
Speaker A: Um, yeah, I think by 2030 just OpenAI and Anthropic will have over 100 gigawatts combined. Um, and then you'll add Meta and Google and so on and so on and so forth. It's a humongous amount of compute that will be dedicated to inference. Um, and by 2040 it'll be terawatts. Right. Um, the curve of productivity that we're going to get. And so inference deployments is going to be huge. And so if you look at like 2040, I think probably more than half of incremental compute will be going in space. But if you've got 2030, I think it's sub 1%.
Speaker C: Do you think intelligence per watt has been increasing? Uh, and then it seems like there's still a giant gap between where we are intelligence per watt versus human biology. And so if we are, do you think we are to close that gap? And if so, where is that game going to come from?
Speaker A: Yeah, I think it often depends on what you're doing too. Right. Like a TI84 is way more intelligence per watt in terms of doing math than us. That's like 30 years old.
Speaker B: Right.
Speaker A: So obviously this is like a dumb, dumb, you know, sort of general intelligence. Yeah, but general intelligence, wow. Um, so one of the things Inference X does is we also measure the power and cost of all of this hardware. And so we offer not just throughput versus interactivity, we offer cost versus interactivity, we offer power versus interactivity. And so as far as has intelligence per watt been increasing? Um, I mentioned it's been a 60x cost decrease for same benchmark level. Um, we've also seen the same on uh, intelligence per watt. It's not been exactly 60x, it's been closer to like 40x. Uh, some of the efficiencies are non power ways. But there's been a humongous improvement in intelligence per watt on an annual basis at least so far this year, last year, year before, year before. And I expect that to continue as far as where we are from the human brain. We're many orders of magnitude away. Thankfully, doesn't really matter. We can devote a lot of power to computers. Much easier to power computers than human brains. Like you know, we have sickness, disease and like food preferences, sleep. Ah, exactly.
Speaker B: Let me just ask one more question on the like, on the general theme in my opinion in terms of like intelligence per watt or intelligence per dollar, like any of these metrics, I think there's kind of three levels of input. You can get hardware improvements where the hardware is more efficient. You can get low level systems optimizations like kernel level improvements, matrix multiplication libraries, things like that. Or you can get high level model level or algorithmic improvements at the highest level. To me it seems like in the last three years most of the games have come from hardware level and some from the model level. Like do you think that that uh, is what, do you agree with that? Do you think that's what it'll look like in the future? Like do you think there's a bunch of juice to squeeze in the say like kernel level?
Speaker A: Yeah, Sean, I completely disagree with you by the way. Great, great.
Speaker B: That's why I'm asking question.
Speaker A: Um, okay, so I, I think, you know, one way is to look at it as these three different layers. Um, and in that sense, like okay, from Hopper to Blackwell, which is all we've had over the last three years. Roughly 30x improvement on deep seq on the most optimized deployment, which is, you know, you can see on Inference x there's about a 30x improvement. But you know, over the last three years, um, we've had way more improvement intelligence per watt. A lot of that coming from the model layer. Right. If you look back three years, it's GPT4 now. It's like you Know, maybe like Quinn, one of the M. Smaller Quinn models that's like you know, 27B parameters total and like 2 billion active is like way better. And so you've got this huge improvement on model, you've got this pretty sizable improvement on hardware. But it's that co design layer and I think that's, that's what's important. Right. If you look at the architecture of you know, any of these models, but Deep SEQ is the most famous one at least uh, that's public and people
Speaker B: have seen Deep SEQ got huge efficiency gains from um, like optimization or kernel level optimizing memory.
Speaker A: Yes, I think it's like kernels of course, but it's actually you build the hardware architecture for the chip. So if you look at the shapes of all the Experts in deep seq, uh, v3 they were all optimized for Hopper and if you look at, for V4 they're optimized for Blackwell and Huawei's chip. And what's interesting is despite the fact that TPU's are objectively an amazing chip, you know, and they run all of DeepMind and they do all the training for Anthropic as well, uh, on the pre training side at least TPU's suck at running deep seq, but they are really, really great at running other kinds of models that don't run well. On Nvidia there is some level of such deep optimization that has been done, um, whether it be shapes, network, uh, IO, uh, patterns, you know, how you do the collectives, how you do um, things around the arithmetic intensity of the attention mechanism. All these different things are co optimized between the model and the hardware and the infra software in between. And it's hard to say you can disentangle the game.
Speaker B: Do you think that? My understanding is that China has done this a lot better than the West. The last few years in the deep sea was one of the first models to really do this.
Speaker A: I don't necessarily think so. I think it's more so that the west doesn't tell people what they do. Right. Like OpenAI didn't tell people that, you know, GPT4O was uh, how sparse it was, what the shape size was, all these things. But GPT4O is roughly the same size, slightly smaller than deep seq. V3 and 4O came out, you know, a little bit earlier.
Speaker B: Right.
Speaker A: If I recall correctly is your view
Speaker B: that like all three of these things have been happening simultaneously, like roughly the same rate and the most the biggest gains are when you just co optimize.
Speaker A: Yeah, I would say there's been more gains on the model layer than on that co op, than on the sort of software infrastructure layer and the hardware layer. Um, but there's been innovations on every layer and really the biggest gain and the beauty of the best labs is when they co optimize all three and that's what when anthropic is even though they used many different kinds of hardware, they don't really inference too much on TPUs. They mostly train on TPUs. Um, and they inference a lot on Trainium and GPUs. And GPUs more jack of all trades. But they've optimized their hardware, optimized their model, they've optimized everything. So they can do that. Whereas OpenAI prior models were optimized for Hopper more, now they're more optimized for Blackwell. And you step forward through time, um, these labs and the same with Google, they've optimized. Gemini 2 was really optimized for the TPU. Uh, uh, V6E or Gemini 3 was and then uh, the next Gemini that's coming out is really optimized for TPV7. Um, and so sort of like a lot of these things are being co optimized and actually when you pull that model and put it running on the old hardware, it's really not that great. Um, and so I think a lot of this co optimization is the most important thing. It's called software hardware co design and that's what's like really exciting about like you know, sort of what, what you know, I think my day to day is like, you know, great. You get to look at one layer. There's all these innovations happening here, there's all these innovations happening on every layer. But the real breakthrough innovation is when you leapfrog a few layers, you co optimize and co design them and now all of a sudden you've, you've taken what could have been a 2x here, 2x here and instead of being multiplicative to 8x it's actually 100x because you've optimized it across all three layers. And so that's what's really exciting about sort of like what you see at the labs which you see at like company like Nvidia, who's not co optimizing on the model layer per se, but a little bit from the model layer all the way downstream to you know, silicon or you look at a company like tsmc, they're co optimizing not Just, you know, fabrication, but all the way from the components and the consumables and the tools all the way upstream to what the designs, their chips are, the customers are telling them. Is this co optimization across many layers of the abstraction stack?
Speaker B: There will always be bottlenecks somewhere in that optimization though. They're like lagging behind and then need
Speaker A: to get pulled forward and band aids to shorten that up.
Speaker B: If you had to predict at any level of the stack, it can be literally anywhere. What are some of the bottlenecks you're kind of tracking most acutely the next year? And not necessarily in the supply chain, not in scale, but in terms of the actual, um, and it can be in the supply chain too, but just like, you know, is it memory improvements? Is it, is it that like, just like scaling?
Speaker A: So memory, Memory is, memory is an easy one that everyone's talked about, but I'm not going to talk about it from a supply chain angle. I'm talking about from a technology angle. Right. Memory, um, capacity and bandwidth have been improving very slowly. The NAN cell was invented like 25 years ago. The DRAM cell was invented like 40 years ago. And there's been no major breakthrough in cell like what a NAND cell is. Obviously NAND is like a very simple gate or DRAM cell. There is stuff that could come down the pipeline that could be hugely innovative. But even over the last five years all ah, we've really done is make the HBM more stacks faster. But actually there's new innovations coming in the next few years where instead of stacking the HBM separately from the chip, you stack the memory directly on the chip and that makes your bandwidth explode. Um, and so there's interesting companies in that space and interesting PoCs that companies are trying to do there. I think bandwidth is one of the biggest. Another one is, um, for the history of like silicon, basically for the last two decades at least. You know, how many watts a chip is, can be easily predicted just by looking at it. For uh, for a data center or desktop chip, it peaks up at 1 watt per millimeter squared. And so if a chip is 100 millimeters squared, generally the power consumption is around 100 or a little bit less. Um, and if you look at the newest Nvidia silicon, the newest TPU silicon, it's still on that range of 1 watt per millimeter squared. So you know, chips are now getting to, you know, 1400 watts. Next generation is 2000 watts for Nvidia, um, with Ruben and such, uh, and you move forward to Ruben Ultra. It's going to be like 4,000 watts or something like that. But really there's increasing the amount of silicon. What's exciting is we're now finally doing things and it's in development right now where you actually can pump the amount of power into the silicon, uh, to be way more, more than 1 watt per millimeter squared. And now that all of a sudden means you need less silicon. Uh, obviously it's running at higher power and less efficient in some cases. But you reduce the amount of silicon and you're able to like, open up like thermal issues. Thermal issues. Um, there's uh, interference of like electrical interference issues. There's all sorts of different issues, uh, that crop up. And that's why it's a hard engineering problem. That's why we've stuck at about 1. But what's exciting is the world is trying to change these things, I think. Interesting. Like in a different part of the supply chain is sort of like, you know, people, people talk about like, energy is hard and you know, we have energy bottlenecks and it's like. Yeah, but there's actually like very simple solutions, you know, one could think of. Right. Um, take the millions of diesel engines for trucks that the US has the capacity to make. Um, you can very trivially convert them to be using for gas, uh, in the assembly line, and then stick them up to a electrical motor like back driving it. So the electrical motor generates electricity rather than the electrical motor causing the rotation, uh, of the wheel, for example, but doing it the opposite direction. And now you've generated electricity by pumping gas into something that us can make millions of. Um, and then. Okay, well that sounds like a pain in the ass to service. Uh, right. Because now you have to have hundreds of these on a data center site. Well, actually you can just pull people out of car mechanic shops and have them run around and repair truck engines. Actually, it's actually pretty trivial to not. I don't want to say it's trivial. I couldn't do it. Um, I think you're making a really
Speaker B: good point, which is that because the west wasn't really thinking about semiconductors or even hardware more broadly the last 20, 30 years, we didn't have much innovation. We don't have the best minds thinking about how do you improve these.
Speaker A: Why would you want to go work in hardware when you can, uh, make ads.
Speaker B: Sell ads.
Speaker A: Yeah, exactly.
Speaker C: Um, okay, I'm dying to ask Nvidia versus uh, tpu. What are your thoughts?
Speaker A: Um, I think everyone wants to pick one or the Other for this but it's really a function of like look, you know, you look two years from now, Google's going to make 10 million TPUs and uh, through their supply chain and Nvidia is going to make you know, many more million tens of millions of GPUs and both are going to be 100/billion$. And well, Google's gonna be 100/billion dollars, you know of TPU created a year and, and Nvidia will be you know, 500 plus or you know, whatever. I'm not making a specific estimate.
Speaker B: It's not a revenue forecast. This is just a thought experiment.
Speaker A: Yeah, our research M trains absolutely.
Speaker B: You know, getting ready for the SpaceX I feel.
Speaker A: Are you guys big in SpaceX?
Speaker B: Yes.
Speaker A: Okay, so that makes sense. Um, we're very lucky to be very large investors. Awesome, awesome. Um, so I would say um, the, the case of sort of like Google GPUs versus uh, Nvidia GPUs, they both have like points that are really like in their favor. Right. You know, Nvidia will be like oh, we have switches and we're general purpose. And TPUs will be like well we're more optimized, we're actually more energy efficient and our network is actually more um, optimized for certain types of network architectures. And so you have these counterpoints that both would really get uh, into. And I could with a straight face argue with you that GPUs are way better than TPU's or TPU's are way better than GPUs. But it comes down to hardware, software, co design. So actually the way OpenAI's models are headed, it would be a terrible decision for them to use TPUs potentially. And the way that Anthropic and Google's uh, models are headed, it's actually a terrible decision potentially for them to train with GPUs. I mean it'd be fine for them to train.
Speaker C: What's the fundamental difference there?
Speaker A: There's various things, right? Like the size of the matrix multiply unit is different as a very simple thing and therefore the shape of the matrix multiply. You do the attention mechanism you use, uh, the way that attention mechanism is structured, the way the experts are structured.
Speaker C: So OpenAI and Anthropic are converging the very different model architectures.
Speaker A: I think they have quite different model architectures. In fact, um, OpenAI's are much more sparse, um, and that has benefits. And then anthropics are, you know, they're still sparse but more Dense in general. And that has different benefits. And there's many other things.
Speaker C: Right.
Speaker A: The network topology, right. Nvidia, all of their chips are connected to switches, NV link switches. For Google they have no switch. Um, but what they've done is they've been able to, you know, Nvidia, the NV link can only connect 72 GPUs. For Google, their ICI can connect 8,000 chips at super high bandwidth but you have to pass through other chips to get there because there's no switch. And so there's like, there's trade offs there, there's positives and negatives and that influences the model architecture. It's not necessarily that you should uh, claim one is better than the other because at the end of the day, how do you say that this is better than that when you can't measure them in isolation? Because it also extends up to the model layer. Right. Um, but I remember for a long
Speaker C: time thinking one, the programmability of Nvidia and then just Cuda as such a big moat. It seems to me that the narrative has kind of changed, at least in my mind for the last three to six months. Model companies no longer care about if we have to write custom kernels for this other chip, so be it. We'll work with four or five chips if we have to. Um, Claude and Codex are actually quite good at doing a lot of that optimization work. And so it seems like some of the. And then it's not like there's 10,000 model companies that each need programmability. There's on the order of 10 maybe mobile companies. And so it seems to me that the fundamental premise of tens of thousands of big customers that need CUDA compatibility, it seems that kind of thesis is changing in the last few years.
Speaker A: Yeah, I mean certainly the CUDA mote and software moat is at least partially uh, disentangled because models are just great at coding and all software gets commoditized. In that case I do think there is some level of open source and what people call the CUDA moat is not actually anything to do with Cuda, but it's like the fact that Deepseek, Kimi and Zippuai and Alibaba and Tencent, all these companies, Xiaomi uh, had an awesome model recently. Their models are a co design for GPUs and therefore if I want to run them on GPUs actually in some cases they don't run really well on tpus. Now Google just has to create their own open source model ecosystem or open source Models themselves. So they have the Gemma models and so you end up with like well that's not really Cuda as a moat, it's that the downstream product is more optimized for Nvidia and in these cases these companies are just open sourcing them or like Nemotron is just open sourcing it and then the users of it, for example to open, you know, the inference, uh, API providers, the RL companies that are trying to take open models and customize them for companies business use cases. All these different companies are downstream of the fact that like okay, well I guess I need to use Nvidia because the ecosystem uses Nvidia even though I don't particularly care about writing Cuda kernels because the models are great at that. But it's like the shape of like well this expert, the demod is this and you know the hidden dimension, blah blah is this. Right. And so therefore it's better to run on Nvidia GPUs than it is on GPUs and vice versa, right? If Google were to actually open source really good models, you know this would be the same thing, right? People take their models and they'd be like oh wow, these don't run that well on Nvidia GPUs. Um, I should actually just rent TPUs or buy TPUs and do it on there. For small teams you're going to want to use all the open source software like vlm, SG Lang, um, Pytorch, all that stuff. But the big labs they don't necessarily need to use all that, right? OpenAI's forked Pytorch long ago and you know, Anthropic and all these other people don't necessarily rely heavily on the open source implementation of you know, these things. They forked things or built it on their own already and so they don't need to rely on the open source and therefore now it's more like, you know, I'll choose the best hardware and I'll co design my model and infrastructure software through and through for that hardware. Uh, that is the best and most cost efficient and you know, I'll have AI help me write all that software.
Speaker C: What do you think of Cerebras?
Speaker A: I think Cerberus is a really innovative company. I think in some spots of the market they're really, really good, um, very fast inference. Uh, I think that's a big market, uh, we use fast mode almost exclusively at semi analysis.
Speaker C: By the way, I love how disciplined you've been about accounting for. I don't know if that was one exhibit you did or if you do it consistently, but accounting for the dollar spent in the ROI on each task. Awesome analysis.
Speaker A: Yeah, yeah. Uh, we do it pretty diligently, so thank you. That was the dark, uh, GDP article that we wrote. Um, and so, and also like track everyone's token spend by day and if someone's like spiked up like what did you do? It's like, okay, thank you for telling me that. That seems worth it. Cool. On with my day. I think fast mode is obviously worth a lot for high end tasks. Right. I could just see so many different use cases where super fast tokens are worth it. I can also see the flip side where there's a lot of use cases where super fast tokens aren't needed and therefore, uh, the market won't pay for them and they'll use GPUs and GPUs instead. I think the big risk for Cerebras is I mostly think the best models are the ones that you want to use fast mode on and small models you necessarily might not use fast mode on. I could see that being wrong with financial markets maybe or something like that, like a Jane street high frequency trading or something like that, um, or medium frequency trading. Um, but ultimately, you know, running really large models at really long context is very difficult on SRAM based chips like Cerebras, like Grok. And so now it all of a sudden is like, you know, what happens then if like the models get too big. Right. If OpenAI's model is not on the order of uh, you know, hundreds of billions of parameters or low trillion parameters, but it's actually 10 +trillion parameters now all of a sudden I don't think that that will fit on Cerebras. Right. And then if that doesn't with a long contacts length. Right. If you have a million contacts length now that makes it really difficult to justify, you know, and is all. So far we've seen the bulk of revenue and usage at the labs be on their best model. Even when the model price has gone up. We've seen that, um, there's some data that shows that even though Fable just released today, they've had incredible amounts of people switch to Fable and Mythos, sort of that next tier model, even though it's way more expensive.
Speaker C: And so, uh, that's volume by dollars. Totally. But what about volume by tokens?
Speaker A: Well, I guess who cares about volume by tokens? It's about the dollars.
Speaker C: Fair enough, right?
Speaker A: If I don't care that there's you know, uh, you know, I don't know, 200,000 Mini Coopers or Toyota Camry sold if, if, uh, you know, I don't know, four 150s are 5x ASP and they sell only half as much.
Speaker C: Okay, fair enough.
Speaker A: And therefore the most lucrative market is pickup trust in America. Mostly being facetious, but like, I do
Speaker B: think this is one of the things that you've done so well and differentiates you from almost everyone else is that you, you care so much about the economics in addition to the technology. And I think very few people bridged those two things.
Speaker A: Well, yeah, thank you. I think, I think it's really fun inside of semianalysis because we have 90 people and like a big chunk of them are technologists, engineers across the whole supply chain. Um, and then a big chunk of is people who are formerly at hedge funds. And you see these arguments. Like people are like, oh, well, that doesn't matter. And it's like, then someone's like, well, but cost. And then someone, the engineer is like, no, no, no. But this technology is the coolest. And you see this, you see this organically, like fight it out. Um, and we're pretty informal and you know, given the fact that I was a forum moderator, you can imagine what the exact. Enjoying it. You don't wrestle with a pig because the pig enjoys it.
Speaker B: Right?
Speaker A: Exactly.
Speaker B: Just on this topic, before going to the next question, are there like trigger topics in semis for you? You know, like if someone's, which is like such a meme, you think this person must be a. Like, if, you know, if it's like, oh, you like, memory is the bottleneck.
Speaker A: I mean, it's true. But like, um, I think, I think moreover, the one that really gets me is people are like, AI has no roi. Yeah, infuriates me, right? Like there's like, what's the roi? Or like denying model progress. Right? There's these people that are like, models aren't getting better. They're not reasoning, they can't think, they're going to dead end and plateau. And it's like, bro, the line has been up and to the right in terms of capabilities this entire time. And they're like, look, this benchmark didn't improve. That's because it said 90%. Look at the new benchmarks. Yeah, you saturated now. They're skyrocketing, right? It's like, I think that's more so the issue and challenge, like, I think, think semis are really complex and I don't fault people for um, lacking like understanding of it. Like I learn stuff every day about the semiconductor supply chain from people and I've been studying it for, you know, arguably 18 years, since I started moderating the forums when I was 12. Right. Like, you know, arguably been studying it for that long. But even then like, and it's like live, breathe and that's all I care about. But there's so many layers of the abstraction stack. It's like, like I learned about a new chemical that uh, does like $100 million of sales like yesterday. And I'm like, whoa, didn't know this one existed and what process it did. And it's like, but it's like, you know, you learn about things all the time. It's like, okay, 100 billion seller sales in a, you know, couple hundred billion dollar industries, whatever. But like, you know, it's like, but
Speaker B: it's essential, uh, it's essential.
Speaker A: And it's like actually every chip requires it. It's like, wow, I guess there are a thousand process steps and you know, it's like, oh yeah, you like semiconductors? Name every process step. It's like, no, come on. What I think is the most funny is when people have all the facts in front of them and then they get the conclusion completely wrong. Um, and that's, that happens in our
Speaker B: job all the time too.
Speaker A: Yeah, yeah. I mean I can't, I get, I think my attitude is not to be mad that you do that. It's to do it as fast as possible.
Speaker B: I think the industry, because it's so, it's just like AI is the most important thing in the world right now. And there's so many near term bottlenecks. We talk a lot about the near term. Are there longer term things that you're really excited about? Like say on a 10 year time frame? We talked about orbital data centers, but like, like Silicon M atomics. Do you think they're underrated or overrated on a 10 year timeframe? Are there other things that on a 10 year time frame?
Speaker A: Yeah, I mean I think space is like super crazy awesome in the 10 year timeframe that I'm, you know, for space data centers and all these sort of mining asteroids and all these things which is, you know, super excited about the vision of SpaceX. Right. Um, again, not investment in advice before you hop in. Um, I think on the semiconductor side tremendous market movements and tremendous like things can happen just when things happen one year later or sooner. And so that's all like technology that like, you know, in terms of like co package optics, like, well, Everyone knows it's going to happen by the end of the decade. The debate is like 27, 28, 29, 2030, but some point along there it's going to happen. I think the more interesting thing is there's companies like um, did you guys invest in Naveen Rao's company?
Speaker B: We did.
Speaker A: Okay. Yeah. So I think he's trying to innovate on the silicon layer, on the software abstraction layer and the model layer simultaneously. And he fully understands that it's not uh, like uh, we're going to do
Speaker B: this in a few years. It's not a two year timeframe.
Speaker A: Yeah, it's not a few year timeframe. It's a long term beta. Um, and stuff like that is like, okay, we're going to bring potentially analog compute with energy based models and all this crazy shit all at once. It's like that's exciting. Probably won't work, but that's exciting and I really look forward to it.
Speaker B: Definitely won't work quickly.
Speaker A: Yeah, it definitely won't work quickly is what I should say. I believe in Naveen and I met him very. I think he's one of the first people I met in the industry funnily enough. Like in 2020 or 2021. Um, actually 2020.
Speaker B: It says something about him. I think he's someone in my experience. He's always trying to.
Speaker A: I baited him on the Internet. I baited him on the Internet. That's the guy.
Speaker B: He's always trying to help the younger generation. He's trying to identify talent and he's
Speaker C: also so ahead of his time with Mosaic. I remember getting pitched, no, it was 2019.
Speaker A: I was still anonymous then actually I baited him on the Internet and he started replying and then I just took it to DMs and then took it to a call. And that was the first person who was really important that I talked to in the entire semiconductor industry.
Speaker C: That's funny.
Speaker A: But yeah. Sorry to interrupt.
Speaker C: That's funny. What do you think is the end state of the ecosystem? Do you think every lab, every hyperscaler just has its own chips? Trainium, um, seems like it's now working. So do you think we end up with every lab, every hyperscaler has its own chips, at least for inference and then maybe for training you go to Nvidia or whoever. What do you think is the end state?
Speaker A: I think everyone will try and they won't stop trying. I think ultimately um, supply chains matter what technology you can bring in matters. And more and more as the industry gets Bigger supply chain diversification happens. Um, right now everyone's chip more or less looks the same. It's a big logic compute die in the center and there's some HBM on the right and left and on the top and bottom, top side is networking and then the bottom side is PCIe and other I O. Um, and that is the exact same structure for Trainium, tpu, Nvidia chips. Um, and most of the startups, not um, Grok and Cerebras, those are doing weird shit. But that's cool. Um, I think as you step forward we're going to get more bifurcation of hardware architecture and model architecture and therefore people are going to co optimize them and some of them will end up in local minimas. If it's gradient descent, people are trying to go to the most optimized solution. Some people will race to a local minima and then the question is how do you scoot back over to the absolute minima and to some extent like a general more Nvidia will always be more general purpose than anyone else's chip in general, at least on a parallel AI compute basis because they have so many customers who care about different things, who will always give them feedback in the design. You know, the minima will always be better than them. But is that minima, a local minima? Like is the TP or Trainium or Groq or Cerebras or whoever's design optimized awesomely for here, but in the end state, actually you got to go over here. And so they're the wrong, um, and maybe they make a great time, they're great for a little bit of time, but then they end up being wrong. It's like that's the real question. Um, and so I think, I think there will be a big market for general purpose AI compute. Um, because you talk to people at labs, they don't even know what architecture they're going to be doing in a year. They literally don't know what architecture they're going to be doing in a year. They have bets, they have many research bets and that's this exciting thing, but they don't know where it's going. Generally they know what hardware they have and they're trying to co optimize but ultimately like if a new breakthrough happens on model architecture, it's like just replace the tension mechanism with something else, right? Who knows? Or you know, all of a sudden you know, something happens, the best hardware will change and therefore like are people going to make five year investments on hardware solely on, you know, an ASIC that is more specialized or are they going to do so they're going to have some bucket of more general purpose compute. And so you see this with like Google's paying $11 an hour per GPU to uh, Xai for GPUs. Right? Like that's insane.
Speaker B: Right?
Speaker A: There's a very high amount of uh, obviously compute is limited and so on and so forth, but it's like very like insane. But at the same, you know, despite the fact that they have TPUs, and so there's like some questions there, like why did they do that? M. Google actually has three different design programs for TPUs. They're making a TPU with Broadcom that's a different architecture than the TP with MediaTek, that's a different TP than the architecture that is, you know, I won't disclose, you know, by research. Um, but you know, they're making different architectures. It's not just like, oh, they're making TPUs with a couple vendors. It's the same architecture, it's different architectures. And the third one is a very different architecture from the first two. And so I think people recognize that the local minima can happen and therefore, um, I think everyone will have their own ASIC program. I think everyone will deploy billions of dollars of their own asics. Tens of billions of dollars in the case of Google, hundreds of billions of dollars a year of their own asics. But ultimately they're also going to have workloads that don't use TPUs. Right. Some of the Google bets that are not Gemini DeepMind actually primarily use GPUs. They don't use TPUs. Um, some of them also primarily use TPUs. Right. It's a bit of a broad thing but like, you know, maybe for drug discovery or for Waymo, you might not want to use TPUs. I won't say which one it is. But like, you know, there's, there's, there's, there's different architecture bets and different paths for AI. AI for science may have different algorithmic patterns than, than general intelligence, AGI models. Um, and so I think we'll see, we'll see diversity continue to proliferate.
Speaker C: Yeah.
Speaker A: And because the market has gotten so big, niches will be carved out. And so that makes it possible for companies to have their niche and actually make money, even if the majority of the pie goes to Nvidia and TPU and Trainium.
Speaker C: Okay, love that. Can we talk about the data center build outs? Like one, it Seems like, I mean, by all accounts, if you look at the charts like dollars per compute hour, we are in the middle of a crazy compute crunch. Um, and it seems like it's both a demand and supply side crunch, right? Like demand for Long Horizon agents, skyrocketing supply. A lot of these data center buildouts are delayed. Um, do you think this, we're in a compute crunch for the foreseeable future or do you think it alleviates at some point?
Speaker A: Yes, every quarter we're deploying vastly more compute than the prior quarter. And there's more data centers built than the prior quarter. Um, this year there's going to be 20 gigawatts, uh, even accounting for the delays. And next year there's going to be more than 30 gigawatts accounting for the delays. Um, of course, delays happen on everything, right? Anything hardware can have a delay. That's just the reality of life. Are we going to have a compute crunch for the rest of our lives? It depends on what happens with models. But like the TAM for Mythos, you know, mythos 5, fable 5 is not just like 2x that of opus, right? The model is so much better and it can do so many more tasks that the Tamford is way larger than that. And yet compute in the world did not double in the last, you know, six months right? From you know, Opus or uh, maybe like seven or eight months since Opus 4.5 launched to now huge, you know, 464748 were improvements. But Fable and Mythos were like a huge step function improvement. The world's compute did not double in that or quadruple or whatever in that same timeframe. But the demand for useful tasks that can be done by AI, the number of useful tasks and the value of them that can be done by AI has. And so now the question is, what happens? Well, Obviously anthropic in Q2 is profitable, their net income profitable, um, excluding stock based compensation. Um, and I think by Q3 they may even be profitable, including stock based compensation. That's like how profitable they're getting. And their margins on a, on a, on an OPUS token at least. Opus for a token is like north of 80% for the API price. They've got a lot of deals where their total corporate gross margins gets clogged down a little bit, uh, because of like how they do bedrock deals and vertex deals and things like that. But ultimately their, their per token margin is so high. Well then if you don't have the capability, they have the capability to pay ultimately every GPU they Buy at a above market rate. You know, they also bought GPUs at above market rate from SpaceX, which is below the rate of Google. But that's because they signed earlier. Um, you know, it's something that, you know, other companies, maybe a venture backed company or a company that's not really got positive, uh, margins can't necessarily do, right? What is the cost benefit ratios? Like every GPU I rent because I'm out of compute capacity, I can immediately turn around and sell tokens on it or every TPU or every trainium, I can immediately sell tokens on it at a positive margin. And if I'm running 75% gross margin and I double the cost of the compute, it's fine. I'm still running 50% gross margin. And spinning up more compute nodes is not really necessarily a human requiring task for them if they're renting them. And so ultimately it's like, well, my NOI still goes up, right? And so I'm gonna rent GPUs at whatever price. At some level. Whatever price I wanna pay, I can pay.
Speaker C: I have almost the reverse question of like at some point does this compute build out, go bump? A night earlier today, I think there was a tweet like Crusoe publicly said one of their customers to halt construction on one of their data center build outs. Like, it seems like everybody in the ecosystem is so levered right now to like, we got to build, we got to go build, we got to build. High leverage, high growth to me is like, makes me very, very nervous as an investor.
Speaker A: Like wait, hold on. High leverage, high growth means small amount of equity has huge upside. You're not a debt investor, you're a
Speaker C: credit, you're an equity investor. Um,
Speaker A: you got to go to the school of private equity equity levered buyouts only.
Speaker C: I actually come from a school private equity.
Speaker B: Oh, awesome. She forgot the school. She had a VC for too long.
Speaker C: Now I just do revenue multiples. Do you see any signs of that? Are you worried about that?
Speaker A: I see what you mean. Right. And that sort of goes back to the model point, right? Obviously if the model's expanding the total economic valuable like work that's sort of the dark GDP report, uh, that we did and that you mentioned earlier. Um, if the work that these models can do does not expand faster than the compute capacity, then that tide turns. Right? And over the last six months that tide has been very much levered in the direction of, um, the models can do more work, are expanding their tam of work, they can do faster Than the compute is increasing and so prices go up. It's very possible that all of a sudden model progress stops. You talk to anyone at Anthropic or OpenAI, maybe they're drinking the Kool Aid, uh, but you talk to basically all of them, they're like, no, no, no, no. Model progress still go up. Um, and so ultimately current methods could stall somewhere. I'm not sure where that would be. It seems like we have line of sight to model improvement. Rapid, uh, model improvement. And in fact, models are improving faster than they were six months ago or a year ago. Because there's, I would call it recursive self improvement. But basically the models are helping write all the infra and launch the next model sooner and sooner and sooner. So you've got this like pseudo recursive self improvement loop going. And so the models are getting better and better and better faster. Um, and so. But ultimately, you know, capital is a big problem, which is why Google raised capital. You know, they've got ungodly amount of SpaceX, right? They own like 5% of the company. I think a little more, but yeah, yeah, maybe. I think at one point they had like 10%.
Speaker B: Mary Page invested a billion dollars at a $10 billion valuation, got 10% of the company. It got diluted like all this. But that was one of the greatest investments of all time. Good job, Larry.
Speaker A: Yeah, so they know they have like $100 billion in the bank that they can sell in, you know, nine months or whatever from the lockup. And they have all the gross profit they do. And yet they still modeled that. And they're like, we need to raise capital. And so they did an offering and it's like, that's insane. So that tells you how much they think they need to spend. But capital is like really, you know, you know, meta's do it. Meta did it, announced that they're going to do a raise. Stock tanked, people don't like it. But you know, that's all these companies are going to raise capital, whether it be debt or equity. At some point, money spigots will have to slow down. But right now, every GPU that Amazon adds, they're making higher revenue. Or every TP or Trainium, whoever anyone adds is making gross profit.
Speaker B: A little bit of a tip on this to turn it into a question for you, but as we talk about this, for me, the thing that's going through my head is almost an alternative hypothesis. For the Crusoe example, I'm going to use an analogy. In oil. Like in oil Saudi Arabia has way lower cost per barrel to produce oil than a lot of other countries. There's also like the purity of the oil. A lot of, you know, Saudi has generally like very low contaminants in their oil which makes the refining easier. All of this the question from me is like when you look at, for every gigawatt that's being put in the ground, I've called the 20 gigawatts coming online today, like how much like how much homogeneity do you see in those gigawatts? Is it something like. And you can tell me whatever metric you think is right. But like are uh, Google's gigawatts two times more valuable than say most NEO clouds because they have optical switches and they have like they've been doing it for a long time and like they know how to do power smoothing. Because I think this could be the alternative hypothesis that some of the people that are, it's like the people that are good at ah, building data centers, they, they should just do it to the max because there's so much demand and there's so much better than it. But then maybe we're starting to see the early signs the people that are like not as good at it kind of getting hit a little. So I don't know the reality. I'm just curious.
Speaker A: Yeah.
Speaker B: How do you think about so far?
Speaker A: Um, there are metrics for this, right? So Trainium uh sells at sub $10 billion per gigawatt rental rate uh to anthropic and to OpenAI GPUs at least before the craziness of the last 6 months usually went around 12 to $13 billion per gigawatt. So the rental rate, and this is from a Neo Cloud versus Amazon even. And now when Amazon sells GPUs they'd also be 13 or so.
Speaker B: And my understanding of that also is that those number like Amazon subsidized that a little bit so that it's like I actually think the numbers were even. Like I think the despair is even more.
Speaker A: It's less than 10, it's less than 10. But there's like some weird basically how much.
Speaker B: Yeah and like look I, my understanding obviously like Anthropic played a big role in making training useful in terms of you know, writing all the libraries etc. And so like I, everything I hear is that training's really freaking good hardware and it's getting way and way like way better. And obviously Anthropic is now using it a uh, lot so hopefully we would see that price go up, you know
Speaker A: like yeah, the deal they did was actually like there's a floor mechanism and like if it didn't do well it would be like cheaper than to the point where it's cancelable. And you know, if it did really well the price is kind of higher. Um, but effectively um, less than 10. Right. Is, is where Trainium shakes out at whereas GPUs. I mean the SpaceX deal again was like 25 or something crazy billion dollars per gigawatt or $25 million per megawatt. Right. A year rental rate with Google I was like that's a crazy divergence. Now obviously if Amazon was selling training today, it'd probably be more expensive than 10 because the compute shortages. But you do see this already in the sense of uh, with data centers, oftentimes the rental price of a data center if you're doing co location right, not compute in there but just power. Here's the data center. Um, you price it generally on a dollars uh, per kilowatt per month. And so they used to be $60 per kilowatt hour per month and now you see things transacting at anywhere from like 120 to 160 um, but different quality data centers. This actually I've seen data centers go as high as 200 um, when the customer is not such a great credit rating and then the data center is a pretty good one. And I've seen stuff go as low as 100 still or in like India go like as low as 80 because the grid's not reliable, the Internet connection's not great and it's a pretty mid data center. But at least it's a data center. Um, and so you see this huge discrepancy there already. Um, in the case of like data center construction, usually the pitfalls, they just fail. There's a lot of people have failed. You claim they're going to, they're like, they're like four guys. They're like yeah, we heard, I bought some turbines, I put the money down for them. I'm going to build a data center. And then they get delayed, delayed, delayed and fail. Um, so you have to like probability, wait time, wait time lag of the teams that suck versus don't. Um and sort of, you know our data center M model does that. We kind of track every data center uh, and try and do this for every single one based on you know, equipment that they're using and all these things. One of the things you mentioned about Google is you know, in a gigawatt data center they actually put like 1.5 gigawatts of hardware. And because they have such understanding all the way from workload to, you know, they're able to slosh the power around. And so instead of, you know, constantly, you know, a gigawatt of compute, which typically runs at like 60 or 70% utilization in terms of power consumption, not utilization of the hardware. Someone's always renting it, um, they're now running it at like, you know, you know that 60 to 70% means it's at a gigawatt and you're using the full gigawatt. Um, you see people uh, doing deals with, including Google with utilities, where they're like, oh, well, I know this grid can sustainably take a gigawatt, but you know, except for three days of the year you can actually do 2 gigawatts. So give me 2 gigawatts and then just tell me to turn off and so they'll do that. And so these sorts of tricks and then you need to have supreme management of workload, backup power, all these things, um, generators on site to figure out how to actually keep it 2 gigawatts sustainably. When people do this, they're able to charge more. Whether it be I'm actually selling 2 gigawatts despite only having 1 gigawatt because those three deal days I'm able to deal with via battery, gas, etc. Or I figured out how to build power on site. Now I have a gigawatt where no one else does and so I'm able to do it quickly. Um, it's not necessarily transacting for a higher price. It's that I'm selling more gigawatts. And sometimes there are levers where you're selling more gigawatts, uh, where each gigawatt is selling at a different price. I think it's more on the data center and energy layer. It's more about just having it versus not and then that being delayed or not. It's more binary. But on the compute side, I do think there's a lot more, um, interesting work there.
Speaker B: Right?
Speaker A: A, uh, gigawatt given to anthropic is objectively worth more revenue than a gigawatt given to OpenAI. And it seems that both of them could sell every gigawatt that they have right now. Given rate limit problems and token max limit and all these sorts of things that OpenAI anthropic, uh, especially since Codex 5.5 came out, it's much better. And then likewise if you gave a gigawatt to SpaceX they turn.
Speaker B: My guess, like, my suspicion is that they probably make better use of the, you know, hardware than most people. Um, just like I think people underestimate how much networking experience they have from Starlink in particular, and also how much is like power management experience they have via from Tesla.
Speaker A: Yeah, people like Brett Mayo are like, incredible. Like, very good. Yeah.
Speaker B: And so I think that like, for me that's actually, I think probably the thing that might be. I don't actually know the answer, but I think that might be missing from the analysis a lot of people are doing.
Speaker A: I think it's also the fact that when coreweave builds a gigawatt, even though their GPU compute is objectively better than Amazon or Google or Microsoft's in terms of performance, we've tested the performance and reliability. Um, the problem is Google sells it six months before they have it up and they need to turn around and take that paper that they signed to get debt, uh, with that credit backing, and then turn around so they can actually pay for the PO that they've already issued, you know, for the order that they've already issued. Whereas SpaceX was like, no, no, no, this is running now buy it. Right. And it's a big discrepancy when you have a balance sheet to do that versus not. And that also helps your revenue per megawatt like be much higher.
Speaker C: Why does the Neo Cloud opportunity even exist? Because if you had asked me five years ago, I would have said the hyperscalers are going to own this. And you mentioned just now core Wave has better performance than the hyperscalers. I. Why does this opportunity exist? Maybe at the macro level and then in the execution level.
Speaker A: Yeah. So in 2023 I wrote a report that had, ah, Amazon really hate me. Um, it was called Amazon Cloud Crisis. So I talked about how Amazon was the best cloud because they had their nitro nics which offered like tenant isolation. All the hypervisor ran on the nic and then you could sell all the cores and they had custom SSDs that they made and they'd buy them raw NAND and they'd have lower cost because they'd buy the raw NAND and build their own SS SSDs. Um, and you know, they had their cost of Graviton CPUs and that drove down cost per core. And so they had all these things that enabled them to sell more cores, have better security, good networking for. But this was all for the traditional cpu, better storage for the traditional, you know, cloud world. But in the AI, cloud A lot of this stuff hurt performance, right? These nitro nics were bad for performance, still are worse performance, although they've caught up a lot because they've had a couple iterations to like, you know, improve them, but they're still worse for performance. Um, a lot of the security stuff doesn't matter because it's not like I'm time splicing users or splicing a socket into many users, right? It's like no one rents a single GPU and an 8 GPU server. No one rents a single GPU and a 72 GPU rack. They rent the whole rack and in fact they write many of the racks and so, and then, and then there's no like, oh, I rent for six hours and I give it back. It's everyone has these long term contracts. So the mechanics m of the GPU rental market meant that a lot of the expertise of the hyperscalers fell away. Um, and a lot of the expertise that they did have were actually, some of them were detrimental, right? Network performance. For Google and Amazon they had custom networks that were better for traditional CPU and for the stuff that they were doing, but actually worse for AI. Um, and then in other cases it's like, well, Microsoft would save money by building their own data centers, but their data center teams were not actually that great. And so when it came time to run, when it was predictable building, it was like fine. When it came time to actually double your forecast for the year, it's like they fell on their face and they had to go get a bunch of neo cloud capacity, I think so performance, I think, you know, I think time to market is another one, right? You know, these massive organizations, no one's getting rich from building this data center faster, right? But you look at Crusoe for example, Chase and all the other people at the team. You know, I was going to name some people at the team, but I'd rather not. You know all these people are getting rich if they fucking deliver these, this compute fast. They're hyper levered equity owners and they're
Speaker B: also all coming from bitcoin. And you're not supposed to say that.
Speaker A: I mean a lot of the data center, like their main data center guy came from Microsoft.
Speaker B: I don't know, I'm just teasing but uh, it's like you learn a lot when you're in a very high fluctuation market.
Speaker C: What do you think Was Jensen playing 4D chess?
Speaker A: Jensen Absolutely hates a world where all the hyperscalers have all the power. There's a reason he's like, blowing money on, like, random AI labs that, like, I don't even know if, like, it makes sense to. But, like, you know, he's blowing money and pumping them up and going to, you know, everyone around the world and saying, you should invest in this company. Because he wants to create a multipolar world. That's why he loves Chinese labs, because he wants to create a multipolar world. A world where open, anthropic and Google models are the only models, is one in which he's screwed.
Speaker C: Yep.
Speaker A: Right. Um, a world in which, you know, the hyperscalers are the only ones building compute is one. He screwed it. And so, you know, of course he needs to point the allocation gun at Neo Clouds, help backstop their clusters, do anything and everything. Because while today a GPU sold to Crusoe and a GPU sold to, um, Core Wave and a GPU sold to Google and Amazon are all the same price for him five years from now, Crusoe and Core Weave existing means Google TPU will be weaker and means Amazon Trainium will be weaker. And more inference being done with, you know, non closed source model labs is better for him. So I think, you know, the NEO cloud ecosystem is, you know, it's these people that are Wild west, these Neo Labs as well. A lot of them have investments from Nvidia. It's the Wild West. Some will fail, many will fail, but, you know, some will emerge as really great teams. Whether it be, you know, oddly Crusoe, who's a bunch of crypto guys who then started building data centers and doing flared gas stuff, or, you know, Core Weave, who initially was a bunch of hedge fund guys.
Speaker B: They were also. And then they were doing crypto guys,
Speaker A: but then they, like, built, you know, there were a lot of people who didn't bubble up. Like them started around the same time, just failed. Right. So I think, you know, uh, I
Speaker B: got to say both those teams are phenomenal.
Speaker A: Yeah.
Speaker B: A lot of credit. And that's your point.
Speaker A: But yeah, I mean, my point is like, he, he, you know, you throw. It's like you throw a bunch of like, bait into the water and the best fish will figure out and survive. Right. Um, and sort of the same way with the Neo Clouds and, and he hopes the Neo Labs as well. We'll see if any of the Neo labs really bubble up. But like, you know, Thinking Machines has a few hundred million dollars of ARR. That's pretty impressive. Even though they've had, you know, in the media, it's like, oh, they've lost all this talent. It's like, well. But Tinker is doing a few hundred million dollars of ARR. Like that's pretty impressive for out of the gate a product that's less than six months old or whatever. Um, and we hope the same happens to other Neo Labs. And so, um, you know, he wants a multipolar world, truly.
Speaker B: Congratulations on the success. Thank you.
Speaker A: Thank you.
Speaker B: Just the last thing I'll say is I've seen a little bit of this. I think the public, they can probably tell from listening to you how hard you work. But like it's clear you've just been working your ass off for more than a decade and it led to the last few years of being in the right place, right time. But I think it's unbelievable what you've accomplished and I know it's just the beginning.
Speaker A: So thank you so much.
Speaker B: Thank you for joining us.
Speaker A: Awesome.
Speaker B: Sa.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.