
Forging The Future with Chris Howard · 2026-06-04 · 59 min
Key moments - from our scoring
Substance score
52 / 100
Five dimensions, 20 points each
Microchip's Brian McCarson walks through his framework for understanding AI's evolution across four eras - machine learning, training, inference, and AGI - and explains why the inference era is critical for embedding AI into physical systems, robotics, and autonomous devices. McCarson, who joined Microchip at the start of 2025 after 25 years at Intel, inherited a data center business struggling with missed product timelines and team morale issues. Rather than pursuing layoffs, he shifted the organization's constraints by treating time-to-market as a constant (matching Nvidia's annual technology refresh cycle) and deploying LLM-based AI agents trained on company documentation, JIRA tickets, spec sheets, and customer support logs. Teams now have 24/7 "PhD assistants" powered by Claude or other large language models that augment engineering work - writing spec sheets, building Q&A documentation, analyzing decades of customer tickets for patterns - while humans retain decision-making authority. McCarson contrasts this with companies using AI purely for headcount reduction, arguing the bulldozer analogy matters: giving ten engineers each a bulldozer drives growth, not just efficiency. The conversation touches on humanoid robotics bias, Kurzweil's singularity, and token cost management as LLM consumption scales.
The machine learning era (statistical approaches to predict outcomes), the training era (deep learning and generative AI), the inference era (deploying models at scale across devices and endpoints), and the AGI era (artificial general intelligence parallel to human intelligence). McCarson believes narrow AGI is roughly four years away.
Microchip deployed LLM-based AI agents trained on company documentation, JIRA tickets, spec sheets, and customer support history to work as 24/7 "PhD assistants" for each engineer. These agents draft spec sheets, build Q&A documentation, and analyze customer patterns while humans retain decision-making authority, allowing engineers to design more products in fewer hours.
Nvidia moved from 4-year technology refresh cycles to annual cycles, making speed non-negotiable in data center competition. Like competing in the Olympics, companies must show up on game day with the best preparation they can achieve - time-to-market became fixed and other variables had to be optimized around it.
Companies using AI purely to reduce headcount and improve balance sheets are depleting the organization; Microchip uses AI to augment staff (the bulldozer analogy - giving ten engineers each a bulldozer drives growth, not just efficiency). The goal should be improving business outcomes and team capability, not just cost reduction.
The training era focused on building large, sophisticated models; the inference era is about deploying those models at scale - in data centers, automobiles, robots, and endpoints where humans interact with them. Physical AI will emerge from the inference era as AI moves from being a data center activity to living inside machines and tools.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains a handful of substantive ideas - token cost subsidization, compute-near-memory biomimicry, planning for inference architecture, and Microchip's multicast offloading feature - but these are interspersed with extended filler: humanoid robot philosophy, horse-and-carriage metaphors, Matrix/Borg riffs, and the host's extended solo musings. Useful insight per minute is moderate at best.
focusing on standardization of the training tasks which are the most expensive and planning for inference. Inference will be considerably cheaper than training
we've invented this multicast feature where the GPU if it needs to send data packets to 10 different storage locations today, uh, or traditionally a GPU would have to send 10 different instructions to 10 different destinations. We've created a technology that allows the GPU to delegate that to the switch
The AI-eras taxonomy (machine learning → training → inference → AGI) and the biomimicry compute-near-memory framing are the freshest angles, but neither is genuinely contrarian or first-principles. The Uber/taxi token-economics analogy is mildly clever but generic, and most other frameworks (physical AI, agentic AI, edge inference) recycle Jensen Huang talking points rather than adding new perspective.
if you look at every living organism on planet Earth, from a jellyfish to ah, a cow to a human and everything in between, they all share a compute near memory or compute in memory architecture
So we are going through a very similar replacement type business model right now as what the world saw with taxi cabs
McCarson is a genuine practitioner - almost 25 years at Intel, now leading a data center business turnaround at Microchip - and speaks with the credibility of someone who has shipped silicon and managed P&Ls. He is not a career podcast guest or pure thought-leader, but he is also not a CEO of a marquee company, and the episode partly functions as Microchip brand promotion rather than fully arms-length expertise.
I joined Microchip at the beginning of 2025 and uh, that was after a long stint, um, mostly over my career. I've worked at a few companies but spent almost a quarter of a century at Intel Corporation
In my career, turning around businesses is something that I've done serially uh, over many, uh, different endeavors
There are genuine specifics - TSMC 3nm, 160 lanes at 64 billion transfers/second, 10 teratransfers per second at 80 watts, token costs allegedly 90%+ subsidized, narrow AGI four years out - but most quantitative claims are hedged ('I've seen several articles suggesting,' 'far north of 95%') or delivered without sourcing. Named examples like Amazon Alexa are illustrative but well-worn, and many strategic claims remain at the level of analogy rather than data.
this switch device here uh, you know, in particular has um, uh, it's a switch that has 160 different lanes, um, that can each handle that parallel traffic. But each of those lanes can do ah, 64 billion transfers per second
I've seen several articles suggesting we're getting a 90 plus percent discount right now on the cost
The host asks a few genuinely useful operational questions (token budget management, how to architect for inference, energy constraints) but frequently derails into personal riffs on religion, the Borg, and Doordash robots without redirecting to actionable follow-up. Bold claims - AGI in four years, 90% token subsidies, space data centers - go unchallenged, and there is a mid-episode subscription pitch that breaks flow entirely.
The fourth era is that Robopocalypse?
And to be clear, those PhDs were your AI agents, right?
Computed from the transcript - who did the talking, and the words that came up most.
In this episode, Brian McCarson (Microchip Technology) breaks down AI's four evolutionary eras: machine learning, training, inference, and AGI. In today’s conversation, Brian makes a bold case that most organizations are building infrastructure for the wrong phase. Based on nearly 25 years at Intel and now driving innovation at Microchip Technology, he explains why the real bottleneck in AI today is not compute power but the high-speed switches, retimers, and storage controllers that connect GPUs, CPUs, and memory at scale, what he calls the "nervous system" of AI infrastructure. He warns that enterprises over-investing in cloud-based training architectures are heading toward a costly redesign within 18 months, and that the winning strategy is to plan now for inference-first, agent-friendly systems that push compute as close to the endpoint as possible.
Transcribed and scored by The B2B Podcast Index.
Speaker A: AI is moving faster than most organizations are prepared for.
Speaker B: We've started to see epic movements in artificial intelligence that uh, I never would have thought would have moved at the pace that they are now.
Speaker A: While leadership teams are adapting, AI data centers are under growing pressure and cybersecurity threats are evolving in real time. Today I'm sitting down with Brian McCarson, corporate vice president at Microchip, to discuss how he's helping shape the infrastructure powering the next generation of AI. AI systems.
Speaker B: Some people are using AI, uh, to deplete their organization to improve a balance sheet. We are focusing very clearly on AI to improve business outcomes.
Speaker A: We'll explore the next phase of artificial intelligence, why many leaders still underestimate the risks ahead, and what it will take to keep humans in control as systems become more autonomous.
Speaker B: I hope that as science starts to progress, we're open to making sure we're not trying to force a humanoid form into a place where it doesn't really make sense.
Speaker A: This is forging the future. Welcome to the show, Brian.
Speaker B: Thanks so much for having me.
Speaker A: Absolutely. So you've been working with AI since the 2000s. Um, what really pulled you into this space back then?
Speaker B: What I started working on, uh, around 2001, uh, wouldn't even really classify as AI by today's standards, but it was semiconductor, uh, industry factory automation. The whole idea that raw materials can come into a factory and finished goods can come out without ever touching human hands. And the technology then, unlike AI today, was more of, more like nested, if then else statements. It was machine learning, it was statistical based, but nothing like the level of sophistication that exists today. But then the human race wasn't really ready for what we're seeing today like we are now. But that got me introduced into this idea that there uh, are technological ingredients that could significantly improve business outcomes and eventually even create things that humans can't. And it became a, uh, flashpoint for me intellectually that started this. I best describe it as a very slow rolling avalanche of innovation over my career that has only started to really come to fruition in the last decade where we've started to see epic movements in artificial intelligence that uh, I never would have thought would have moved at the pace that they are now.
Speaker A: Yeah, I mean I've been hearing about AI since, I don't know, the 80s. Right. And even back in the 2000s like you said, which actually don't feel that long ago. But um, you heard about AI, but it wasn't like you said, it wasn't really AI, it was more of a programmatical scripting kind of thing. But the last five years have been absolutely amazing for AI evolution. And we're now on that up into the right exponential curve, um, in AI. So I think the next 10 years are going to be like 100 years. Right. So how's your view of AI evolved from those early days? You know, now we're in these generative inference driven systems.
Speaker B: The best way to describe how humankind has transitioned in artificial intelligence is to break things apart into eras not unlike how humankind has done that. To describe evolution of natural intelligence on Earth, we, we break it into these geologic eras. I think in artificial intelligence you can find a similar evolutionary framework. And that first era that I was introduced to early on is really that machine learning era. That's where we started to understand how there's scientific and statistical approaches that could be used to predict outcomes and improve the behavior of systems. But it was really the second era, which is the training era, where I think things really took off. And the advent of this was really the development of deep learning capabilities. When we went from machine learning to deep learning, we started to be able to apply artificial intelligence methods in ways that we hadn't even imagined before. And visual analytics, image and video analytics I think really kicked off that major transition. And this relatively small company at the time, Nvidia, was a, uh, world leader in graphics processing and had built this amazing gaming business around uh, how do you improve rendering for gaming in real time? So you have very crisp graphics. And I'm certain they had uh, the best intentions going into it, but it's, it almost feels like it was the perfect storm to bring them into a uh, world of deep learning with all the technology ingredients that they were putting together. And that training era, uh, uh, really exploded when the world had its chat GPT moment in you know, roughly about I guess three, three almost four years ago now. And that chatGPT moment, generative AI had already existed before, but it wasn't accessible to humans. It wasn't accessible in an app. We had some features, you know, Snapchat filters and video editing features and things like that that were fundamentally kind of artificial intelligence y but nothing like what the ChatGPT moment for humankind was. And that's really when this generative AI movement exploded. And from that we've seen so many different advancements in training capabilities that it's created a whole new industry around where AI can improve business outcomes, where AI can improve uh, uh, personal outcomes for people and different experiences. But now A new chapter has emerged.
Speaker A: So what you're saying is that in the beginning it was machine learning, then deep learning and then generative AI, at least access. Those are the three first eras.
Speaker B: Yeah. So the first one is machine learning is the first era, and then the second era is the training era. And that started with deep learning and ended with generative AI. Uh, okay, now generative AI and deep learning will continue on. But we've now entered a new era, in my view, which is the inference era. And this is where now that we have models that are large models, uh, small models that are so advanced you can really get exceptional outcomes. Now it's about scale. But scale will need to occur through inference, retrieval, augmentation, ah, inference at endpoints, inference in an automobile, uh, in a kiosk, at an endpoint where a human will interact with it, inference in a robot. Uh, this is the era of inference that we're in. And then once we mature in that, we start to enter in the physical AI, uh, area of innovations. And that's still within, in my view, the inference era. But Jensen Wang, CEO of Nvidia, famously talks about how we're moving into physical AI and that's following our agentic AI or inference, uh, phase that we're in right now. And I think the whole meaning behind that is when AI stops being a data center activity and starts to become something that lives inside a system. A machine, a tool, an automobile of a robot. Uh, that's when the physical world is encompassed in AI.
Speaker A: The fourth era is that Robopocalypse?
Speaker B: Is that the one artificial general intelligence? Uh, the AGI era? Yeah, AGI, ah, that's when, uh, artificial intelligence systems start to, uh, become parallel with human intelligence systems. Um, today we're starting to see little pockets of that, but we're not fully there yet. I do not feel like we have entered AGI. Now there's celebrities in the space, in the AI space, billionaire celebrities that would disagree with me on that. But, uh, I have not seen us reach even narrow AGI yet, which is defined as when you have AGI kind of capabilities in a small narrow area. And AGI is when you have that broad general intelligence. I think we're four plus years away from narrow AGI, but I map it out as those four eras. So you have, um, the learning era, the training era. Ah, now we're in the inference era and we're moving to the AGI era.
Speaker A: And how does that fit into Kurzweil's singularity? Do you see the singularity in Those eras or is that beyond that?
Speaker B: I think the singularity is going to be in the AGI era.
Speaker A: Okay, and singularity being the merging of human and AI. Right.
Speaker B: I don't know if I'm intelligent enough to imagine what that would look like. Uh, but certainly we would need to be approaching narrow AGI levels before that's going to be on the horizon from a predictability perspective. And I think narrow AGI is about four years away.
Speaker A: Interesting. Not that far really. Yeah, I mean, I don't know what'll happen either. Maybe we'll be the Borg, maybe we'll be merging with robots. We'll all have robotic bodies. I don't know, maybe, maybe the robots will be the next generation and they'll talk about, you know, those humans in the past. But uh, I thought there's also some parallels where you think about, okay, if we bring in religion to the whole thing, but how God created man in his own image. Right? And you're like, well, why would he do something like that? Right? Why would he use his own image? And then I'm looking at these robots and I'm like, huh, huh. I kind of get it. We're creating those robots in our own image. Everyone's after the humanoid robot, right? After humanoid robot makes a lot of sense. And we are a human operated world.
Speaker B: We are, yeah, yeah, definitely. I can, I can see that. But I do think there also is some bias there as well, um, in that, you know, there's this assumption that we, we have the right physical construct to be able to do those tasks. Um, and you know, when it comes to humanoid robots, there's, there's only a couple of circumstances where I could see the need for a humanoid style robot. Whenever you're replacing a human's work or when you need a human like emotional experience, and so you want to create some relations, uh, with that, that human that you're interacting with. Um, uh, otherwise if you look at, you know, what Doordash has for their little robot, these cute little red things that drive around the city and deliver Taco Bell to people at 2 o' clock in the morning or whatever, ah, we have those in Phoenix, Arizona. Now, um, those aren't humanoid shape, they're wheeled vehicles. And that's that ideal function for delivering a human is not an ideal, ah, design or morphology for delivering Taco Bell to someone three miles away. I hope that as science starts to progress, we're open to, uh, making sure we're not trying to force a humanoid form into a place where it doesn't really make sense. Just, just purposes of trying to match us.
Speaker A: Yeah, well, I mean I agree with you. I do think that it does make sense in some cases where, yeah, moving through a human built world in a robot form, all the access is designed for human interaction. From that standpoint, that's true. On the other hand, you don't need a human form for a lot of things. And maybe we shouldn't give them all human forms because we'd rather have the robotic arm sitting there mounted at the table and not trying to take over humanity, but we also the R2D2s of the world and in addition to the C3POs. Right. You know, so that's a perfect analogy.
Speaker B: Right on with that one.
Speaker A: I don't know if we have to learn beep boop beep, you know, as a language but you know, so pivoting to, you know, where we're at now and what you're doing over at Microchip, when you stepped into Microchip's data center business, what state was it in?
Speaker B: So, some background. I joined Microchip at the beginning of 2025 and uh, that was after a long stint, um, mostly over my career. I've worked at a few companies but spent almost a quarter of a century at Intel Corporation. Um, so not a stranger to the semiconductor sector and advanced technology. When I came into Microchip, uh, I had this amazing opportunity to work with a phenomenal team of employees around the world. Microchip, uh, had uh, made a number of acquisitions including the acquisition of Microsemi uh, many years back which had a lot of data center oriented products. And uh, that business had really thrived during the COVID years but had some challenges. Um, there was an oversupply of products into customers. I think the company had lost its way a little bit in terms of um, making sure customer support and customer service was the highest priority in all that the team was doing. And that needed to be corrected. And uh, there had been some pretty significant technological challenges that the team had encountered which going through those challenges resulted in several products being multiple years behind schedule. So when I came in the team uh, had some morale challenges. They were feeling a little down about having some disappointed customers being late to market. And it required a technological, behavioral and business oriented turnaround. In my career, turning around businesses is something that I've done serially uh, over many, uh, different endeavors. And so I have a playbook that I follow of how to first go into it with an open heart and open mind and really understand what is happening in the business, understanding what's happening with customers, and then assess the true health of the business in that regard. And what I found was, uh, I just have an absolutely remarkable team with remarkable technologies that happen to have a myriad of issues that came together the wrong time and became self compounding. So what we started off with is just this recognition that all engineers by definition are good at optimizing variables. You give them an equation, they're going to say, what are my constants? Okay, constant of gravity. I get it, I'm stuck with that. Speed of light in a vacuum, I'm stuck with that. I'm not going to change those things. But tell me what my variables are and I'll optimize the variables to get the best outcome. And unfortunately one of the variables that the team had was time to market. And you can do that in some businesses. In the data center business, you can't do it. Especially not when this incredible company called Nvidia is running at what they describe as running at light speed. So it used to be data center technology changes operated at a four year cadence. So every four years you'd have a major technology refresh. And Nvidia just threw that whole playbook out the door and said, no, uh, we're going to do major technology refreshes annually. And so that created a whole new reset for the whole industry. And success has been measured on how well companies were able to adapt to that new cadence. But they are clearly setting the pace. They're the conductor of how the world of data centers is evolving right now. And so we stopped allowing time to market to be a, uh, variable. It became a constant. If you want to compete in the Olympics, you don't get to choose when the Olympics are held. The Olympics, the date's there. You show up for the race with the best level of physical fitness and preparedness that you can, and you show up on game day to race. Uh, and so this team had to make that adjustment. And when time to market became a constant and sacrificing quality to hit time to market wasn't an option, then the team really started focusing on the core things that were holding them back. And a, uh, few things were able to be orchestrated well. Once there was the clarity of purpose and the clarity of variables and we were able to deploy a lot of AI capabilities to really help streamline the work. Um, we needed to have our employees have as much efficiency as possible in reducing menial work so that they could focus on the highest level strategic and intellectual tasks. And it turns out that really improves job satisfaction when you handle it that way.
Speaker A: But how do you do that part? I mean, how do you. Okay, I think most businesses are struggling with this right now. They know they need to use AI, uh, they need to bring that into their teams. Right. They need to improve their operational efficiency. But there's so much information out there, so many tools out there, plus some resistance to the. Obviously for most people, like, uh, how do I, how do I change what I do? Right. You know, how did you decide what path to bring that team up to speed on? We need to leverage AI in order to increase our time to market. Right.
Speaker B: So we had to distill it away from the buzzwords AI for the sake of AI is worthless and it can actually be negative. What we were concentrating on was making sure we had an end goal and agreeing that we were going to use AI to help with that end goal. Now, I've talked to C suite executives at companies whose end goal is to reduce the amount of humans that they have to make the same level of revenue. Their goal is to, uh, eliminate jobs, to improve profitability. Uh, that is not our goal with
Speaker A: AI, which is interesting because a lot of people say that that isn't what they're trying to do, but there is a number of people that are trying to do that. Right? You know, using AI to.
Speaker B: All you need to do is go look at news articles about what companies have laid off 5, 10, 15,000 people because of AI.
Speaker A: Right?
Speaker B: Exactly what their CFOs and CEOs had in mind. When in our case, our use of AI was to help facilitate the work, we needed to have industry leading products to market with the best quality and to improve our customer support in the process. So how do we speed up without having to go hire a bunch of people? Um, and instead what we did is we hired an army of PhD assistants that work 24 hours a day, seven days a week on the specific tasks you want them to work on. And so every person in the team has this personal assistant that's PhD level intellect that can answer questions for you at the drop of a hat. They can help. Let me take that off your plate. You give me the data, I'll go construct that into a spec sheet for your customers. I'll go build the Q and A for you once you give me the facts. So we're just providing the facts and letting the AI systems actually improve some of the more menial tasks.
Speaker A: And to be clear, those PhDs were your AI agents, right?
Speaker B: Those are the AI agents. So PhD level academic intelligence. Now, sometimes false, sometimes accurate, but more often accurate. Much better than Vegas odds, I would say. Most of our chatbots that we use, uh, especially the, you know, the more the latest ones from, from anthropic, uh, when you orchestrate some of those, LLMs really have, uh, fantastic level of accuracy. I mean far north of 95% accuracy.
Speaker A: So if you're a business owner and you wanted to do something similar to what you did, you chose a platform, whether it's Gemini or Claude or whatever, right? And then created these PhD agencies. How do you define an agent? How do you say I want a PhD to work alongside Bob over here and help him with his job? Right, yeah. You need to create what, a bunch of markdown files and kind of are you doing that? Who's actually creating that agent? Who's using the clay, the AI clay to create that PhD to sit next to Bob?
Speaker B: That's a great, that's a great, great question. So the substrate, the canvas, uh, and the paint is provided by these large, uh, corporations. The OpenAI is the end tropics, the rules of the world. What they're offering is this LLM level capability which keeps improving. Every month you get a new release and there's new features, new improvements. Um, we brought those capabilities in house in a locked down firewalled infrastructure and allowed them to learn about all of our documentation, all of our JIRA tickets, all of um, our spec sheets, all the logs that we've had, and created our own augmented database of how can we create a, uh, an encyclopedia, an organization of all of that information in an easily accessible way. And when we have the entire history of every customer ticket we've ever had issued and how we resolved it, um, you can start to find patterns and intelligence that would take a normal human, uh, years to go through. Decades of customer, uh, feedback and tickets and design criteria and spec sheets. And some of these AI systems are able to compile some of that information in days, um, and then use that to help collaborate. So we never have an AI making a decision. We have AI making recommendations for the human. The intent is never to replace the human in making, doing the design work. The intent is to allow our humans to maybe design, you know, in a few years, twice as many products and fewer hours in the week with better outcomes because each of them has, you know, one or even multiple, uh, AI based PhD assistants that are off working 24 hours a day, seven days a week on their behalf to help improve things. And they could come in and in a supervisory way Check the quality of the work, make sure that the context was clear, the approach was clear, the outcomes were clear.
Speaker A: Yeah, augmenting the staff, right, instead of
Speaker B: replacing the staff capability, it's not about replacing them and potentially being able to grow the business to even higher levels than you could have with the existing staff. Uh, but microchip is in a growth trajectory, not an atrophying trajectory. Uh, and it's just like, you know, I, not to be judgmental but you know, some, some humans are using OIC and they're kind of wasting their bodies away. Uh, others are using it for really legitimate medical purposes. And I, I, I've known people who have no, uh, longer have diabetes anymore because the right level of dose, they're not doing it to just to be skinny or whatever. Uh, some people are using AI to, to deplete their organization, to make it as lean as possible, to improve a balance sheet. Uh, others are using it to actually strengthen and make the organization more healthy. M so we are focusing very clearly on AI to improve business outcomes, not just a balance sheet.
Speaker A: Right. I heard one analogy. When you think about development, right, or you know, say a construction analogy. You've got a developer using a shovel and you've now given him a bulldozer, right? Uh, of course he wants to have the bulldozer, but that also means that you might not need 10 people with shovels, right? But if you could actually give those 10 people each a bulldozer, you're going to make a lot more progress a lot faster.
Speaker B: That's right.
Speaker A: If you are enjoying this conversation, don't miss what's next. Subscribe to Forging the future@softech.com FTF and we'll send you our ultimate guide to Edge AI. Learn how to get started with Edge AI, how Edge AI is reshaping the business landscape. And plus you'll get a checklist to evaluate your Edge AI readiness. Thanks for tuning in and subscribe now. How did you handle the token question? Because I see this a lot right now with organizations. Okay, that all sounds fine. And well, everyone has a PhD assistant. But tokens cost money and there aren't really a lot of controls for managing tokens individually. At least it seems that way. Maybe some platforms are better than others, but this guy is gonna use, uh, he's gonna use his PhD and that PhD's gonna cost you $10,000 this weekend. Cause he was on a roll.
Speaker B: So we are going through a very similar replacement type business model right now as what the world saw with taxi cabs. And Uber, you know, the cost today for me to go from my home to my airport in a taxicab, inflation corrected versus in an Uber, is the same now. It used to be much cheaper in an Uber, but that was at the expense of profitability for the company that was trying to disrupt an incumbent business model, which is the taxicab services, by undercutting them so deeply they would go out of business and then you could later raise prices. Uh, and so the consumer ends up paying roughly the same thing that they did before. But there's this brief period of time where everything seems cheaper. Now what we have today is actually a much better experience because I can, on an app, essentially get a taxi directly to me. I can select what type of taxi very convenient. But the cost really isn't that much better now than it was before. For, um, now what we're seeing right now with AI is that the cost for one of these licenses right now that you have as a user, even an enterprise license, is a tiny fraction of the amount of spending that goes in to actually utilize that. And, uh, I've seen companies already say, oh my gosh, you know, for 100 bucks a month, I can eliminate four people's jobs in the company and letting those employees go at, ah, not realizing that in about 18 months they're going to see the cost of that service go up by 10 plus X. And we're already starting to see that transition with token costs. It's going to be a slow inflationary process. It's not going to instantly overnight, uh, be a huge crisis. But it's this idea of create a dependence on a new system, undercut the competition, and then you're going to switch back to a more expensive business model later. And I've even seen some cases now where companies are hiring employees back to do some of the basic functions so they don't waste tokens on those tasks in their AI system.
Speaker A: Interesting.
Speaker B: So there still is always, there is now and there always will be an economic equation to this. But my belief is, uh, right now we are in an era where the economics of deploying AI is being heavily subsidized by the companies that are driving this trying to stimulate growth. And we are the beneficiaries of that subsidization at this moment. But it's not a sustainable level. Um, you're not going to have OpenAI spend, whatever the latest number is, half a trillion dollars on data centers or something, uh, and expect to have a free chatbot. Uh, that's just not how economics work.
Speaker A: Yeah, those tokens actually cost a whole lot more than what they're charging right now.
Speaker B: A whole lot more than what they're charging. I don't know for certain, I have no insider information on what that looks like, but I've seen several articles suggesting we're getting a 90 plus percent discount right now on the cost, which makes
Speaker A: it even more important to try to figure out how to leverage it right now, uh, because things are pretty cheap and you can use it.
Speaker B: So one of the things that my team is doing at the moment, which I would encourage anyone to do, is focusing on standardization of the training tasks which are the most expensive and planning for inference. Inference will be considerably cheaper than training. And so starting to architect for inference is going to lead to the best business outcomes, the cheapest cost per token, and the best and most scalable economics. So if right now you're like, it's so cheap just to do this stuff in the cloud, let's just build our architecture to do that. You are, you are setting yourself up for a very hard redesign that's going to have to come in 18 months. If you're taking advantage of cloud infrastructure today but planning for lower cost scalable inference later, then you're setting yourself up for long term success.
Speaker A: And by planning for inference, how would I plan for inference in that case? I'm doing that training and getting that data structured. Or the data, what does it mean?
Speaker B: Well, yeah, I think that's a really great question. Let's think of it from this perspective and we'll use a smart camera analogy. Uh, so many people have different, various homes, uh, automation, uh, products where you have this intelligent camera and you can program it to say, hey, if you see my face or my car, you uh, can go ahead and unlock the door. Um, or, or you don't have to send an alarm, but if you see someone strange that you've never seen before, then go ahead and alarm and let me know or alert me. Um, in many cases that's actually programmed onto the flash that's on that smart camera itself. Uh, you're programming in two or three faces or whatever you have the memory to be able to store, um, on that device. And so when someone comes into field of view, the camera doesn't have to wake up, try to reconnect to the cloud, send the image to the cloud, have the cloud process it, send it back to the camera, uh, determine, oh, this is a violation, or whatever, I need to send my alert function. Uh, it's doing it right there locally and that's a much more cost effective way of doing it. If you have a small amount of consistent data that you're going to process.
Speaker A: The edge AI essentially right at the
Speaker B: edge AI at the edge is primarily an inference task. 99% of edge AI use cases are going to be inference based. And so when you can store a certain small amount of information at the edge and infer on it, you'll get a much better outcome. I would, you know, give huge kudos to uh, Amazon for what they did early on with their Alexa devices. Um, they programmed a device that in the amount of time it takes the word takes to say the word Alexa is about the time it takes to boot up the device and connect to the cloud. Mhm. And they pre programmed in some very common words that would be used and stored those directly on the device. You know, the hundred most common words would uh, be stored there locally, uh, so that it could infer quickly and then repeat back to you the question, which sounds like the robot validating you, the AI system validating your question, but it's also giving it enough time to connect to the cloud, answer the question, get the feedback, download the connection to the sound clip you want to hear, the song you want to hear, answer your question, share with you the recipe, whatever. Um, so they built in a natural harmony between the Amazon cloud and the Amazon Alexa device in your home. And so that's an inference and training balance. And they did that, you know, ten plus years ago. Ah. And I think that creates a really nice architectural blueprint for all of us to follow when we're thinking about scalability. Are you going to be able to infer this information? Um, or are you building incomplete dependence on one cloud service provider or one uh, LLM provider and it's all going to be cloud based and you want
Speaker A: to make sure that the uh, models can do a little bit of thinking on their own, right?
Speaker B: That's right.
Speaker A: Well, you've also described microchip as the nervous system of AI infrastructure. What do you mean by that?
Speaker B: So when I, when I think about microchip and I think about the business that I manage right now, the data center solutions business, uh, we're not competing with the Nvidia's, the intels, the AMD's for the GPU, the DPU, the CPU, I think of those as like different hemispheres of the brain, if you would. Maybe one's like the medulla oblongata and one's the left hemisphere and the right hemisphere or whatever. Um, that's where most of the Intelligence occurs, but the connection to the physical world, the connection between the left hemisphere and the right hemisphere goes through the corpus callosum, right? It's this huge network of wires, uh, cellular wires that connect these things together so the two hemispheres can communicate. Uh, what we build in Microchip's data center solutions business is some of the world's fastest switch technologies that connect the CPU to the GPU. In some architectural configurations. We produce uh, timing devices that help clean up signals. We call them retimers. They help clean cleanup signals, amplify signals to make sure if a packet has to go a long distance, it arrives at its destination in the right format, there's no signal losses. Uh, we build memory controllers and long term storage controllers. So when, when a GPU needs to communicate to a, uh, memory bank, when it needs to communicate to a storage controller, um, when it needs to communicate with the cpu, uh, we're creating that highway system of switches, retimers and storage controllers that enable that communication, that transfer of information. So we're really an infrastructure company that enables the logical systems, the GPUs, the CPUs, the DPUs to be able to perform their magic. And uh, that's really where I think we're starting to see now, uh, more and more um, need to study that infrastructure. As Nvidia has been running at light speed and innovating GPUs at such a furiously, uh, uh, incredible pace, um, they've outpaced the innovation roadmap M of memory storage and interconnect technology. And so now you ask yourself the question, do I need this latest generation gpu? If I even put it in my system, would I be able to get good performance? Uh, you're getting to the point now. It's like do I put a Ferrari V12 engine into my Toyota Corolla? I could possibly not designed for that and I'm probably not going to get great performance out of it either. And if you don't build the infrastructure around that brilliant engine to be able to handle all that power, some of it's going to waste. And as more and more companies have been building data center solutions, they're realizing, oh my gosh, these data centers are huge power hogs, it's a huge impact to the energy grid. Uh, and if you're not going to fully utilize the GPUs because they're really expensive by the way, obviously Nvidia is doing a great job monetizing their leadership position. Um, I have to be able to utilize that in an efficient way. And so then it's like saying, well, okay, now I built the entire car around it. I have a full Ferrari, but I only have dirt roads with potholes on them all over the place. Like, am I ever going to get the full use out of that vehicle? And so a lot of what we're building is that highway system, that infrastructure, that connective tissue, that nervous system to allow all these innovative applications, these innovative uh, uh, features, the autonomous vehicles, autonomous robots to connect into data center infrastructure through our devices.
Speaker A: Interesting. So in a data center, everyone's always bragging about, oh, we've got, I don't know, 24,000 CPUs, GPUs, latest Nvidia chips and everything. Right. But what you're saying is, of course, uh, that's great, but if they're not actually able to talk to each other in a meaningful or fast way, it sounds like a lot like parallel computing. Remember when you had chips that could do parallel processing and then the challenge was okay, that's fine, but the programmer doesn't know how to actually do organized parallel code. And there was a lot of inefficiency or even in the compilers they weren't actually compiling in a way to take full advantage of these parallel processes. So they're in the computer sucking up electricity and power, but they're only running at 10% or only 1 of the cords was being utilized and the other 16 are sitting there not doing anything. Right. Does your solution help with that? I mean it's more than just connecting uh, an RJ45 network cable between the two processors and calling it a day. Right?
Speaker B: Yeah.
Speaker A: Is there more smarts in that interconnectivity between these GPUs? It sounds like there is.
Speaker B: Right, sure. So if you look at um, our, our latest device, um, you know this, this switch device here, uh, is, is built on um, uh, the, the most advanced node that's scaling in production right now in manufacturing in the world. So TSMC's 3 nanometer node, um, uh, there's work going on in like chiplet designs and things like that that are, that are more advanced but in real scale. 3 nanometer node is it right. Know, um, and you know this device, uh, you know, in particular has um, uh, it's a switch that has 160 different lanes, um, that can each handle that parallel traffic. But each of those lanes can do ah, 64 billion transfers per second. Um, so what that adds up to is 10 tera transfers per second on one device that runs at about 80 watts. Um, but we put even more intelligence into this device as well. Um, so it has a, ah, built in security uh, features on it that uh, allow you to make sure there's a healthy uh, level of across network security. Um, but then other cool features like uh, we've invented this multicast feature where the GPU if it needs to send data packets to 10 different storage locations today, uh, or traditionally a GPU would have to send 10 different instructions to 10 different destinations. We've created a technology that allows the GPU to delegate that to the switch. So it can send one packet of information to the switch with 10 addresses and then the switch sends the 10 packets on behalf of the GPU. So the GPU gets right back to doing work again. And so this uh, although this device is extremely power efficient and actually reduces the overall power level of a data center just by its own merits, there's also knock on effects of how do you improve GPU utilization and remove some of the menial tasks that a GPU would have to do. Uh, and so we've almost been trying to treat these switches like they're an agentic AI assistant to the gpu. How do we make the GPU GPU smarter and more efficient? By offloading some of the lower level tasks so the GPU can really focus on what it's brilliant at.
Speaker A: Which I assume is great for inference workloads, right?
Speaker B: Yeah. For inference it's especially important to have good switching. Uh, in inference uh, scenarios you tend to have much smaller packets that are moving back and forth more frequently versus when you're trying to train a massive LLM. Um, uh, there's much less network traffic in comparison.
Speaker A: Speaking of power, what are some of the energy constraints today? Reshaping uh, data centers, a few things. Besides firing up Three Mile island,
Speaker B: We have a few major supply chain issues that are happening right now, uh, in the data center world. Um, certainly the demand for compute has skyrocketed lately. Ah and most of the world uses a foundry model. So uh, it's creating some constraints with those foundries. So the accessibility to manufacture on these latest nodes to save power and improve performance is somewhat constrained. There's more demand than there is capacity. Uh, but even if the semiconductor world was able to build to 100% of that perceived demand, um, there's not enough concrete trucks, metal fabricators, um, at least I'm just talking about the United States right now, uh, to go build out all the data centers that are planned, um, or proposed. Um, and even if they did, the energy grid is not Stable enough in all regions to be able to even handle that workload. So um, you've got an upstream uh, uh, supply problem of being able to produce enough chips to make these data centers that are forecasted and then an actual construction of that and then an operational risk. And now there's a behavioral challenge that's come into to it where uh, some regions are starting to oppose new data centers coming in because they see it as potentially increasing electric costs for homeowners. Um, um, there's been you know, some data center providers that didn't have enough of a stable grid so they put on like propane power generators and stuff to operate, which is terrible environmentally. Um, it's really, it's just a horrible greenhouse gas emission source, um, just to bypass the length of time it takes to uh, get the energy grid to keep up. And energy distribution is highly regulated as it is right now. Um, you think about California which has had rolling blackouts in the summer, hot summer months for years without adding uh, double the electric consumption of a bunch of new data centers. That becomes an almost insurmountable problem to overcome. So we're really dealing with a culmination of issues where um, our infrastructure in constructing buildings, empowering buildings, uh, of normal capacity can't keep up with what we're doing, let alone the massive power demands of data centers. And the regulatory requirements for getting a new nuclear reactor built is on. You know, it's unheard of sadly. Five years to get uh, your reactor permitted, uh, let alone start construction, get it operating.
Speaker A: Well we need to solve the problem before they turn us all into batteries.
Speaker B: Yeah, that's right. And solar is, is a terrific source of daylight sourced energy but you would have to find a way to store that for operations at night as well. So there's, there's a lot of grid stability energy, uh, supply supply issues that we're going to have to overcome in some of these regions to be able to facilitate the demand for AI.
Speaker A: Uh, it's interesting right, because all of those things are more of a physical limitation right now. Right. It's not really the technology problem. It's uh, how do we get enough concrete, right. How do we get enough power energy. Right. Um, not like how are we going to make these computers think for themselves. Um, how are we going to build more inference. I mean we're doing all that but I mean we're constrained by some role, physical limitations.
Speaker B: And, and this is, this is not uncommon when humankind goes through major technological transformations. Um, you know we, we as a species encountered this with the invention of the automobile.
Speaker A: Yeah.
Speaker B: You know, and that was. The invention of the automobile was absolutely devastating to the horse and carriage industry. I mean people built entire careers around horse and carriage industry and they were completely upended and out of work and had to find new roles. I think it still bettered human society as a result of it. But um, it was a major disruption, the Industrial Revolution. Major disruption, uh, the advent of the Internet, the advent of mobile computing. These are major inflection points in humankind and none of them as substantial as what we're going through right now.
Speaker A: Well, it does always feel insurmountable because like you said, if you bring up the horse and carriage example and then you think, oh, well, we have these cars and the cars are much better running on roads and we don't have any of these roads anywhere in the United States or the world. And now you're like, yeah, we're going to pour concrete from Los Angeles to Jacksonville, Florida. Right. It's just one example. We're going to pour i10. Right. And you're like that, just, oh, uh, how. Doesn't even seem possible. But you look at now the road network that we have and over time we've solved it. Now we don't have to pour a concrete, uh, road, but we do have to build a data center. So I'm sure we'll figure it out.
Speaker B: But we will for sure. But it's going to create new market opportunities and significantly disrupt other markets.
Speaker A: Do you think you're going to take Elon's approach where. Well, we're going to, we, we don't have to build concrete data centers. We need to send these things up to space. It's cooler up there. You know, you don't have, you know, we have access to sunlight. We don't have to worry about people worried about nuclear. I mean, what are your thoughts there?
Speaker B: I, I'll, I'll. Yeah, transparently he's not wrong. Um, you know, cooling data centers is, is non trivial and uh, it's easier to do in the, in the frigidness of space if you have the right technologies behind it. Um, uh, access to a continuous, uh, solar power based on the orbit of your satellite is very feasible in space. Um, and most of these uh, satellites orbit, you know, at 140 miles above the Earth's surface. And if you think about it, most data centers today are farther than 140 miles away laterally, uh, if you're traveling along the surface of the Earth, so there's actually a signal Benefit to that too. And satellites use lasers so you know they're using light uh, to transmit data. So there's a speed variation. Essentially it's you know, over the air fiber optics I guess.
Speaker A: Right.
Speaker B: Laser based. Right. Invisible lasers. But they're communicating down directly through uh, uh, an optical link. So you get a speed advantage too. Uh, the, you know the problem is the cost of deploying one. But uh, there are some benefits to it and Certainly I think SpaceX and uh, is you know, commercially positioning itself in a way that he uh, could make his own business thesis possible. Uh, but certainly he has a vested interest in trying to push that narrative. Sure. Uh, given what he's doing and uh, what he wants to do.
Speaker A: Well, other than moving to space, I mean looking ahead in the next three to five years, I mean what infrastructure decisions today will separate those winners from.
Speaker B: Yeah, I think what, you know, if I were a uh, CFO or I were in procurement or a CEO of a company, uh, my team was coming to me with these huge infrastructure bills, um, I would ask the question about what is our end state, what are we trying to build for, what do we want to have in five years? And I can guarantee you if, if the team honestly goes off and looks at that, they would have a different ratio of inference versus training demand than they do today and it would be a much more inference based solution that they would want to have five years from now. So if you don't want to over invest in one thing that will become less useful for you five years from now, you need to plan for the future. And I would plan for an inference based future. I would plan for agent friendly architectures, I would plan for trying to put as much compute as you can as close to the human, as close to the endpoint, the vehicle, um, the conveyor belt, the airplane, whatever it is that you're in business of. And ah, try to put the compute close to that, um, not to diminish the need for the data center, but to properly partition the workloads that are best suited for the data center to be in the data center. Uh, and if you need a data center on wheels, which is essentially like an autonomous vehicle, but it's just doing inference instead of training a build for that. Uh, make the decision now to understand what you want later. Uh, you know, it's same way that you would tell anyone who was planning to buy their first home, um, well, are you planning on having four kids in the next four years? Uh, if so, a studio condo may not be the right choice. For you, you're really not, you're not buying for what you're going to need later. And you're deliberately building in major transitions you're going to have to do in the coming years. So have some foresight into thinking about what, what your infrastructure needs to look like in the near future and start building towards that outcome.
Speaker A: Are there any upcoming architectural differences that you see that would be necessary for that era?
Speaker B: I've been fascinated for many years about the field of biomimicry. Uh, the idea that you can look at natural evolutionarily based systems and start to understand how natural intelligence NI could uh, be used as a template for artificial intelligence AI. And one thing that's interesting, if you look at every living organism on planet Earth, from a jellyfish to ah, a cow to a human and everything in between, they all share a compute near memory or compute in memory architecture. Um, all biological systems have learned that it's good to think close to where the memories are. And if you store the memories far away and the thinking is somewhere else, you end up wasting all these calories and biological architectures in traffic management. And if you put the compute and the memory near each other, you end up with much better performance of that, that intelligence. Uh, but today's architectural systems, um, I mean they're very distinct. Compute is in one location, you have a rack of GPUs and then you have a storage server somewhere else. So we keep compute and memory in two different locations and there's been a variety of different studies about what kind of inefficiencies that creates, um, and uh, ways to work around that. But something I find very fascinating to me is how, uh, how much more energy efficient and how much better outcomes you can get in inference by positioning the compute and the memory right next to each other when you don't need a huge bank of GPUs in an Edge device, uh, for Edge AI, uh, and you don't need a ton of memory, uh, you can do that pretty effectively. So what I would suggest is, you know, as, as corporations and individuals and scientists and engineers are thinking about the future of AI. Think about the efficiency of having your compute and your memory almost seamlessly next to one another, uh, and what kind of speed advantages you can get from not having a bunch of time wasted in transmitting packets back and forth from the thinking machine to the memory machine and instead trying to make those operate more closely and harmoniously, much like all natural systems do today.
Speaker A: That is interesting. Yeah, we don't have a separate storage organ on Our bodies. Right. It's all up here.
Speaker B: Yeah. And even the way our neurons are architected, ah, these synapses actually ah, form connectivity patterns and they connect to memories, uh, where the computation occurs. Um, and instead our minds, if you look at the human brain in particular, have involved, have evolved in a way that there's portions of our brain that handle all language processing. There's portions of our brain that are responsible for motor skills. There's portions of our brain that are responsible for deeper level thinking. Um, our brain has been architected in a way that allows uh, for speciality of intelligence to occur in a, in a compute near memory architecture. So all the memories needed for language skills are stored right next to where the computation for language skills is dedicated. Uh, and so it's just interesting. Biological systems have a way of selecting for efficiency and certainly I think there's lots we can learn from that. And I love that we're starting to see some of those biomimicry elements start to merge into advance our architectures. And I think inference is going to be where that happens a lot in the coming years.
Speaker A: We need to learn from nature, that's for sure.
Speaker B: Yeah.
Speaker A: Well, thanks for bringing on the show, Brian. Fascinating stuff. And thanks for uh, doing what you can to get that infrastructure built for those inference data centers.
Speaker B: Oh, uh, you bet, yeah, my pleasure. Thanks for having me on.
Speaker A: AI is moving fast and the pressure on power and performance is only growing. Brian McCarson is stepping into that with a turnaround mindset, rebuilding Microchip's data center business as the nervous system that keeps AI powered infrastructure running at scale. Now the real question becomes how fast the systems underneath can keep up. If you enjoyed this episode, follow Forging the Future for more conversations at the edge of innovation. Sa.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.