
AI at Scale · 2026-06-15 · 23 min
Key moments - from our scoring
Substance score
47 / 100
Five dimensions, 20 points each
The infrastructure demands of AI deployment are forcing a complete rethinking of data center architecture. Chris Sharp, CTO at Digital Realty, and Steven Carlini, Chief Advocate for AI and Data Centers at Schneider Electric, outline how power densities have jumped from 10-15kW to over 200kW per rack in just years, requiring data center operators to design not for today's chips but for silicon arriving 2-3 years hence. The conversation moves beyond traditional metrics like PUE to token economics - the actual revenue output per watt - as the measure that matters. They explain why power and cooling must be engineered as integrated systems rather than separate domains, discuss the shift from training to inference workloads as the monetization layer, and address practical frameworks like 'build the base, rent the spike,' which combines private AI deployments on open-weight models with public cloud bursting. The episode tackles supply constraints, grid connection delays, permitting bottlenecks, and why infrastructure planning must account for interconnectivity and data proximity - not just compute placement.
Power density is increasing at an unsustainable pace - from 10-15kW to 227kW per rack - while silicon innovations arrive every 12 months, creating a structural mismatch between permanent infrastructure and rapidly evolving hardware that requires planning 2-3 years ahead.
Cooling historically consumed 40% of data center power at legacy densities, but at 227kW per rack, the electrical and thermal loads are so extreme they must be engineered together from the start, using precision techniques like direct liquid cooling to avoid hot spots and stranded capacity.
Token economics - specifically tokens per watt and tokens per dollar - now drive infrastructure decisions because tokens represent actual revenue output; not all tokens are equal in value, and token pricing fluctuates based on compute scarcity and reasoning capabilities.
The 'build the base, rent the spike' approach uses private deployments with open-weight models (like Hermes on Nvidia Spark) for baseline workloads, delivering 30x cost reduction versus public tokens, while renting public cloud capacity for demand spikes and unpredictable bursts.
Grid connection acquisition from utilities (unaccustomed to growth after 20 years of flat demand) and permitting processes remain slow, creating bottlenecks that prevent data center operators from moving at the speed required by the AI race.
Our reviewer’s read on each dimension, with quotes from the episode.
A handful of genuinely useful framings - token economics replacing PUE, power-and-cooling as one system, the silicon/concrete mismatch - but they are surrounded by significant filler, repetition, and high-level generalities. The episode rarely goes deep enough to give an operator an actionable or non-obvious idea they couldn't have reached themselves.
PUE, that was the miles per gallon you got out of your data center. The tokenomics of what you're getting out of your infrastructure today is that new measurement
concrete is permanent. Silicon is not right. And that mismatch and that fundamental kind of alignment of the rapid innovation of silicon, it's really just changing the industry
The token-economics-as-new-PUE analogy and the 'age of inference' framing are modestly fresh, but 'build the base, rent the spike' is a direct cloud-era recycling, and most observations about density increases and planning horizons circulate widely in data-center coverage. Nothing here is genuinely contrarian or first-principles.
build the base, rent the spike
Not all tokens are equal in value and tokens have a longer and longer lifespan, if you will, just to touch at the highest level because of reasoning
Chris Sharp as CTO of Digital Realty is a legitimate senior practitioner who has operated large-scale infrastructure at real volume; Steven Carlini carries the title 'Chief Advocate' at a vendor whose podcast this is, placing him closer to thought-leader/sales evangelist territory. The asymmetry and the branded context dilute overall caliber.
Chris Sharp, the Chief Technology Officer at Digital Realty
Steven Carlini, Chief Advocate AI and Data center at Schneider Electric
There are real numbers - 227 kW per rack, 8 - 10 kW historical baseline, doubling density per year, cooling at 40% of power draw - and named silicon generations (H100s, H200s, Blackwells, Vera Rubens). However, the striking 30x token-cost reduction claim is completely unsupported, and most evidence stays at illustrative rather than analytical depth.
227 kilowatts per rack, when in the past we were talking about 10 to 15 kilowatts per rack
I have an actual private AI running an Hermes agent behind me on an Nvidia Spark, running a Kimmy K2 model. The ROI of the amount of tokens I can produce off of this is just game changing. It's like a 30x kind of reduction in rate
The host asks broad, open-ended questions ('what has changed?', 'what hasn't changed?') and never challenges a single claim - including the extraordinary 30x cost reduction figure. The episode is a branded Schneider Electric production with a Schneider employee as co-guest, structurally precluding any meaningful pushback; it ends with the host reading back a bullet-point summary of what was just said.
What would you say has been the major change when it comes to powering and cooling data centers in the last few years?
We should be checking in with you on this every 60 days. Um, it sounds. Because it changes so fast.
Computed from the transcript - who did the talking, and the words that came up most.
“AI is often framed as a software revolution. But the real constraint isn’t software. It is whether we can power and cool it at scale". In this episode of the AI at Scale Podcast, Chris Sharp, Chief Technology Officer at Digital Realty, and Steven Carlini, Thought Leadership Leader at Schneider Electric, explain a fascinating shift in the data center industry. AI is not only about AI models or chips. As data centers are becoming AI factories success depends on physical systems that must keep up. There is a clear message on planning: decision makers need to design for what is coming, not what exists today. Building only for current demand will not hold up. For C-level executives, this is a practical view of where AI can slow down and what needs attention now. In this episode, you will learn: why power and cooling must be designed as one system to support AI innovation how rising power density is changing data center design what token economics means for cost, value, and ROI why planning ahead is critical to avoid future limits how infrastructure decisions affect AI performance and growth
Transcribed and scored by The B2B Podcast Index.
Speaker A: Concrete is permanent. Silicon is not right. And that mismatch and that fundamental kind of alignment of the rapid innovation of silicon, it's really just changing the industry
Speaker B: from an AI factory and a legacy data center. You know, things have really been turned upside down. You know, we have, you know, the IT part of the, uh, data centers being a small part where the physical infrastructure part, outside the chillers, the UPS's, the generators take up most of the space.
Speaker C: Hello everyone, I'm Tom Kroczak and this is the AI at Scale podcast by Schneider Electric. In this show we focus on artificial intelligence and its impacts, looking at it from various different perspectives, but always with a very practical approach and based on valuable insights from true experts in the industry. In this episode, over the next 20 to 25 minutes, we will be talking about the technology behind it all. We will be talking about AI infrastructure and what is happening in the data center industry. From power to silicon to cooling to managing data center infrastructure in the era of token economics. And I'm very excited to introduce our two special guests today, Chris Sharp, the Chief Technology Officer at Digital Realty. Chris, thank you for being with us here today.
Speaker A: Thank you Tom, excited to be here. Uh, appreciate the opportunity.
Speaker C: Chris brings a lot of deep expertise in data center infrastructure, digital strategy, global technology ecosystems, and he specializes in building scalable high performance architectures that enable the cloud and digital transformation. Our second guest is Stephen Carlini, Chief Advocate AI and Data center at Schneider Electric. Stephen, welcome to the show.
Speaker B: Uh, very happy to be here, Tom. Huh.
Speaker C: Steven is a global expert in integrated solutions for data centers. Uh, he is an author and speaker at many global forums and industry conferences and has been guiding the development of many industry leading solutions to real customer problems. Together we will be talking about AI infrastructure and the fact that AI is transforming the traditional data center in what we call the AI factory, an environment, a very complex environment in which thousands of accelerators operate as one integrated system, turning electricity into precious tokens and heat that needs to be removed. Um, but in the meantime, as new generations of silicon emerge, as new chip architectures emerge, the, the requirements for powering power distribution and precise cooling of data centers continue to grow. Chris, I'd like to start with you, uh, with a simple question. What does this mean for AI infrastructure and the way it's designed today?
Speaker A: Yeah, I think there's a couple pieces. So thanks for the opportunity, Tom, again. I think one of the things I've always been watching in the market is that there's a structural mismatch. Um, one of the Things we're always looking at is like concrete is permanent, silicon is not right. And that mismatch and that fundamental kind of alignment of the rapid innovation of silicon, it's really just changing the industry, uh, and just creates a high risk and failure models kind of coming out to market. So as you start to look at bringing in these new chip sets and every 12 months there's material jumps in not only their performance, but that performance then translates into power densities. And I think we're seeing racked into these jump from it was not Too long ago, 8 to 10 kilowatts was pretty dense in a coload environment. Um, but you're seeing 20kW up to 200 plus kilowatts in a single design. And so that design cycle is just so quick that we're always looking at that next generation, if not three to four years out as another step function that our customers are going to require to be successful. Because that performance and being able to translate that power, those watts delivered to that, uh, to that infrastructure, that accelerated compute then translates into token production which is actually where you can start to monetize the infrastructure. So if you're not planning for uh, against what three years out, you're going to fail in the not too distant future on being able to deliver those capabilities. And so that mismatch is something that we're always watching and allowing our customers to future proof their designs where don't just look at today's chipset, you really want to be aligned to two to three years out as that evolution of silicon kind of comes to market.
Speaker C: Super interesting and it takes a lot of forward thinking there. Steven, in your experience, what has changed over the years? You have been active in the industry uh, for quite a long time and you have witnessed many transformations and changes. What would you say has been the major change when it comes to powering and cooling data centers in the last few years?
Speaker B: Yeah, for decades we were on a very slow and pragmatic pace with enterprise and cloud data centers and things really didn't move that fast. As Chris was mentioning, the densities that we're seeing in data centers right now are on uh, a traumatic trajectory. It's when we move from CPU based x86 architectures to now accelerated compute, GPU based. What we're seeing is we're seeing a doubling of the power density per rack roughly every year. And that puts tremendous strain on a data center operator's ability to power and cool that. So we have to think ahead as these data center Projects sometimes take a couple of years. So what's out today we have to design for what's coming out in the future.
Speaker C: Right, that is very clear. Uh, Stephen, back to you. Uh, you mentioned uh, the, the risk, uh, the failure model that uh, businesses trying to scale AI capacity can find themselves in. Um, what would you say is the biggest disconnect between ambition and the operational reality today?
Speaker B: Well, the operation reality. As Chris was saying, there's a desire to monetize these AI factories. So there's been a lot of focus on training AI models, but now the focus is shifting to deploying these models as working models and monetizing them and adding value. So companies are, are desperately trying to come up with these models where not only do they add a little bit of value to processes and automation within their industries, but they really are trying to uh, take a giant leap towards AI providing a lot of services. The disconnect is how do we uh, cost effectively build the IT and the data center structure to be able to support uh, these AI inference engines going forward.
Speaker C: That's clear. What you said about power and cooling, I wanted to zoom in on that a bit. Um, in the past these two domains, these two systems were treated separately, were treated um, as separate systems. But with that forward thinking approach that you described, Chris, um, it sounds like now we need to think about the entire data center, the entire AI factory as a whole, uh, and have an upfront strategy for the delivery of power in a very precise way at the right density, uh, and the removal of heat. Uh, Steven, would you want to comment on that?
Speaker B: Yeah, it's in the past, you know, you have your power systems and you would build your cooling systems. It really, you know, wasn't a dramatic concern. Traditionally cooling takes about 40% of the power of the data center and you're designing for that. But with these rapidly increasing densities, you really have to design your power systems to be able to accommodate these high density cooling workloads. We're talking about workloads at 227 kilowatts per rack, when in the past we were talking about 10 to 15 kilowatts per rack. So the power and the cooling have to be designed together as a system enabled to be able to function properly.
Speaker C: Chris, would you agree?
Speaker A: Yeah, no, I vehemently agree. And I think where you were going, Stephen, what are the ways I always talk to both the Enterprise and the hyperscale? Because I get to spend a good bit of time with all of them. There's three core components, there's hardware, software and power. The hardware, uh, I think we touched on it at a high level. The chipset sets, the design that chips are iterating so fast that it's very hard to keep up with that. So you have to continue to watch how that comes together. I uh, would tell you the software and Steven, you said this. Well, the workload or the model that's starting to come to market, be very mindful of that. You touched on it. But I really want to emphasize we're in the age of inference, right? And inference is that monetization, coal face like where you really want to start to generate revenue out of those capabilities. And so that software we're always watching and that's where it also translates in the token economics and the output of that. But that power piece, that's exactly where you're at, Steven. In the core of your question, Tom, is power and cooling are one system. And it does take a material amount of the overall electricity coming into the facility to do the electrical distribution and then cooling that as these rack densities continue to increase. I think that's where we're always watching the fact that you need to be very granular in understanding exactly how you need to cool these chipsets that are bound to that workload. I think it's about precision now. Before people used to talk about just generalized systems throughout the data center. You're going to have hot spots and you're going to have colder spots. It's all around that power density, being able to very granularly apply the right type of cooling technology for that generation of chip. And so I'll just give you my piece that uh, this is what keeps Steve and I busy in the industry is making sure that you're putting that step function of capabilities in place as your customers evolve. And that evolution, you can't go ahead of them because you'll strand capital. So you can't build a fully uh, dense facility day one. So you have to have modularity in your electrical distribution to support that density and then match that with the right type of cooling, be it direct liquid to chip or rear door heat exchanger exchangers. Those can get us up into that kind of 200, 240 kilowatt per rack range and you can continually automate, modulate that and deliver that to our customer requirements.
Speaker C: That is clear and a fascinating evolution. Um, I wanted to come back to what you said about monetization and token economics. Um, what would you say is the biggest challenge and how should business leaders think differently, uh, about the Cost per token about uh, the infrastructure in the story when it really becomes a crucial component of the monetization strategy, um, and without will clearly fail. Um, in this rapidly scaling environment, how can companies think about being future ready?
Speaker A: Yeah, absolutely. One of the pitfalls is the market changes so quickly with the whole AI landscape. And I think this is where people get a lot of fatigue and hopefully listening to us today, they can kind of understand some of the basics and foundational elements associated with the token economics. And I've given many talks about what are the constraining factors. But even in my talk that I did 60 days ago has drastically changed because of the lack of compute right. If you've seen token pricing in the last month, it's increased materially so the longevity of you being able to run older infrastructure longer has become more viable. 60 days ago I was saying, listen, you want to stay abreast of some of these new silicon innovations because you want for every watt to get the maximum amount of tokens out of their environment. Now with the scarcity of being able to place these silicon out into the market, get them turned up, get the actual models running, the prices increase. So it's elongated the actual viability of running that infrastructure. And so what I always try to express to everybody is watch exactly how your environment is going to be performant. So many of us are probably been in the industry like Steve and I for many years where it was like pue, that was the miles per gallon you got out of your data center. The tokenomics of what you're getting out of your infrastructure today is that new measurement. And so as these tokens come into market in the software they're built upon drives the value, right? Not all tokens are equal in value and tokens have a longer and longer lifespan, if you will, just to touch at the highest level because of reasoning. And there's a new skill set which just blew me away the other day around council, where it brings up multiple types of agents to look at a problem from different perspectives. It's that type of advancement in technological kind of capabilities that these token economics are really starting to place a different type of outcome against your infrastructure. So it's not only hey, here's my full stack and how it's going to be operating, but then what is the value of those tokens and how are they being delivered? Because the last piece I would tell you on the token economics, that is a pitfall that is somewhat uncorrectable, that if you're not thinking about the delivery and who's going to be consuming your tokens and the ecosystem of value you're trying to be a part of. The interconnectivity piece of that is often missed, not just getting a single network and putting it out there. You really want to understand who the consumption factors are, the type of consumption and how that's going to work. Those tokens, and this is my last piece here, and I'd love to hear your take on this Steve, is that it's not just about those tokens. Those tokens are generated based on data being able to place that algorithm, which is most of the uh, just simplistic sense, this AI kind of capability close to that data set and the offtake or the output of that, the value is the tokens. Being able to watch that entire equation is everything to our customers.
Speaker C: Thank you, Chris. We should be checking in with you on this every 60 days. Um, it sounds. Because it changes so fast. Stephen, any comments on that?
Speaker B: No, it's just a whole new way of looking at things. I mean we used to have the processors and you could tell they're going to do this much performance for what. Now we're thinking of things with tokens and we have to look at things like, here's the output I want. These are the tokens that I need for that output. What are the right processors? As Chris was mentioning, there's different. There's still a 1/ hundreds, H2 hundreds, the hoppers, the Blackwells and now the Vera Rubens. Do I need the token output of uh, Avira Rubin to run my application? If so, then okay. And that's the way that companies are going to start evaluating. They're going to start looking at these tokens per watts and tokens per dollar to make these decisions. So it's a whole new way of looking at the market and the deployments of these uh, accelerated compute it stacks.
Speaker C: Thank you, Stephen. Um, Chris, what would you say is a practical framework that businesses can apply in the era of token economics and the fast uh, scaling of data center infrastructure?
Speaker A: Absolutely, Tom. And this is one of the things where it's a common understanding in the industry where it's build the base, rent the spike. And you might hear it in the market today around private AI. And so one of the things we're seeing is the ability to buy off the shelf kind of compute. And I have an actual private AI running an Hermes agent behind me on an Nvidia Spark, running a Kimmy K2 model. The ROI of the amount of tokens I can produce off of this is just game changing. It's like a 30x kind of reduction in rate to kind of public, uh, token consumption. And so I would say just beware of there's different ways to generate tokens that will benefit your company, benefit your outcomes to your applications. And so that's one of the things we're always balancing is the ability for people to utilize some of these open weight models, utilize private compute in a very secure, completely private environment for your highly sensitive types of applications. And so we're always watching that market, but always emphasize that to a lot of individuals where just like in cloud, we all went to public cloud and then we kind of came back to private and being able to build that base rent, that spike has a huge ROI that can assist a lot of our customers. Because token expense is not only a CFO level visibility, it's starting to become a board level visibility of the expense associated with a lot of this compute and token generation.
Speaker B: Yeah, I think a lot of the AI models and uh, the private applications are still early and I think as the industry matures and as these applications start providing more value, you're going to see more and more of your private IT stacks. And the companies that we talk to really want to control their own tokenomics by deploying their own. So I think the industry is going to go through maturation, uh, phase. But we're still at the beginning of that. But it's been interesting to see how it plays out over time.
Speaker C: That's super interesting and appreciate, uh, your thoughts on this, both Chris and Steven, a lot of key messages here and I want to make sure we have enough time, uh, to recap on what you said. Uh, but I want to close with a question to both of you. In this rapidly changing environment, in this fast evolving technology space, what hasn't changed? What are still the principles that um, still apply? Is there anything that hasn't been redefined yet and we still need to remember about?
Speaker A: Yeah, I would give you my take and then I'll hand it over to you. Steve is like the fundamentals of watching the complete infrastructure stack, you have to watch the full stack, right? There's space, there's power, there's the connectivity, there's the cooling that still maintains the same. It's just the rate of change within it that's putting a lot of pressure on it. The fundamentals are there. I think one of the other things that many of the successful deployments I've seen and been a part of, you plan really early and you understand your supply constraints and you understand and, and this happened in cloud, right? Uh, there's so many similarities to what we see with how AI is maturing over time in that you got to be buyer beware, don't just go buy some chipset. Steve, you said it right, like CPUs, we went through that era and those things, they increased in compute and capability. But you have to be buyer beware where if you buy this infrastructure and you're not planning and understanding where it's going to be deployed and what the success criteria is on what you're trying to achieve with it, you're going to have a bad outcome. And so balancing that all together, the fundamentals haven't changed, but the drastic nature of how the constraints are against those fundamentals, that's what's really exacerbated at this time.
Speaker B: Yeah. And uh, from an AI factor in a legacy data center, things have really been turned upside down. We have the IT part of the data centers being a small part where the, the physical infrastructure part, outside the chillers, the UPS's, the generators take up most of the space and the whole, the whole uh, purchase process. These data centers are being designed and built as systems. You're not doing the old bid spec thing where you're buying the cheapest components that you can find and putting together as a data center. And what needs to change in the industry is the power acquisition part of things. You know, we're still, uh, data center operators are still going through a legacy process of trying to get grid connections so they have a system in place that's really been tough uh, to, because you know, utilities over the last 20 years have seen no growth and now we're starting to see data centers trying to come online and they need the utility power for growth. Uh, the other, the other part is the permitting process for data centers. You know, those are still, still too slow. So we have this need to be able to uh, um, move fast in the AI race. But we still have some legacy issues that uh, are being addressed and hopefully will be corrected so data center industry can actually move faster.
Speaker C: Great. Thank you, Steven. So, let's recap today's conversation. First, um, it's about designing for the future forward thinking. If the infrastructure for AI is designed around yesterday's technology, it will fail tomorrow. So companies need to constantly reinvent and think forward. Number two, power and cooling, inseparable one system that needs to be designed together upfront to enable uh, the compute capabilities in the future. Number three, token economics. A fascinating area that we have all been watching and witnessing the growth of, um, the fact that AI infrastructure plays a tremendous role in the value creation process in, uh, the way that tokens are generated, is something that we heard about today, and I really, really appreciate your insights on this. And last but not least, build the Base Rent the Spike, a practical approach to, um, managing, uh, and scaling, uh, AI capabilities in private and enterprise environments. Uh, Chris, Steven, thank you very much for being with us today.
Speaker A: Thank you very much. Appreciate the time
Speaker C: and, uh, I look forward to speaking with you again. That's all from our side today. Thank you very much for listening. For more episodes of the AI Assistance Scale podcast, make sure to visit se.com and follow us on LinkedIn. Uh, and stay tuned for more until next time. Thank you very much.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.