Screaming in the Cloud · 2026-09-10 · 37 min
Key moments - from our scoring
Substance score
74 / 100
Five dimensions, 20 points each
Matt Rehder, AWS VP of Global Networking, walks through the infrastructure decisions that enable AWS's famously reliable and performant networking. The conversation centers on AWS's deliberate over-provisioning strategy - building far more capacity than immediately needed to absorb failures and maintain line-rate throughput between instances within the same availability zone. Rehder details the multi-year transition from traditional hierarchical network design (with distinct core, aggregation, and edge tiers) to a flatter topology called Resilient Network Graph (RNG), which mathematical models show is more efficient and reliable at scale, though most hyperscalers have never successfully implemented it. A critical enabler is AWS's 15-year journey building proprietary networking hardware - custom switches and devices that eliminate unnecessary vendor features and firmware bloat, allowing the company to optimize purely for their use case. The entire AWS network now runs on a single standardized top-of-rack switch design across all tiers. Equally important is the virtualization abstraction layer introduced with VPC around 2007-2008, which decouples the physical network operations from customer-facing virtual networks, enabling rapid innovation on the physical side without disrupting customers. For operators evaluating multi-cloud or hybrid strategies, this episode reveals why AWS's networking resilience and performance are genuinely difficult to replicate - they result from systematic, decade-long investments in custom hardware, simplified architecture, and massive over-provisioning that most organizations cannot justify economically.
AWS keeps intra-AZ data transfer free and uncharged to make the network invisible to customers and encourage high-throughput workloads within zones. The company over-provisions internal capacity and absorbs the cost as part of broader AZ scaling economics, rather than metering internal traffic.
AWS over-builds network capacity far beyond immediate demand, deploying redundancy and resiliency throughout the infrastructure so multiple simultaneous line-rate transfers do not compete for limited resources. This approach is more cost-effective than implementing quality-of-service or traffic engineering mechanisms.
RNG is AWS's new flat network topology that randomly interconnects thousands of switches instead of organizing them in hierarchical tiers (core, aggregation, edge). Mathematical modeling showed this design is more efficient and reliable than traditional hierarchies at massive scale, though it required years of simulation and real-world testing before production deployment.
Custom hardware design allows AWS to eliminate unnecessary vendor firmware, features, and complexity tailored to general use cases, reducing the number of failure points and the attack surface. Over 15 years, this has enabled standardization on a single switch design across the entire backbone, which would be impossible with off-the-shelf products.
VPC, introduced around 2007-2008, virtualizes the customer-facing network separately from the physical infrastructure. This decoupling allows AWS to evolve, replace, and optimize physical networking components without changing how customers perceive or interact with their networks.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains genuine technical depth on AWS networking architecture - particularly the RNG flat network redesign, network overbuilding strategy, SRD transport protocol, and the virtualization layer separation - that would be novel to most operators. However, significant portions consist of Corey's anecdotal stories and throat-clearing that dilute the insight density. The technical content is strong but interrupted frequently by tangential narratives.
We build more network capacity than we need, both from a reliability perspective so that we have a lot of redundancy and resiliency. Many devices can fail, many links can fail. We still have sufficient capacity for all of the customer traffic.
the entire network that you see as a customer is fake. Yes. It is an emulated, virtualized imagining of a network designed to look like networks looked 25 years ago
The RNG flat network architecture and the decision to overengineer rather than apply QoS/traffic engineering are genuinely counterintuitive positions that challenge conventional wisdom. The SRD protocol wrapping is a clever technical innovation. However, the core framing - resilience through redundancy and simplicity through abstraction - represents executed strategy rather than novel thinking. The episode lacks contrarian or first-principles challenges to existing cloud paradigms.
mathematical theory will show you that as is actually the most efficient way to build a network, both from a cost perspective, but also from a reliability perspective. The problem with building a flat network is all of the interconnectedness
If you just had more network capacity you don't actually need that stuff, and your network runs much, much more reliably
Matt Rehder is VP of Global Networking at AWS with 15+ years tenure building AWS infrastructure from junior engineer to executive leadership. He has hands-on experience designing and deploying network architecture at planetary scale and participated directly in major decisions like RNG deployment. This is a practitioner at the highest operational level with direct accountability for outcomes, not a consultant or theorist.
I joined Amazon as a very junior engineer back in 2008, and my initial job was kind of doing monitoring of the network and, you know, seeing things breaking and learning to fix it. And I've stuck around long enough and learned enough that I managed to grow into this position
And today we're in production in multiple data centers. Uh, and this is now the new default network for all core services for AWS. For any new data center we're building, we're, we're - basically, we've switched to this new RNG flat network
The episode includes specific technical details: RNG architecture, SRD protocol, VPC launch timing (2007-2008), 15-year journey of in-house hardware, unified switch design across all layers, packet-per-second trade-offs, and EBS/EFA use cases. However, it lacks quantified metrics on capacity gains, latency improvements, cost savings, or deployment timelines. Named examples are sparse - no specific data centers, no benchmark numbers, no customer success metrics provided.
Since, since it's really, I think, since about 2007 or 2008 when we introduced VPC, everything from that point was virtualized
it took us about 10 years to fully get to the point where we could do that, but we're bull-headed, and we just kind of continued to push forward
Corey asks strong foundational questions and genuinely pushes on assumptions (e.g., 'why do it at all if already reliable?', questioning SRD adoption barriers, probing on networking education gaps). He demonstrates expertise and forces Matt to defend positions. However, follow-ups are sometimes diverted by Corey's own anecdotes and nostalgia, and he doesn't press on soft answers - e.g., when Matt says SRD works 'almost all workloads,' Corey asks exceptions but doesn't probe further on real-world adoption friction or competitive positioning.
Why do it at all? Exactly. Yeah. Because at some point it's no longer your bottleneck.
One thing that I have always wondered about is the reason that the internet exists is the idea of interoperable standards...Does that mean that there is a path to a more efficient protocol for some workloads?
Computed from the transcript - who did the talking, and the words that came up most.
AWS VP of Global Networking Matt Rehder joins Corey Quinn to pull back the curtain on the massive network infrastructure behind AWS. They explore resiliency at scale, AWS’s move toward flatter networks, the advantages of building custom hardware, and why AI is making networking exciting again. Show Highlights: (01:08) Meet AWS Networking Lead (02:01) Why AWS Avoids Global Outages (05:23) RNG Flat Network Explained (10:08) Overbuild Capacity And Custom Hardware (17:37) VPC Virtual Network Origins (19:50) Why TCP Still Wins (20:38) SRD Inside AWS (21:54) Opt In SRD Transport (25:31) Networks As Utilities (27:24) Learning And Growing Engineers (29:22) Training Talent In House (31:27) AI Rekindles Networking (34:26) Where To Learn Networking
Transcribed and scored by The B2B Podcast Index.
Matt: We build more network capacity than we need, both from a reliability perspective so that we have a lot of redundancy and resiliency. Many devices can fail, many links can fail, we still have sufficient capacity Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. Somehow, I have managed to breach containment and go talk to people at Amazon.
My guest today is Matt Rehder, who is the VP of Global Networking. Matt, thank you for joining me. Matt: Yeah, thanks for having me. Corey: This episode is sponsored in part by my day job, Duckbill.
Do you have a horrifying AWS bill? That can mean a lot of things. Predicting what it's going to be, determining what it should be, negotiating your next long-term contract with AWS, or just figuring out why it increasingly resembles a phone number, but nobody seems to quite know why that is. To learn more, visit duckbillhq.
com. Remember, you can't duck the Duckbill bill, which my CEO reliably informs me is absolutely not our slogan. So what is it you do exactly? Networking, what's the point of it all?
"That feels like something old people care about," said the old person. Matt: Yeah. Uh, I mean, it… The simplest answer is we make all the servers talk to each other. Um, but, um- Otherwise, Corey: they're just expensive space heaters.
Matt: Otherwise, they're very expensive space heaters. Yeah. So it's, it's really interconnecting all the servers, interconnecting all the data centers around the world, and then connecting them to the, all the people around the world on the internet, is, is the gist of it. Corey: Yes, I… You were kind enough to recently give me a tour- Mm … of the networking lab, one of the networking labs.
I'm guessing there might be more than one. We have more Matt: than one. Corey: Because we're Amazon. Matt: Yeah.
Yeah. We didn't show you the super-secret one, but we can, yeah. Corey: Oh, exactly. Yeah.
The, the real secret one. Yeah. That's where the cables are longer so we can fit more data into them before it comes back around. No, yeah, that's how networking works, right?
My networking background myself- Mm … is such that it's just enough to know what I don't know, and… but it does guide me toward asking some of the right questions. Mm-hmm. There's a whole different sense of scale that I don't think most folk appreciate. Yeah.
Uh, one of the enduring qualities of AWS that has also led to frustration, but I maintain is the right answer, is your networking is phenomenal. It i- You have a harsh separation between regions. You have, knock wood, uh, never yet seen a global networking outage. There's never been a global rolling outage of things going down.
Error, failures are region-bound in virtually every case I am aware of. And that was a choice, and it does mean that when you log into an AWS account, each region makes it look like there are now 31 AWS accounts. Go hunt and find it down. But the resilience and durability story are unparalleled by any other company.
I include every other hyperscaler in that list. Resilience is great, and it's been that way for a long time. Now it seems that- Systems have gotten good enough across the board, especially since modern AI systems are not that reliable. We're not seeing five nines of uptime on anything, but companies are putting it in their critical path to a point where it feels like resilience is something companies pay lip service to largely.
It is an afterthought more than a lot of other things. It just isn't an area of focus. My sense is that this is something that happens when you've been too long without a really good, really notable outage. It's sort of like a self-inflicted, we've been too good.
People forget this stuff can break. How do you see it? Matt: Well, first of all, thank you for the kind words. Uh, we care a lot about resiliency.
That is our number one guiding principle, especially for our network, is massive reliability. And it, and obviously we wanna contain outages to within a region, but we never wanna have them span to a region. We go to the level of isolation, even with inside of individual data centers, we're, we're building isolation at every level to get to that. In, in terms of resiliency, I, I've definitely noticed a trend.
There's, there's increased risk appetite, I would say, from a lot of customers, who they wanna move faster, uh, and they're willing to, to trade off some of the multiple nines of availability that traditionally have been a primary focus for all customers. Mm-hmm. Uh, but I think, I think that's a, a smaller number of customers. I think most customers still care deeply about resiliency.
Uh, I talk to lots of them all the time- Yes … and they're, they're very focused on resiliency. They're pushing harder and harder for higher availability, higher reliability on the network all of the time, which is great. Uh, that's, that helps us all be better. Uh, I think for us it's a, it's a balancing act of how do we build great services that have all the level of resiliency that some customers want, uh, but don't slow ourselves down with all the resiliency, uh, in order to move as fast as we possibly can.
So that's the, that's the challenge that we're facing right now. I- I- Corey: it's interesting. The, the resilience demands are much higher in cloud than they are in the era when most companies ran their own data centers- Yes … just because of the correlation risk. Like, great, today my bank is down, tomorrow your grocery store is down.
Well, you're bad at running websites, and so am I, the end. When AWS has a problem, suddenly lots of companies are suddenly impacted, even those who have previously done a great job of planning for resilience. But, oh, no, we have a critical path dependency on a third-party vendor who did not realize they themselves had that ven- Yep … dependency exposure. Everything's interconnected.
Some of these are circular. And the fact that this is not broadly appreciated, understood, et cetera, by even by many engineers and technologists, again, is testament to you folks getting this very right very early on. Yeah. Which brings us to, I guess, the showpiece thing that you wanted to show me in your networking lab, which is the, I guess, the embodiment of the jellyfish paper.
Matt: Mm-hmm. Corey: Can you explain that in a way that… I, I mean, I could make an attempt at it- Yeah … but I'm gonna sound way dumber than you will. Please take it away. Matt: Yeah, no, a- as you said, uh, like, when AWS was building the cloud, we, we knew customers were trusting us with their business, and it's a, it's a tremendous responsibility.
And in order to convince them that this was a good idea, we knew we had to have nearly perfect availability. We had to be far superior to what they've been able to do. And again, that's always been a guiding principle, especially at our infrastructure and network layers. Network is underneath everything.
If you have a network issue, it affects many services, many customers. And we've had a very, very reliable network for many, many years. We're very proud of our network. But we saw an opportunity to make it even better, uh, and particularly in terms of its availability and resiliency.
And so the new network is called Resilient Network Graph Resilience, or RNG for short. Uh, and the idea is, is the way networks have been built for really the past 20 or 30 years, they're hierarchical, they're factory networks for people who know about networks, and it's really stacked layers of switches. And so there's a tier one, there's a tier two, there's a tier three. They're all interconnecting into each other, and that's how you get big scale networks.
Oh, and Corey: three, you have aggregation switches that are on, on leaf. Yep. You have Matt: edge switches. Distribution cores, edges.
Yes. You Corey: have a core switch. You have an out of band switch for the managing- All the things … the servers that like to break, and Matt: mm. Yes.
All the things. Yeah. And so many layers of switches. And for a long time, people have asked, "Why do you need all those layers in your network, Matt?
Why, why can't you have a flatter network?" And, uh, mathematical theory will show you that as- is actually the most efficient way to build a network, both from a cost perspective, but also from a reliability perspective. The problem with building a flat network is all of the interconnectedness that needs to happen to, to plug everything together. At minimal scale, you can take a, a switcher that might have, like, 32 ports on it, and if you have a small number of switches, you can connect them all together Easy, right?
Uh, if you have many, many thousands of switches, you can't plug them all together. You don't have enough ports to. So you need to effectively randomly interconnect them with some links, and then you have to figure out how to route through multiple different switch hops to get to your destination. There's, there's no structure like there is in a hierarchical network where everything is very, like, clean and, and orderly.
So Matt has taught us for years this is the best way to build a network. Many people have tried to build networks like this. No one, to my knowledge, has ever actually built a network like this at any sort of scale. We decided, uh, three or four years ago that we think we can actually pull this off.
And so we started working down that path and doing a bunch of research, doing a bunch of simulation. This is a very novel concept. Most of the network engineers on my team thought this was a really bad idea. "Don't do this, Matt.
You know, this is not how we run networks. It's gonna break. It's gonna be weird." We, we eventually convinced ourselves through a lot of simulation that, no, like, this will work, uh, and it will work more reliably than the network that we have today.
And so we invested, built a- built an engineering team, ag- did a lot of work, actually started to build these things for real. Uh, learned a lot of lessons, and that's, that's one of the other keys to resiliency is learning a lot of lessons and actually taking the learnings from those lessons and then trying to, like, bake them into your product. And so in order to build this for real, we had to build this in smaller scale, we had to test this out, we had to see how it was going to fail, the problems we were gonna have, and then iterate quickly into that before we could deploy at scale.
And today we're in production in multiple data centers. Uh, and this is now the new default network for all core services for AWS. For any new data center we're building, we're, we're - basically, we've switched to this new RNG flat network from the traditional networks we used to build. Corey: Yeah.
Uh, my take on this is not that i- it's not gonna be possible for you to do… A real brave take on my part given you already done it and proven it. Yeah, with the benefit of hindsight, I, I sound real awesome. No, my question is, why do it at all? Why Matt: do it at all?
Corey: Exactly. Yeah. Because at some point it's no longer your bottleneck. Your, your network is already ridiculously reliable.
So at some point, making it even more reliable- Yeah … seems like that's no longer the bottleneck. That is no longer the area of focus that- Yeah … of contention that is customer exposed. Why spend the investment, research, energy, hardware time- Yeah … et cetera, on that rather than other areas that are more, I guess, directly perceived as impactful to customer workloads? Matt: Yeah.
I, I see any outage in the network as a problem, and while our network is extremely durable and customer applications are highly reliable, we still see behind the curtain, and we still see the opportunities and the risks, and it's just we wanna make this better for our customers. Like, to me, the network is still in the way in the sense of you know it's there, you feel it's there. Sometimes you still have network degradation. If we can make it even more reliable, it becomes more invisible, and it enables more customer workloads.
And so it's part of our customer obsession, I guess, at the end of the day of like, it's just not good enough. And I think it never will be good enough, really. It's just a constant, uh, work that we're doing to continually try and drive improvement in this. And not all customers will care, but some of them do and, and therefore, it's worth it for those customers.
Corey: Somewhere between 10 and 20 years ago. Yeah. I wound up d- helping with a cluster build-out for more or less a 256-node, uh, HPC cluster w- alike, where they wound up putting a whole bunch of servers into racks in a data center cage that they had. Great, awesome.
And the idea was that customers would then pay to basically load their data into this, do all kinds of number crunching. Yeah. And it was a multi-petabyte cluster, which at that time was reasonably impressive. Today, you're like, "Ha, that's cute, I have S3 buckets like that," which, different era.
And a question I had for them during the planning process was, okay, great, I see the aggregation switches at the top of each rack. I see the, uh, the core networking structure here. Great, awesome. So what… how are these petabytes of data getting here exactly?
"Oh, customers will ship us a pile of disks." Awesome, great. "And then we plug them in over there on the bench in the rack. See, we thought about this.
Stop with your impertinent questions." Cool. Just one more, um, what is the sustained throughput rate between all of those switches? And that… and when you saturate, even assuming line rate, how long does it take to fill all of those servers throughout there?
And the answer distilled down to, "Oh, no." Because it's a narrow pipe problem. How do you do this? Uh, it is… and that was not well understood.
These are smart people, I'm not trying to dunk on them. Oh, yeah. Yeah. It, it's the sort of thing that's really obvious the second time.
The first time, you have questions on this. And even now, companies that were born in AWS and are exploring, "Well, what if we move this workload to our own data center? Let's go ahead and build that out," they are misled by a very strange aspect of AWS that I confess I don't fully understand myself, and I'm hoping you can shed some light. Uh, in my experience- And you can do this concurrently with basically everything you imagine.
Pick any two points in your, uh, between, in the same availability zone, in the same cluster, et cetera. You can get damn near line rate network transfer between those points. Mm-hmm. And then when I have a bunch of things doing that simultaneously, I don't see a subsequent degradation.
It is still right there- Yeah at line rate. And the only answer I have on that is witchcraft, which, m- yeah, from a certain level of ignorance, everything seems like magic. Yep. I get it.
How do you do that? Matt: Yeah. We overbuild the network. I mean, it, the, it is that simple.
We build more network capacity than we need, both from a reliability perspective so that we have a lot of redundancy and resiliency. Many devices can fail, many links can fail. We still have sufficient capacity for all of the customer traffic. And so it's, it's really is that overbuilding and building as much capacity as we possibly can to create that illusion of elasticity or that illusion that basically there is no network there.
These servers are all just directly wired together. You can send as much data as you want between them all the time. Um, and that's really been our mission for the last 15 years is if you wanna do that and you wanna do that at massive scale, you have to do it reliably, but you also have to have a cost structure so that you can affordably do that. A lot of the reasons that there are network problems, it comes down to constraints or lack of capacity, and then people get creative in that they wanna, like, maximize the usage of the network, and you get into things like quality of service or how do I do traffic engineering or how do I move this traffic around?
If you just had more network- You don't actually need that stuff, and your network runs much, much more reliably Corey: if I Matt: have a more- Corey: Well, how do you even QoS? Well, I don't know. If you have the capacity to put it all through with the same, uh, latency target- You Matt: don't- … you don't need it … you don't need it. Yeah, and that, and that's been our mission for many, many years, is we really try and keep our network very simple.
I mean, that's another secret to our resiliency. We try not to do creative or fun things in the network. Keep it as simple as absolutely possible, have a lot of capacity, have a lot g- more capacity than you actually need so things can fail and you're still fine, and running a giant network becomes something you can actually achieve. It- but the, but the key there is it's the cost structure and the ability to have that much capacity.
Like, how do you actually pull that off? Mm-hmm. Uh, for us, it's we started investing in building our own hardware. And so taking control of our own hardware, we did this about 15 years ago and started this journey, meant that we could be really prescriptive about just the things we wanted in the hardware and in the software on our devices, and not take along everything else that when you buy from a vendor is gonna come along.
It's like- Well, Corey: you'd have that section of the firmware in case you wanted to plug it into a Mellanox thing later. Exactly. So yeah, we know we're not gonna do that. Yep.
Why bother, uh, taking up the RAM? Matt: Delete that code. Yeah. It's, it's one less thing that can fail.
Again, uh, few- fewer features, fewer functionality, simple, tailored to your use case specifically, uh, is, is really the key there. And so over the last 15 years, we've, again, iteratively matured. We keep making it better and better and better with every generation, and you get to this point where today 100% of the AWS network is built with devices that we, we design ourselves. And, and the other fun secret about the way we build our network, we actually use the same switch everywhere.
Uh, so again, most you get these hierarchical networks we talked about before with core and aggregation and edge and all these other layers, but also each one of those layers would use different type of router for, for different fit for purpose. And there's good advantages and reasons to do that, but then now you have more complexity, you have more device types to manage. They all have their own little nuances. We said, "Well, why can't we just use the basically our top of rack switch, and why don't we just use it everywhere?"
And we've achieved that at this stage. It took us about 10 years to fully get to the point where we could do that, but we're bull-headed, and we just kind of continued to push forward and get to that stage. But so now, like our entire internet network, our backbone network, runs on the same top of rack switch that sits in the top of rack switch and connects to EC2 servers. Corey: Well, I, I have memories, uh, uh, misspent youth of driving a van with a core switch in it that cost more than the van we had rented to do this.
Yep, yep. It's like, "Well, I'm pretty sure we're not insured for this, but that's the boss's problem," certainly not mine. It, uh… And we didn't hit anything, so fortunately it was no one's problem. Yep, yep.
Yeah. It, it's the, the world has changed. The, the way you address these things has changed and, and it's, it's wild. Uh, one caveat, I want- uh, that I know I'm gonna get comments on this, otherwise if I don't say this.
Like, well, yeah, you talk about economies of scale, that's why data transfer is so expensive. I, I get it, but in the context of inside of a VPC, inside of a subnet, you're getting that full magic line rate between two endpoints, assume both EC2 instances, that is free. There is no additional charge metered to customers- Correct … until it starts crossing other boundaries. Correct.
So it's, it's not just that you've overbuilt and, uh, made it super awesome. It, it is not charged explicitly for the most common use cases where those things matter. Correct. And I wanna make sure that nuance is clear.
Correct. Because otherwise it's, well, yeah, if I was charging X dollars per gigabyte, I too would invest in making sure you could shove as many things as possible through it, but that's not what's going on here. Matt: No. That's, that's not charged, and as availability zones get larger, again, there's the magic of elasticity that comes from AWS.
Behind the scenes, those are actual data centers filled with devices that all have to be interconnected together as we scale out these availability zones, which means more and more and more network, and that's also a driver for us to drive cost efficiency. And so we can keep it free effectively, like we don't want to charge for that because we want to get the network out of the way and let customers just move their data at whatever speed they possibly can. Corey: And that's valuable.
Uh, as long as people can predict that it's not gonna cross those chargeable boundaries, which is where it's… It's not even that it's too expensive, it's that I thought it was free and it's not, that scares people. Yes. Uh, one other bit of magic in here that I talk to folks, even folks who have networking backgrounds there, right? You look at the typical… You, you start inspecting the traffic that's going over the wire.
You look and you see the physical, uh, the back address on the physical layer, Layer 2. You look at the IP logic around Layer 3. You look at TCP and all the stuff that's being built in as those things happen. Mm-hmm.
The entire network that you see as a customer is fake. Yes. It is an emulated, virtualized imagining of a network designed to look like networks looked 25 years ago, and it is entirely living on top of what the reality actually is. And that is wildly exciting.
Was it always that way? This episode is sponsored by my own company, Duckbill. Having trouble with your AWS bill? Perhaps it's time to renegotiate a contract with them.
Maybe you're just wondering how to predict what's going on in the wide world of AWS. Well, that's where Duckbill comes in to help. Remember, you can't duck the Duckbill bill, which I am reliably informed by my business partner is absolutely not our motto. To learn more, visit duckbillhq.
com. Matt: Not in the very, very, very first days. It, it, when EC2 originally launched, you could see the actual network. Within about a year, we very quickly realized this was going to be a bad idea, uh, and that's when we invested in building VPC.
Mm-hmm. And since, since it's really, I think, since about 2007 or 2008 when we introduced VPC, everything from that point was virtualized, and the whole idea was make a very simple virtual network for customers, abstract it and separate it from the physical network. That way your physical network can change, and you can do whatever you want on the physical world behind the scenes, and you're not really disrupting customers. You're changing the way customers perceive their network to function, and it's been very powerful for us to have that separation.
Like, I like to tell people I run the real network at AWS 'cause it's the physical network. It's the network network. Corey: I love the fact that it just sounds like you're talking smack. It's amazing.
Like, well, you know, those fake network things. I don't know. Matt: No, and we have, we have amazing teams who build all of the network services the customers use, uh, and it's super powerful. Like, I, I work with those teams, but my teams are not tightly coupled, right?
Like, we can build our network relatively separately from the services that are delivered on top of that network to customers, and that lets us all move faster. Corey: One thing that I have always wondered about is the reason that the internet exists is the idea of interoperable standards. Mm-hmm. And the one Like, there were a bunch of early things that came out, like, oh, AppleTalk.
Sure, great. Uh, IP- ISP, uh, what is it, IPX? ISP- Yeah … SPX, that. Yeah.
Things that I don't even remember because that, they did not win. TCP on top of IP did. It's Matt: very durable. Corey: But if you look at the protocol definitions and see how they are structured, what they are built for, it's the reason the internet works.
It is designed for a wide variety of network environments with a wide variety of ever-changing conditions- Yep … in those environments. Inside of a virtualized network like you have built, a lot of the things that those services and protocols have been built for historically will not happen. You can state that with a certainty. Like, we didn't build the ability for the simulation to ask about the nature of itself or whatever the, uh- Mm-hmm … the challenge is.
Does that mean that there is a path to a more efficient protocol for some workloads as long as it doesn't have to start dealing with the rest of the broader internet? Yeah. How do you think about that? Matt: We already have that.
Um, we're- Corey: You do under the hood. I know that much- Yeah … historically. It was SRD that you were talking about- SRD, yes … a few years back. Yes.
You announced this at re:Invent, and it was great. Like, we have this amazing TCP replacement protocol- Yep … that we use at the time to power, I think it was EFA and a lot of the higher EBS, uh- Yep. Matt: Everything, everything behind- … resources … EFA runs on SRD as well as all EBS traffic, uh, within AWS regions. Yeah, so all the storage traffic.
Corey: Yes. And, and there was like, "Wow, that's really interesting. Can I learn more about it?" And the answer was basically no.
And cool, so why are you telling about it? Like, honestly, it's really neat, and we're proud of it, and it do- it's how we do these things. But outside of the context of how we run, what we do, how we see it- Yeah … it is not something a customer would even find useful, much less make sense to even go too far down the path. Yeah, yeah.
I'm talking about the other side of it. When, once you're in- inside of the virtualized network, once you have… I have an EC2 instance that's talking to another, and I wanna send a bunch of this data over- Yeah very quickly. Yeah. Like, as the connection stands up, you start seeing TCP window scaling.
Yeah. Starts slow, speeds up, et cetera. Yeah. If I have certain guarantees around this, I could see, and these are famous disasters, I could write my own protocol and put traffic over that.
I've been involved with a project once that did it, and I was there- It's hard monitoring it, I said, "Turn it off." Yeah. It's really hard. But I could see you folks putting something like that out for- Matt: Yeah.
We're gonna… We, we have. So SRD, you can go to an EC2 instance right now- Mm-hmm … on your, your ENI or, and you can enable, basically, SRD transport under the covers. It'll still look like TCP that you're communicating with, but behind the scenes, we wrap your TCP packets in SRD, which effectively makes them extremely reliable, lower latency, higher throughput, and send the TCP bits. And so from a TCP perspective, it just kind of looks like you're on a magic network with infinite bandwidth and no packet loss.
But you can… But that way, there's no- That one Corey: snuck completely past me. Fantastic … Matt: but there's no change for the customer, and that's, that's the big thing. Like, you can adopt EFA, right? But if you are adopting EFA, there's a lot of work for you to do in your application so that you can make it EFA capable, and a lot of customers just won't do that or can't do that, and we wanted to, like, how do we make the network better for all of them?
And so right now this is an opt-in feature- Yeah that you can go turn on. Again, we're busily working behind the scenes to be, 'How do we make this the default? Like, how is this just the way it works for everyone?' And there's a bunch of technical challenges, but we're, we're chipping away at that.
Corey: Today at least, what are the workloads for which that makes sense, and which is one of those- Matt: Almost all workloads, honestly. Corey: Okay. Okay. Is there any exception cases?
Ooh, do not use it Matt: for Corey: that one Matt: thing? Or- The, the, the exception case is there's, there is a tax of, uh, maximum packets per second. Mm. So in order to do that encapsulation on the hardware layer, the peak PPS you can get is reduced slightly, and that's the reason we haven't gone and turned it on by default- Okay for people yet.
Again, we're chipping away and r- and reducing and kind of removing that bottleneck. But for the vast majority of workloads, unless you're very, very sensitive to peak PPS, uh, enabling that functionality gets you more reliability in your connectivity. And, and the real way you see this is if you're looking at, like, tail latency for your application. You're looking at your P99, P99.
9 latency. That's gon- usually gonna be dominated by some type of packet loss or some other, uh, other issue. And if you enable this with SRD based transport, mostly that disappears. And that's, that's the reason we've done it for EBS.
Tail latency matters a whole lot when you're talking about storage workloads. And so we were on a mission to be like, 'How do we drive that tail latency down?' And that's where we came up with SRD and moved all, uh, all of the EBS traffic over to SRD. Corey: Okay, I've seen it all now.
All I have to do is talk to one more AWS customer, and I see things that are new. Uh, the challenge I always ask, okay, where is this not an appropriate fit for, is because very often I will talk to folks who hear about these things. They're like, 'Great, we're gonna go and implement that everywhere.' Great.
What does your workload look like? 'Oh, it calls out to, uh, an AI inference provider, and it waits for the response, and the response comes back, and we're getting fast speeds. We're getting almost 200 tokens a second on that. It's-' I promise the network is not your bottleneck right now.
No. It's fine. This is not the place to optimize. Yeah.
There are, there are better paths forward for you. Yeah. I, I, I'm excited for the day where this does become the default, with the obvious caveat then that I can see a lot of workloads that people are gonna try deploying somewhere else get disparate results as a result of this once that becomes a default, and start to wonder. I mean, sure, the, the blame is going to accrue in directions that are not aimed at you because you're the good example in this.
How come it's so crappy in our data center? Probably our terrible networking team. Which is not true. It's just this weird- Yeah … almost under- under the hood magic.
Matt: Yeah. And, and, and really, like something like SRD and our capability to do that, it's, it's similar to our network investment and we wanna own our own devices, we wanna own our own software. And we've done on the, on the server side for years with Nitro. We have our own hardware.
We, we abstract all of the customer bits with VPCs, which means we can do interesting thing behind the scenes and we don't have to change, you know, many, many years long standards like TCP. You wanna change TCP? Good luck. And you wanna try and change everyone's applications.
The way to make meaningful change is to leave TCP alone, but just basically magically make it better by doing something behind the scenes. Embrace, extend, Corey: and yeah. Matt: Yeah. Corey: Yeah.
So as you talk to people about what you've been doing for the last, well, 15 years at least- Yeah uh, what do you find is the biggest misconception that customers have about AWS's network? Matt: About AWS's network? Most people don't know it exists. Corey: I agree wholeheartedly.
Yeah. Yes, yes. It- Matt: Yes. And so it's, it's just the lack of understanding of the sheer scale and all of the investment and all the work that goes on behind the scenes to just make all of this function and work for customers.
And, and friends and family and everyone else, they just see it as, well, there's just, you know, some magic happens to the internet, things happen, so. Corey: Uh, a while back, uh, in one of his iterations, Peter DeSantis, before he transitioned to his Amazon role now, was the SVP of Utility Computing. Mm-hmm. And I love the term because it, it encapsulated an experience I think lots of folks have, even if they didn't notice it.
Uh, when something is a utility, uh, like, uh, electricity or water- Mm-hmm … you don't turn on the switch or turn on the faucet and then wonder if it's gonna work this time. It is expected. Yep. And we've all gone through that as consumers.
There was a time at some point where when I go to google.com in my web browser and it doesn't work, in the early days, Google must be having an issue. And at some point, it switched for all of us, "Oh, my Wi-Fi must be having problems." Yes, it's my problem.
Yeah. Exactly. To the point where on the rare occasion, uh, I can't think of one in the last 10 years, but toward the end of that period when Google would actually be down and have an issue serving, I didn't realize that was even a possibility. How could that happen?
It, uh… And I really think the network, adding AWS to the site, is in so many ways almost a victim of its own success. Matt: The electricity example's exactly what we have in our head. Like, we want the network to be like the light switch. You flip it on and off, you expect it to always work, and it's extremely rare and surprising when it doesn't work, and that's, that's been our mission, is make the network invisible.
Get it out of the way. Corey: A, a question I often get from folks is, "Great, how, how did you get started?" How did I become me? Yeah.
And the answer, I started off as a Unix, followed shortly by Linux- Yeah … systems administrator. And during the, uh, the global financial crisis in 2008, suddenly salary freeze, everyone hates their job, but no one's hiring. What do I do for the next year? Well, I started learning networking, 'cause that had always been an area- Yeah … I hand waved over.
And I did dabble as a network engineer after that. But by and large, it made me a much better systems person. Yeah. Because once I understood what was on the wire, I could reason about it.
I could understand- Yeah … what was going on. The challenge now is that we have complexity on top of complexity on top of complexity and so on and so forth. It is dizzyingly high now. And I worry on some level that networking as a whole is no longer as central and core to modern engineering understanding of this.
There are publicly traded companies built in the cloud that do very well, and their internal networking teams, their entire professional expertise is more or less around configuring things like Transit Gateway- Yeah, yeah … and Cloud WAN, and these are all abstractions on top of it. They don't understand that even the old days of MDI-X autosense, which was a consumer feature then came to us, like, great, when I plug the ethernet cable in, it's not lighting up. Why? Oh, that's a s- that's a rollover cable.
Yeah, you got to cross that over, yeah. You need a straight through. It's… Wait, you mean there's different standards for the, for the wiring standard? I have a mug that just has the colors of the ethernet B spec.
Mm-hmm. And every once in a while, someone sees it, they're like, "Yes!" Like, okay, I found my people. It's great.
But now I sound like an old man talking about the Great War- Yeah … once upon a time. Where does the next generation of all of this come from? Matt: Yeah, we're, uh… It's something we grapple with. And, like, again, we've been very successful such that most of our customers don't have to deal with the vagaries of networking, and you like it.
Yeah. Most people look at that, and they're like, "That's terrible. I don't wanna do that." I Corey: like it, and I'm far enough away from it now that I can wax nostalgic about this.
If I were on the phone this morning with Cisco TAC because of a weird routing issue- Man … I would be considerably less charitable. Yes. But that's 10 years in my past. Matt: It… For, for us, it's really grow your own.
Uh, and so, like, I mean, I joined Amazon as a very junior engineer back in 2008, and my initial job was kind of doing monitoring of the network and, you know, seeing things breaking and learning to fix it. And I've stuck around long enough and learned enough that I managed to grow into this position, uh, and it's kind of the same path we have with all of our new hires. We're, we're hiring lots of junior engineers, either from… We hire lots of people from our data centers, actually, some amazingly talented people who were there.
I was just talking to someone who was visiting one of our data centers, and it's someone who was working at a gas station three years ago, and they were just wowed with what they knew. And I'm like, "We should go hire that guy for the networking team," 'cause they actually touch the physical equipment. They understand how this stuff works, and they, they often become our best engineers in the future. Um, but it is a deep investment in us in terms of, like, hiring younger talent, junior talent, and, like, going through the process and training and growing them and teaching them how we build networks 'cause there's… At the end of the day, there's no accelerator for this.
It is increasingly complicated, as you said. It's much, much more complicated than it was 15 years ago when I was starting. But you still need to know the same foundational things that I had to know 15 years ago, but you also have to know this increasing mountain of new stuff over the top, and that only comes from just- Learning the basics, learning the practical bits, and then slowly adding the knowledge on top with experience Corey: I mean, 'cause you know that. Customers don't have to.
I, I wanna highlight just how rare that is. It, it seems that the idea of developing talent internally- Yeah … has almost gone out of fashion in the industry. It's, "Oh, no, this person doesn't have experience. We can't hire them for the role 'cause it'll take six months to get them up to speed."
Yeah, I get it. Nine months later, the role's still open. Yeah. So what are we really doing here, friends?
Yeah. The idea of teaching people and, and arming them to do this, 'cause there was a time where every company that wanted to be on the internet, which let's face it, is probably most of them- Mm-hmm … needed to have a networking- Yeah if not person, team. Team, Matt: yeah. Corey: Where it was you need people able to handle this stuff.
It gets really weird really fast. Yeah. And now you don't need that. I mean, I, I went through a parallel evolution running email systems.
Yep. Now most companies do not run their own mail servers. Nor Matt: should Corey: you. Thank God.
Yeah. Yeah, n- and networking is, is still continuing to evolve, and I'm not naive to th- enough to think that we are at the end of history, and, like, in our final form. What's the next step? What's the next phase?
Where do you see this going? Matt: Where do I see this going? Uh, it's- It's really more of the same. I mean, the, there's so much, there's been an acceleration in investment in networking because of AI, which is exciting as someone who does networking.
The, there's bandwidth demands are higher, there's more capacity to build, there's more things to connect. And so I've seen more energy go into networking in the last three years than I had in the previous 10 years. And so there's just accelerated in investments both at the silicon and the hardware level, but as well as the software, the protocols, everything else. It, it feels like networking is exciting again i- in a way that maybe we made it fairly boring by the time you got to 2019.
And that it's just working. I will say Corey: when, in my misspent youth, there were times I wished it had been a lot less exciting- Yes … some nights. Yes. But yes.
Matt: Yes. Um, but yes, uh, and so I, I think it's just gonna be more of that. Like, we're gonna continue to find… The world has an insatiable appetite for bandwidth. Yes.
A- and I think we're only scratching the surface of that. And if we were able to 100X the internet and network capacity, it wouldn't take long for applications to find uses for that. And I think there's still a lot of things that are not done, or they're done in ineffective ways because of either, uh, not enough capacity of network and/or, like, costs of networking. And so the more you reduce that and eliminate that, I think you just open up new opportunities and innovation.
And so I think it's just gonna be this constant march for more and more and more and more capacity. Well, Corey: we, we've seen that in real time. The, the thing that drove that home to me more than almost anything else was the day that GDPR took effect because a lot of US-based websites, uh, suddenly were not fully compliant, so they turned off all the tracking and all the rest. Yeah.
And suddenly the internet was blindingly fast. Super Matt: fast. Corey: It was like, wow, you're loading an article, and, like, the total load was something like- Yeah … 50K. Yeah.
It was wild to me. And you look at the same article with all the other stuff turned on, it's why is loading that 50K article taking 25 megabytes- Yeah of nonsense? It - As soon as you have the capacity, people find ways to fill it. Or just- I'm not opining on whether it's all useful … or Matt: just look at streaming video, for example, right?
Yeah. Like, I mean, you, you can get very good quality with the, the latest codecs, but it's not the best quality. Like real broadcast level quality video is, is like an order of magnitude or two higher bandwidth than what's streamed to you even if you're getting like the 4K Ultra whatever on your Fire TV. And so that- the reason is bandwidth.
There's always need for more bandwidth, and there always will be need for more bandwidth. And so the more we can do to- Mm … to make it cheaper and more reliable and just more prevalent, we're only gonna be benefiting our customers. Corey: No, like you have the, the news anchor trucks that show up at the scene of something going on. Yeah.
They have the big satellite dish on the roof. It's like, do you think that's because they haven't quite figured out that you can tether to your cell phone? There are some very dedicated, very concerning bandwidth requirements around this. Yes.
And it's one of those areas, networking more so than many, but complexity passes everywhere, and it just slips below the baseline surface of awareness. Thank you so much for being as generous with your time as you are and explaining how some of this, this magic all works. If someone's watching this or listening to this, depending upon their point of view, and they wanna learn more about how networks work and how infrastructure happens, where should they start? Where should they begin learning about this, this secret thing that still drives our entire world?
Matt: Yeah. It, it's, it's again, it's going back to the, the raw material. And then there's a book written in I think the '70s or '80s. It's The TCP/IP Illustrated Book by Stevens.
It's where I'd suggest you go. Like that's where I send anyone who's wants to get into networking and learn something new 'cause those same fundamentals like you talked about of all those standards They still matter, and the, the entire world is built on y- you know, versions of them that have, have evolved over time. But that fundamental skill set will, will take you a long way, and you have to have it to really understand how these things work. Corey: And we'll of course put a link to that.
I love that book myself. It's a great book. Though I will say the, some of the bandwidth references in there of like, "Oh, that's a really fast 10BASE-T connection there," that might not have aged super well. Just Matt: shows you how far we've come.
But yeah, I, I just had someone, uh, actually one of our data techs this morning was messaging me asking me, "How do I learn about networking, Matt?" That's exactly what I sent him. Yeah. "Go get this on Amazon."
Corey: I have no idea if it's still up, and I'll check it. If not, my apologies, but there was a great site, WarriorsofTheDotNet. Mm-hmm. Was always a video that explained in a like five or six minute cartoon approach of how packet switching routed networks work in an extraordinarily accessible way.
I hope that's still out there. Yeah. That was always a fantastic primer. Like, yeah, some of their terminology choices you can argue with, and like, well, there's a lot more to it there.
Yeah, but you're going from zero to one, and you're giving people a chance to see what's next. Yes. And, huh, that's curious. I wanna dig deeper.
The thing that actually tipped me over the edge was subnet masks. Okay, it's this weird number string, and all I know is that when I get it wrong, some things work and other things don't, and I don't understand it. Maybe it's time I stop hand-waving over it. Matt: Yes.
Corey: And for my sins, I learned how networks work. Yeah. 'Cause I didn't know how networks work, and I didn't believe it was possible. And now that I know better, I do not believe that networks work, and I'm astounded that it's possible.
It feels like it works in spite of itself, but clearly it does. It Matt: is tremendously complicated, but it's, uh, it gives us a lot of pleasure to, to make these things simpler for everyone. Corey: Thank you so much. Matt Rader, VP of Global Networking at AWS.
I'm Corey Quinn. Stick around.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.