The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/DevOps Accents
DevOps Accents artwork

#69: Is the Internet Really Decentralized?

DevOps Accents · 2026-01-25 · 46 min

0:00--:--

Key moments - from our scoring

Substance score

44 / 100

Five dimensions, 20 points each

Insight Density9 / 20
Originality9 / 20
Guest Caliber8 / 20
Specificity & Evidence10 / 20
Conversational Craft8 / 20

While TCP/IP and foundational internet protocols were engineered to be decentralized - with distributed DNS root servers, multiple routing paths, and no central ownership - the modern internet has become concentrated in the hands of a few major players. Hosts Leo, Pablo, and Kirill discuss how services like Amadeus (which handles 50% of global flight and hotel bookings), Cloudflare (serving 70,000+ edge locations), AWS (especially the overloaded US-East-1 region), and Let's Encrypt (securing 50% of all HTTPS traffic) create single points of failure. When these middlemen go down, entire systems collapse despite the underlying internet being technically up. The episode uses examples like the Spain power outage affecting global payment systems, the AWS US-East-1 October incident, and the complex supply chain behind chip manufacturing (TSMC, ASML, Zeiss) to illustrate how infrastructure is hyperconcentrated. The discussion covers practical strategies for disaster recovery, regional redundancy, and the trade-offs between cost and uptime requirements for different business types.

Key takeaways

  • →Even though internet protocols are decentralized, the applications and services running on top are increasingly centralized through cloud providers and middleware companies, creating hidden single points of failure like Amadeus, Cloudflare, and AWS regions.
  • →AWS US-East-1 remains the most overloaded and incident-prone region due to legacy adoption, and avoiding it in favor of other regions like Frankfurt can provide resilience.
  • →Disaster recovery strategy depends on business criticality and acceptable downtime; most applications can tolerate hours of downtime, while critical systems like banks and hospitals require active-passive or active-active redundancy across regions.
  • →Certificate authorities like Let's Encrypt control 50% of all web encryption, and outages cascade globally because they sit between DNS and users, affecting edge networks and all downstream services.
  • →The complexity and concentration of supply chains extends beyond software to hardware (TSMC, ASML, Zeiss), meaning geopolitics, regulation, and single-country disasters can disable critical infrastructure worldwide.

Topics in this episode

CloudflareTSMCASMLLet's EncryptAmadeusAWS US-East-1ZeissDNS root serversCertificate authoritiesEdge computing networks

Questions this episode answers

What happened during the AWS October incident and how can companies protect themselves?

US-East-1, AWS's oldest and most overloaded region, partially controls authentication for all AWS services globally. Companies can protect themselves by avoiding US-East-1 and deploying to less centralized regions like Frankfurt, or by maintaining disaster recovery sites in separate cloud providers with acceptable recovery time objectives (RTO).

Why does Cloudflare going down affect the entire internet if it's supposedly decentralized?

Cloudflare operates 70,000+ distributed edge locations worldwide, but a single bug or outage in their centralized control plane can bring down all of them simultaneously. They sit between users and DNS, making them a critical dependency despite having geographically dispersed servers.

How much data and infrastructure redundancy do companies need across regions?

It depends on acceptable downtime and business criticality. Non-critical services may tolerate 5-24 hours of downtime, requiring offline backups and manual recovery procedures. Critical services like banking and payment systems require active redundancy across multiple regions with failover in minutes.

What companies represent hidden single points of failure on the internet?

Amadeus (50% of flight and hotel bookings), Let's Encrypt (50% of HTTPS certificates), Cloudflare (massive CDN and edge computing provider), and Visa/Mastercard (90% of payment transactions) are major middlemen that can bring down entire systems despite the decentralized internet design.

How do semiconductor supply chains create internet infrastructure risk?

Chip manufacturing for AI and cloud infrastructure depends on TSMC (Taiwan), ASML (Netherlands), and Zeiss (Germany). A geopolitical event, regulation change, or natural disaster in any of these countries could interrupt GPU production and disable global AI and cloud infrastructure.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

9 / 20

The episode has a handful of genuinely interesting structural observations - particularly the TSMC/ASML/Zeiss hardware dependency chain and the Amadeus travel-platform single-point-of-failure - but these are buried in extended dopamine-addiction metaphors, political tangents, and conversational padding that dramatically dilute the idea-per-minute rate.

If Amadeus has a bad day, it's not only you can't buy a ticket, it can turn into Airlines can't access system, they need four reservations, airports struggle to print one, boarding passes and even baggage handling can get weird
to manufacture Those advanced chips, TSMs, they need uh, extremely specialized machines. And those machines are produced by a single company, asml, um, based in Netherlands. And to produce these machines they need uh, very complex, very precise optical systems. And those optical systems, they are produced by a single company, Zeiss from Germany

Originality

9 / 20

The episode's best original framing - 'we didn't lose the centralization. We outsourced it' - is a clean and memorable synthesis, and connecting physical semiconductor supply chains to internet resilience is an underused angle; however, the bulk of the discussion (DNS fragility, CDN dependency, multi-cloud strategy) covers extremely well-trodden ground without a genuinely fresh angle.

the Internet is kind of decentralized at the protocol level but we centralized at the um, convenience level. So we didn't lose the centralization. We outsourced it
lots of the benefits we get from the cloud, it's actually really easy to get all of them also with open source software on your own servers, uh, these days. But just for convenience, we committed to pay the premium

Guest Caliber

8 / 20

There are no external guests - the episode is three co-hosts who run a small cloud consultancy (MKiDev), giving them genuine practitioner credibility but at modest scale; Kirill provides the most technically grounded analysis while Pablo frequently defaults to the dopamine metaphor rather than operational depth.

mkdef M was not down because we are in Frankfurt region. Because it also makes no sense for us to be in USS1 region. So that's already a good um, idea just not to use the most overloaded cloud providers region
My name is Leo Sushiev. I'm a co founder of MK idev where we help teams design and run cloud platforms

Specificity & Evidence

10 / 20

The episode earns credit for naming real companies, services, and rough figures - Amadeus, TSMC, ASML, Zeiss, Let's Encrypt's ~50% certificate share, Cloudflare's 70,000+ edge locations, DigitalOcean/Hetzner/Fly.io as concrete alternatives - but almost no hard data with sources, timelines, or cost figures are offered, and key claims like the Spain power outage lack dates or quantified impact.

last encrypt, this company, they have 50% of all the certificates. So it's the certificate authority of 50% of all the web pages
Cloudflare has I think more than 70,000 Edge locations, if not more

Conversational Craft

8 / 20

Leo occasionally drives the conversation toward practical operator territory with useful follow-up questions, but there is minimal pushback on vague or questionable claims, Pablo's political and dopamine tangents go largely unchallenged, and the episode never reaches genuine productive disagreement or forces a guest to defend a position under scrutiny.

And what if Frankfurt is down?
do people waste money on this uh, resilience theater and if they're like a cheapest real win that everybody should adopt

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker C42%
  • Speaker A31%
  • Speaker B25%
  • Speaker D2%

Most-used words

internet40cloud33cloudflare22decentralized20sure17problem17happen17service15data15part14system14single13infrastructure13application12example11happens11

Episode notes

The internet is “decentralized”… until one cloud provider has a bad day. This time on DevOps Accents, Leo, Pablo, and Kirill break down where decentralization actually exists, where we’ve centralized by convenience, and why outages feel inevitable in today’s cloud-driven world. In this episode: What “decentralized internet” really means - protocols vs. reality Why major outages (Cloudflare, AWS, DNS) expose hidden centralization Single points of failure: clouds, CDNs, certificate authorities, and power grids From infrastructure to geopolitics: how political borders affect the internet Disaster recovery trade-offs: cost, downtime, and realistic expectations for teams Practical advice: how much resilience is enough - and what not to blindly trust

Full transcript

46 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Internet is decentralized, sure, but here is the part that messes with your head. Even if the protocols are distributed, uh, the stuff we depend on is often centralized in places you have never heard of. Let me give you one example that most people don't know. There is a company called Amadeus. They are based in Spain. And if you have ever booked a flight, booked a hotel, checked in online, or basically tried to go on vacation like a normal person, there is a good chance you touched Amadeus without knowing it. And the scary part is not that if the down, uh, your booking is not working. The scary part is how deep it goes. If Amadeus has a bad day, it's not only you can't buy a ticket, it can turn into Airlines can't access system, they need four reservations, airports struggle to print one, boarding passes and even baggage handling can get weird because there are systems in the chain that help plan and optimize how bags get loaded, basically like Tetris. And a lot of ground operation depend on software like that. So you can end up in this absurd situation where the Internet is technically up, your phone has signal, the airport has electricity, but the trip still collapses because one invisible platform in the middle of everything is having an um, incident. And that's why today's episode is about one does decentralization actually exist? Or did we just outsource Internet to a, uh, few giant platforms, cloud, CDNs and um, middlemen? No one notices until they break. We'll break down what decentralized even means, why one provider outage can create a domino effect, and most importantly what normal companies can do to protect themselves from single points of failure. My name is Leo Sushiev. I'm a co founder of MK idev where we help teams design and run cloud platforms that are faster, more reliable and more resilient with the right architecture, automation and platform engineering choices from day one. And this is DevOps accents podcast where together with Pablo and Kirill, we strip buzzwords down to reality. Let's get to it. Welcome to DevOps Accents, a, ah, podcast

Speaker B: on everything around DevOps, public cloud and

Speaker A: cloud native topics with your hosts Pablo, Leo and Kirill. Hey guys, uh, I, um, last time, no, not last time, it was during the end of the year recording we touched on this topic a little bit. I mean the um, uh, decentralized Internet. And uh, today we decided to cover this uh, in this episode, uh, link to the previous episode in uh, the description so you have to refresh your memory and uh, people love saying the Internet is decentralized. It's Cool. It sounds nice, makes you feel safe. And then I know one cloud provider has a bad day and suddenly your Spotify is down, your discord is down, and some random SaaS dashboards are also down. And you're sitting here like, I don't know, like, uh, that's a true story, uh, about me. I was like, is it my wi fi dead or did I even pay for my Internet? Or what happened? Or is it some, uh, cleaning service tripped over a cable at some data center and suddenly everything down? Um, so the main question is, is the Internet actually decentralized or did we just centralize it into the few clouds? But before we start yelling at clouds like old men, uh, I think we need to agree on what decentralized even means, because I think we are mixing concepts probably. I certainly do. Um, so this is my confusion. When people say the Internet is decentralized, do they mean the protocols are decentralized? Like, no one owes tcpap? Uh, I mean, on paper, the raw Internet, to me, the concept is like roads, uh, many routes, many networks. And, uh, nobody is CEO, uh, of roads. Right? So nobody owns the Internet. Uh, but it maybe sounds simple and kind of naive, but when people say the Internet is decentralized, what do they usually mean? And what part of that is actually true today?

Speaker B: Okay, I think that first, if the people is not using, like, uh, mental health problems, the Internet to use the instantaneous thing for one hour or three hours, nothing happens. You know, because these days when Internet is down is because the people cannot use X, they cannot use Instagram, they cannot use this kind of tools. That. Because this is the people who discover first that Internet is not working. Apart from that.

Speaker A: Yeah, it's down detector. Right? So you go on the website down detector and all that's popular services. Oh, no, it's not working. It's not working.

Speaker B: You're using instead. And then it's not the scroll, infinite scroll is not working anymore. So I don't have my dopamine. So then like a, uh, drug dealer, I need my drug dealer. Where is my drugs? You know, I need to more dopamine. But okay, this is different. So the point is that I think that the people imagine that Internet is. So nobody is able to stop Internet, because this is one of the mottos that many times has been told by people. There is no way that you stop Internet because there is no central point that you can do something to stop all the Internet. And we discovered that. Okay. And not that we discovered. So there are many, many points that are Completely, uh, a problem. In case that. I don't want to talk about DNS now, but okay, there are many other points that in case that this part of the Internet goes down, this part of the global infra goes down, many things are going to be down.

Speaker A: Mhm.

Speaker B: Everything these days is more and more and more interconnected because at the end one application use tools of another companies and these companies use tools of another companies and this tool use tools of another company. And if something goes down, this is like a domino, many pieces of these other components are going to be down. And at the end I don't get dopamine. So it's like a summary.

Speaker C: Uh, yeah, but this Instagram part is, I mean people notice it because there are maybe like billions of people using this service. But when Instagram is down or like AWS is down, then there are many other things that are down. Like your local municipality website could not be working because of Cloudflare, but it's not going to make news because okay, who cares about your municipality website outside of your city? But I think Internet like okay, decentralized means uh, there is no center of the Internet, so there is no single system or point that can go down and then the whole Internet is down. I think the Internet originally is super beautiful in this way because everything was actually engineered and built to be super decentralized like DNS. There is no single DNS server in the world that if you bring it down there is no DNS. There are this like, I forgot how many? 12, seven. I always forget huge root DNS servers that are uh, intentionally on um, different continents and countries. And then everything else goes like a tree from them. So it's really hard to kill DNS on planet Earth then the routing as well. So you can like M. Networking starts at this private network, not with the interconnected global network. So then everything is built with lots of small networks that are decentralized. Each of these networks can have their own routing DNS and so on. And then at some point they all connected to the big Internet. So the foundational, uh, technology is not centralized. It was built with decentralization in mind, but it is now centralized in many ways. Like it could be on local level, politically centralized, because it is obvious, uh, it is possible to just isolate or turn off the Internet for one country. Right? Because then the political boundary is your Internet boundary and you can see that, okay, like in countries like North Korea, um, it's just super isolated. And then in countries like China and Russia, it's more and more isolated. And there is like somewhere a switch where someone can say no more. Global Internet. It is so decentralized because you're not depending on any server in USA and you can just still use your Internet in Germany. But decentralization does not matter for someone who lives in a place where someone decides to isolate you. But what the problem now is that more and more of Internet higher upper levels is owned by private companies like cloud providers. So instead of having this fully distributed decentralized system, we have big chunks of this system owned by a private company that builds their own software on top of this decentralized part. And if some um, developer in this company just makes a bug, then say 10% of the Internet is down because we go through these upper levels.

Speaker A: That's what I wanted. Like, that's what I want. I uh, mean uh, when I say like it doesn't feel like it's decentralized and getting back to the road, uh, analogy, uh, analogy, uh, it feels like we have just instead of many roads, we have three giant highways, uh, two landline operators and one GPS application and that's it. And any of them glitches, uh, we all end up in a ditch. So this is how it feels. This is the issue of, I don't know, monopoly, or is it just uh, the nature way, how it should work on technical level?

Speaker C: I think the technical level is not a natural way because it was not like this before. Um, you could think about it more maybe of modern payment systems because you have two companies that control how many than 90% of all transactions within MasterCard and they are private companies. And if servers of these companies are down, then global payments are done. Or if these companies decide that you're not going to be able to use their network, you will not be able to use their network. Uh, so there is no decentralized foundation there. It's just the system was built initially as something super centralized for credit card processing.

Speaker B: That is a quite example because the topic that Kirill is telling, there are many points like that. For example, these certificates authorities of those, if those certificate authorities are not working. So at the end nothing is working. But there is one interesting case that this one in Spain, the power went down. So this was a few months ago. And um, how long is that kidding? I thought I never stopped. So, but, but when this happened in, in, in Spain, the power went down. It's not only that the power went down is all the data centers of the companies went down. And the problem, the time that took to turn on everything, um, and this Makes us, for example, like Kira was telling before, Visa was not working. So you cannot pay anything. Ah, in many places even you cannot have an appointment with doctors and hospital because everything these days is digital. And then not only the external data centers that are in usa because if power is down, completely down. So at the end in one country there are many things that are not working in many other countries. So at the end everything is interconnected and everything is interrelated. And this is uh, a problem these days because if you want to be a bad guy, there are many themes that you can do. It's not only the DNS as before, many years ago everyone was thinking about, okay, DNS and the 12 or 13, I take down those and then there is no Internet for sure, there is no Internet. But just start to think these days and um, if you find some kind of tool, like for example Cloudflare, that most of the people doesn't know what Cloudflare is. But if Cloudflare is down completely, many companies are going to suffer and many other companies that are using the previous companies, DNS is like that.

Speaker C: But with DNS, if these root servers are down, the Internet is not going to be instantly down because it's rarely that your queries end up on one of these huge root locations, right? Because it's all super cache and distributed. Like if your DNS resolver is like 888. Did I say 4 8. So the Google one and then they have their own DNS cache is naturally like always fall back to the root one. Right. So it's still going to work. But with Cloud 4 Layer, the issue is that what they offer is something that it's hard to offer this kind of service, this edge system, edge content network, but now also edge computing network, um, without relying on DNS. Uh, so what they have is basically they sit between you and DNS records in a way. Uh, and then it's kind of like. So with AWS you have a data center usually or you have like regions and then you either reach out to one region or another and you can build the system in a way that it's kind of makes use of multiple AWS regions and Google regions. But with Cloudflare or any CDN provider, it's trickier because the bug on their end or some outage on their end can influence all of the global endpoints. They have all of the Edge servers. And this is what happened multiple times in the last six months is that if you think about like Cloudflare has I think more than 70,000 Edge locations, if not more. So it's like servers around the world. There's like one in Munich, one in Telde, one, uh, in Belgrade. And all of them are down at the same time. So by nature they built a super distributed system which is around the whole planet Earth. But the bug in their software can still bring all of this down.

Speaker B: For example, I was reading right now that last encrypt, this company, they have 50% of all the certificates. So it's the certificate authority of 50% of all the web pages. And if this one is down, completely down, and the root certificate disappear and cannot be recovered. So in this case you have 50% of the encrypted traffic, that all the traffic mostly is encrypted, is down. So you know that at the end you have many hiding spots that if something is down, you have no idea why, but everything is down. They say you don't know, but the people that is getting dopamine, they don't know how that is working. And then the problem is that then everything is down. And this is uh, a huge problem because it's almost impossible to know all the little problems that could happen because you can prepare whatever documentation about. Let's prepare in case that something happened. Or this one, or this one. But there are things that you cannot control.

Speaker A: There is like ah, one, um, I don't remember where I read that, but it was kind of surprising to me. And eye opening. It's not related to Internet directly, but it's related to artificial intelligence. And it has so many points of single failure. Do they call that like that or single point of failure? Okay, it doesn't matter. So, uh, for AI, we need GPUs, and uh, most of the time we mean Nvidia, uh, when we say GPUs, uh, their chips power most of the serious AI workloads running in clouds today. Um, but the thing is, Nvidia does not actually manufacture those chips. Uh, those chips that Nvidia uses, they are produced by single company, uh, tsmc, based in Taiwan. Yeah, they have different fabrics, they are located there and are controlled by a single company. And it gets even more interesting because manufacturer, uh, to manufacture Those advanced chips, TSMs, they need uh, extremely specialized machines. And those machines are produced by a single company, asml, um, based in Netherlands. And to produce these machines they need uh, very complex, very precise optical systems. And those optical systems, they are produced by a single company, Zeiss from Germany. So they are very like, they rely on each other. If one of them breaks any link in the chain, uh, has a serious problem, including geopolitics, logistics, regulation, even natural disaster. And bam, we don't have AI. It can happen in a minute and the supply chain will break. And this happens not only in um, infrastructure, not only Internet. So the Internet may be decentralized in theory, but the world it runs on absolutely is not. So the problem goes with its roots way deeper than uh, we think. So when shutdowns like the most recent one with Cloudflare happens it's. I don't know, it doesn't sound so surprising to me anymore after I read that. Or even in hardware they are so dependent.

Speaker C: Yes, because TSMC is actually a good example because okay, this cloud and Internet is not as severe with dependency on tsmc. But the reason more and more people use something like Cloudflare is because Cloudflare invested so much time into building this edge computing cloud, CDN system and so on. So it's very convenient and nice to use it. So the better they get with their product, the more people will want the usage. With TSMC it's even worse because it took for me like more than 20 years to develop all of these supply chain components like the lenses from those, the lasers from another German company, uh, and then components like from around, I think like more than 150 different suppliers involved in building one ASML machine. And then of course if you want to decentralize chip manufacturing, which China is trying to do it because uh, things you cannot get access to the latest uh, um, UV spectrum versions of ASML machines. You just need lots of time. The sad thing with the cloud is that it's actually not as complex as creating um, sub 3 nanometer semiconductors.

Speaker B: So

Speaker C: lots of the benefits we get from the cloud, it's actually really easy to get all of them also with open source software on your own servers, uh, these days. But just for convenience, we committed to pay the premium of the cloud provider and not manage infrastructure on our own. And the outcome for sure is that more and more workloads are running on these cloud providers. So then more and more depends on these cloud providers. Luckily we have three hyperscalers on that one.

Speaker D: There are no challenges that we couldn't overcome. Whether it is immediate infrastructure problems or planning a future project, we won't simply answer your questions. We become a part of your team to help you complete the mission. Our solutions consider the interests of your business and the combined expertise of the industry. As our staff is made up of more than a dozen experts in different areas who share decades of field tested experience and knowledge with you.

Speaker A: But the teams, if we talk about them, um, how they can, I don't know, defend themselves from what's happening. I know, let's try to focus on like particular examples. For example, with AWS on uh, October last year, what happened? Like there was a big incident, uh, and um, it started the domino effect and it was probably connected to DNS or something like that. You were uh, the thing that you were discussing before, uh, talking before Kirill. But when things like this happens and my service is running on that aws, I don't know, branch or whatever, is there a way for me to prevent this or to react to keep my services alive?

Speaker C: Only if you have DNS outside of aws. Because the issue was actually uh, what was done was US East 1, the very first AWS region. And lots of people still, especially in the US default to using this region. So it's the biggest, craziest region of whole aws. The one with the most outages because it's so overloaded and huge and still the one with couple of global services. So the whole authentication of AWS is still partially centralized in US East 1, even if you use any other of their data centers. So this is like mkdef M was not down because we are in Frankfurt region. Because it also makes no sense for us to be in USS1 region. So that's already a good um, idea just not to use the most overloaded cloud providers region that there is.

Speaker A: And what if Frankfurt is down?

Speaker C: If Frankfurt is down, then ideally you maybe have some disaster recovery site with deployment of your application elsewhere. But this gets to the discussion of is it okay for your system to be down? Because in the end you don't have to provide 99 availability all the time for your service or website. Things can be done. It's not like super critical for most systems. Of course if you're providing some software for hospitals, uh, then you need to think about these things a bit more. In most cases if all of AWS is down, you just sit and wait until they fix it. And it's in most cases even can be excused

Speaker A: again. Yeah, because in most cases, uh, the reason why we noted this is because a lot of major services are there and if some small service is down, whatever. But for example, if we can, let's imagine um, some servers, they had this and then they have a reserve like um, on some other service and something happens with US East 1 and then they automatically switch to another service, does it mean that they have to have all the service all the data, everything uh, duplicated into another service. And they have to pay twice for the same service just in case something of that happens. And they are so afraid to cut this dopamine stream, uh, for their users. I don't want this to happen. That's why they have to pay twice. Do they really pay twice or do they have like, okay, we have this reserve and when we need it we turn it on and it starts working.

Speaker C: It uh, depends on again on your disaster recovery strategy. Because every company is different this way. So again like for MKDEF website for example, disaster recovery strategy would be our website can be down for half a day, we would not die from this. For more important applications like insurance, banking, they have procedures how to recover from things like this. And then you have options where you say it's acceptable that it will take us five hours to recover from a disaster like this. And then you plan how you have this offsite backups and the system to spin up your environment in another region within five hours. If you say that your requirement is to be back in 30 minutes, then for sure you have a bit more in the standby mode running all the time to switch over to. Or probably for systems like Instagram or Visa, they have uh, two requirements of disaster recovery until like two minutes. So it's always distributed. There is like no single data center that can die.

Speaker B: But what Kirill is telling is true because at the end is how important is that your application is up? No, because at the end the important thing is money. So because if you have a beetle system, normally it's not connected to cloud or uh, in that case even you have another system locally, you know, in the hospital normally because vital they have their own data center plus connection to cloud. You need to have something like a ah, backup connection in case there's something happen. But if it's not a vital thing, the only thing is that it's money thing. So you are not going to make money these hours. But that's all. But the problem is that how much money is 10 hours for a big company? So could be uh, a big disaster. But at the end is that how much money do you want to spend then? If I spend this amount of money every month in case that one day something happened, I can be sure in 99% that nothing is going to happen. Or if I spend this amount, I'm going to be sure in 99.99% that nothing is going to happen. So it depends on how much you spend. You have a certain security that nothing is going to happen to your services, but nothing can give to you 100% security that nothing is going to happen ever. This is for sure, because, uh, a dinosaur can go to all the data centers and eat the services or, I don't know, Superman can throw a rock in all the water cables and then with the lasers in the ice and then can cut all the Americans like they did with the Nord stream. So they can do many things.

Speaker A: Yeah. And what are the chances something of this happens is 50% if it happens or not?

Speaker B: I don't know. Today, 16th of January. I was watching yesterday, uh, a podcast about politics and they were telling that only 3.8% of the year is happening and they capture Maduro. Maybe a war in Iran is starting and I don't know how many things could happen with Greenland. And today 16th of January. So we are in 2026. Everything could happen maybe one day. Trump, um, says, I don't want to have wires in the middle of the ocean, please cut off all of them. And then he cut all the wires. You never know. So this is the thing, this is what I mean, that you don't have 100% security that nothing is going to happen. And this is the reason that it's 99.9999 or something like that.

Speaker C: Signature is like with the cloud, you give up on being able to recover a lot because there are many things that you can do on your own and protect yourself. Because, okay, you can say the DNS server is down, but if it's your DNS server, you can just fix it. If you just said, uh, all of my applications are on cloudflare. If cloudflare is down, that's it. You build your environment around a single cloud provider. So if it's down, there is. You have no control, you cannot plan or anything. You're doomed.

Speaker A: Uh, exactly. And I still like to get down to the smaller, not smaller, but anyway, real teams and propose some kind of solutions or strategies for them. Because, uh, naturally you cannot buy, uh, a second house every time you have a disaster at your previous house or can just go multi cloud. So it's not an option for many. Uh, but realistically, what can we propose? Because I understand, as I understand that when a new infrastructure, um, is built, we have this discussion both as a customer or internally, uh, what other disaster recoveries we can propose and what we can expect and what the business consequences of something this happens and what kind of strategies out there. Uh, and maybe someone who listens to us, they could pick up some Ideas from there. So for example, of course, uh, being multi cloud in most cases, not the option. But do you have, I don't know, kind of a minimum viable service mode that you might propose that is

Speaker B: uh,

Speaker A: still alive when something happens elsewhere on the bigger cloud but you have a reserved version on a smaller cloud that doesn't cost you a lot but still works. You can still get payments, something like that. Or is it always all or nothing?

Speaker C: You can always do something. It's uh, about getting payments. If your payment processor is Stripe, then if stripe is down, it's down. You cannot do any of that, uh, decentralization. Okay, you can support Stripe and ADN in parallel, but then if Visa is down, then both of them are not working for processing Visa cards. I think like for small teams it's nonsensical even to go into this direction. Like, okay, what you can do is you can select maybe let's say three different uh, infrastructure providers. Like you take DigitalOcean, Head, Snare and Flyio. You make sure that all of them support some standard way to run applications like Kubernetes or just containers. And then you deploy your applications three times in all three of them and you create a DNS record that distributes traffic equally between these three. And then you figure out the data storage part because that's actually the hardest one in general is how do you make your database distributed across the three. But then at least you have maybe ideally also pick Syria that are within different governmental um, jurisdictions. Like one is maybe owned by Germany, another one is owned by us. Third one is owned by Japan. So you're also safe from this kind of legal implications. But it would be stupid to go this way if you're a new startup. Okay, if you're a startup that uh, has, I don't know, maybe you're a military startup. So it's super important to think about these things from the day one then for sure. But if you're building like, I don't know, AI food analyzer, then you would not go this way.

Speaker B: Target is a tool that can be used to analyze the food.

Speaker C: As for me, the issue that I realized lately is that, so I can see myself not using aws, I would be able to run pretty good solid infrastructure without a cloud provider, not just on VPs. Like I have skills, and now everyone probably also have skills with the ChatGPT and Claude and Gemini to configure your application to run in a couple of virtual machines with the load balancer and database. It's like it's Actually super easy to do. Like you don't need to have like 20 years of infra DevOps cloud experience to this. But with tools like Cloudflare I cannot actually imagine what I would do myself to get to the same result as what Cloudflare has. Because Cloudflare does not have any regions or whatever fastly. Any other Cloudflare is probably just the biggest one right now. But um, this kind of system because they're regionless by the design. So every feature that Cloudflare is building, it's not bound to physical location, it's everything runs on edge. When you deploy application to Cloudflare this Cloudflare workers, you actually have your application deployed to all of this like 70,000 servers of Cloudflare. When you use a database of storage, it's also distributed across the whole globe. So when my user goes to my application hosting Cloudflare, it's processed maybe like 5 km from the user right there on edge. And it's super fast and convenient. And if I expand to any market like I built something for you and now I expand to US with Cloudflare, I don't care because it's already global uh, by design. So I don't need to spin up another data center in us. It's just going to be processed by edge locations in us. And for this kind of thing I actually kind of imagine what I'm going to do with open source and buying VPs because then lots of work to get to the same uh, point and lots of uh, open source distributed databases that I would need to configure to make it work and learn how to use them at a planet scale.

Speaker A: Another stupid idea that I had is probably might work is hosting uh, microservices on different clouds so that if one shuts down, uh, the whole application still runs for you, but maybe some part of it does not and you still have a chance to put some message. So I'm sorry, we are experiencing some technical issues. That's why this part of our application is not working right now, something like that. And you just distribute all that between three, four, five, uh, cloud services and you just kind of secured.

Speaker B: The problem with this idea is that at the end you're talking about microservices that they need to talk each other and if every one of those microservices, uh, is a part of the application, imagine something stupid like you have a web page, um, and you have five widgets and every one of the widgets is connected to a uh, microservice only to code this area and then every one of those is in a uh, different data center. First ah, there is a delay and second there should be a communication so you need to talk between the microservices so they need to have an information about what is happening and I don't know the data needs to be in some place and um, later on is the privilege. And so this idea, so of this decentralization talking about microservices is given several problems. The first one is how big is your application? Because at the end if you're a startup this idea is absolutely wrong. It's a bit crazy. But not only that is the idea of how are you going to communicate between those service is needed to have five times the same infra in all the places. What is the performance that you're going to get? This kind of question is the one that are going to come to you when you design that. And even the problem is that you have now imagine five data centers or five cloud environments that if any one of those has a problem, you're going to show that there is a problem. And if you have everything in one cloud, of course if there is a problem everything is down. But statistically you have more options to have a problem if you have five times the same infra that if you have only, only one. Plus you need to have a team that is able to work with five different cloud providers.

Speaker A: So I think the main problem is that we don't know um, a real single point of failure. We mostly have. We are just guessing things like that may happen, something else might happen and we just. So it's just a matter of how prepared are we and how we calculated all possible uh, disasters that might happen to our infrastructure. But then again in this case do people waste money on this uh, resilience theater and if they're like a cheapest real win that everybody should adopt if they um.

Speaker B: Of course because is not your money. This is the difference between a startup that the decisions are made by the people who put the money or uh, a company that the decisions are related to a budget that someone decide. So if they give to you 1 million budget you say okay, I can expect 1 million. If you are creating your company and you have 30k for sure that you decide. I am going to try to spend one every month or two, uh, thousand every month because then in this way I can live for 15 months. So and there are many, many configurations that are crazy for sure.

Speaker C: I think this, uh, on technical level there are many things you can guess in advance and Protect yourself. So I'm actually call this not so much concerned about the technical part because if there is time and resources to implement technical measures to increase availability and faster disaster recovery and so on, then it's always doable. But one of the issues with decentralization we have is also the on kind of a political level. Right? Because if everyone in Germany is using AWS for the workloads and then okay, first of all there is a risk that all of your IP is going to US because CS company so us can say this super cool startup is incredibly cool AI product is hosted on US east one or even Frankfurt. We can just get off all of their AP because we have quantum computers to decipher all of the encrypted volumes. Or US can say unless Denmark gives us uh, Greenland, then we're going to shut down all of AWS workloads owned uh, by Danish companies or I guess Dutch companies is the right. And this is like already not on a technical problem. Right. And that you never know when it's going to happen. And it, it's going to happen for sure. You can see it with many countries and random things where either your government decides to block some service from other country or other country decides to block their service for your country. And of course one easy answer is like okay, we need to have digital sovereignty and then everything in Germany we need to have like a German uh, cloud provider with German infrastructure owned by Germany and so on. But then the downside of it is that if there is a new uh, leadership in German government that says okay, very good that we have everything in our German cloud provider, now we can fucking isolate the whole Germany from external Internet then it's also an issue such can happen both ways. So it's good to be independent from other entities outside of your country. But if you're fully independent then you always risk to get decentralized out of global net.

Speaker A: Yeah. So my uh, takeaway right now is the um, Internet is kind of decentralized at the protocol level but we centralized at the um, convenience level. So we didn't lose the centralization. We outsourced it. Uh, before we wrap this up, uh, uh, my last question would be uh, is there one dependency you would stop uh, blindly trusting? Not because it's bad, but because it's too important to trust blindly. What would it be like a ah, final advice to those who listens to us.

Speaker B: Don't be a drug addict. Dopamine so you can stop using and abusing those tools.

Speaker C: My advice is to have I'M not a fan of doing this whole self hosting and offline backups of my own homework. Homelab is open source software. Even though maybe I should but, but the very least, don't bet everything on single cloud provider. Especially like on a personal level if you really like Google, like Gmail and uh, Google Drive and everything, still make sure that you have a backup of everything to Apple or Microsoft or whatever. So if you have all of your photos in icloud, make sure to back them up to Dropbox or box.com or anything else. Just think about like another good cloud company that offers similar services and duplicate everything there at least the data.

Speaker A: All right, thank you. So this is us, uh, I think this is all that we have to say on that. And if you're thinking about if your infrastructure is resilient enough or did you do it write uh, if you have any failure uh points in your infrastructure and some shortages, you know how to call all the link in the description, make sure to check uh, those we can help you with making your infrastructure even better, even resilient. And when something happens to cloudflare, you make sure that you stay online. And also all the links in the description, uh, I will put there uh the links to the disaster that happened to Internet in last years to refresh your memory and also the link to subscribe to our MKDF Dispatch newsletter. Uh, every two weeks we post uh, our thoughts on um, things that happened. We share everything for you not to miss and of course don't forget to be subscribed to DevOps accent and give us likes and we will see you in two weeks. Thank you guys for coming and thanks you for listening to us.

Speaker B: Thank you.

Speaker C: Thank you.

Speaker D: Do you think your project infrastructure is well set and maintained? We know for sure there is always room for improvement. If you are uncertain where to begin, let's first do an audit of what you already have. We will review your setup from every angle, performance, cost, security, high availability and automation and provide you with a detailed roadmap of which direction your infrastructure should go. Generate concrete tasks for you to implement or even take on your infra entirely if you let us of course.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysisTraining Data · on TSMC95 / 100
  • Eric Ries on Why Good Companies Go BadPodcast Archives · on Cloudflare92 / 100
  • OpenAI launches Daybreak with Cisco, Cloudflare and CrowdStrike, Vapi wins Amazon Ring as it raises $50M Series B, JPMorgan picks Mistral as $430bn sovereign-AI rival, Isomorphic Labs banks $2.1bn led by Thrive CapitalThe Daily Marketing Brief · on Cloudflare88 / 100
  • Conversational Commerce with an All-In-One Booking Engine | with Adam DeflorianGAIN Momentum · on Amadeus86 / 100
  • 270: How We Actually Use Claude to Run Our Businesses (With Real Workflows)Creator's MBA · on Cloudflare85 / 100
  • Battle for the AI Data Center: Deep Dive on the Semiconductor SupercycleTechSurge: Deep Tech Podcast · on TSMC84 / 100

More from DevOps Accents

All episodes →
  • #68: The Current Reality of AI Coding Assistants
  • #67: 2025 Year In Review
  • #66: Is Kubernetes an Engineering Choice or a Must
  • #65: How Tech Hiring Really Works Now?
  • #64: Chat GPT is making you dumber (not a clickbait, there is a research for that)
Explore the best B2B Engineering & DevTools podcasts →
All DevOps Accents episodes →