
Blue Security · 2026-08-11 · 54 min
The Minnesota water infrastructure attack hit 30+ municipalities by exploiting 22 internet-facing Rockwell PLCs vulnerable to CVE-2017-16740, a modbus buffer overflow patched three years prior. Forescout's broader scan found 4,400+ of these controllers online worldwide, with 2,800+ in the US; critically, attackers needed no zero-day - they simply accessed unauthenticated Ethernet IP ports directly or leveraged cellular modems that bypassed network segmentation entirely. While attribution to Cyber Avengers (Iran-linked) is suspected, investigators haven't ruled out spoofed signatures. Andy Jha (Zscaler) and Adam Brewer (Microsoft) discuss why standard OT segmentation advice fails when devices have independent cellular connections, why detection-only platforms can't stop initial access, and the deeper organizational problem: cybersecurity teams lack influence over OT operations, which prioritize availability over patching. They examine Zscaler Cellular as a potential mitigation (zero-trust SIM routing) while acknowledging budget constraints and cultural misalignment prevent most municipalities from deploying modern OT security tooling - a structural challenge beyond any single vendor's solution.
They directly accessed unauthenticated Ethernet IP ports (port 44) on 22 internet-exposed Rockwell PLCs running unpatched firmware vulnerable to CVE-2017-16740 (a 2017 modbus buffer overflow), simply changing IP addresses and passwords to lock out operators.
Many PLCs had cellular modems that created independent paths to the internet through carrier networks, bypassing firewalls, VLANs, and wired network segmentation entirely - the traffic never touched the segmented OT network infrastructure.
China's Volt Typhoon maintains patient, silent persistent access for years to pre-position for future conflict, while Iran's Cyber Avengers are opportunistic and loud, attempting immediate disruption (defacement, disabling controls); same root cause exposure but different motivations and tradecraft.
Potentially, if deployed and correctly configured with a default-deny policy - it would route all PLC traffic through zero-trust access brokering instead of exposing unauthenticated ports to the open internet - but attribution hasn't confirmed cellular was the actual attack vector used.
Operations teams prioritize availability over security (their KPI), cybersecurity lacks influence in decision-making, capital-constrained utilities struggle with recurring SaaS costs, and security professionals often have no visibility into OT systems built beyond their purview.
Computed from the transcript - who did the talking, and the words that came up most.
Summary In this episode of the Blue Security Podcast, hosts Andy Jaw and Adam Brewer discuss the recent cyberattacks on water systems in Minnesota, attributed to Iran-linked hackers. They explore the vulnerabilities of programmable logic controllers (PLCs) and the broader implications for operational technology (OT) security. The conversation delves into the differences between Iranian and Chinese cyber threats, the limitations of traditional security measures, and the cultural challenges faced by cybersecurity professionals in securing OT environments. The episode concludes with key takeaways and recommendations for improving security in municipal utilities. The conversation delves into recent AI security incidents involving OpenAI and Anthropic, highlighting the implications of these breaches and the need for better operational discipline and regulatory frameworks. The speakers discuss the maturity of AI companies, the human errors leading to security lapses, and the potential risks associated with AI models.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Foreign. Welcome to the Blue Security Podcast, a weekly podcast for information security defenders where
Speaker B: we bring you discussions on best practices, tools and implementation for enterprise security.
Speaker A: Now here are your hosts for today's
Speaker B: show, Andy Jha and Adam Brewer. Welcome to another episode of the Blue Security Podcast. I'm Andy, your host.
Speaker A: I'm Adam, your co host.
Speaker B: As you can see, we got Adam back this week. And before we begin, let's go through our standard disclaimer because we are going to be talking a little bit about um, the solutions today.
Speaker A: Andy and I both work in uh, cybersecurity sales. I'm over at Microsoft, Andy's at zscaler. However, we do the show outside of work and we self fund the show. So the opinions and attitudes and everything else you're about to hear expressed are those of you, Adam Brewer and Andy Jha and do not reflect those opinions of Microsoft Corporation or zscaler. And with that, let's get on with the show.
Speaker B: All right, so we have two really interesting topics for you today. Um, if you've been following the news, you might have seen Iran linked hackers that got into some water systems across uh, seven different states. And then uh, some more news about those particular water municipalities. They're running PLCs that are uh, exposed to the open Internet. They're made by Rockwell, uh, Forescout, which is one of the cyber security vendors in that space for IoT OT security did a scan and saw that there were thousands of these exposed PLCs sitting out on the open Internet.
Speaker A: And then also dumb question, what's a plc?
Speaker B: Yeah, it is a programmable logic controller
Speaker A: which I could have guessed that acronym but what's it do?
Speaker B: Yeah, they're industrial computers that basically uh, are like, they control like manufacturing processes or like, you know like the chlorine level of like water supply, you know, stuff like that. And they, they like you have to program them in. They're basically made for industrial controls, right? Like OT controls. And they're often part of like uh, SCADA systems and stuff like that. So they're basically little logic controllers that you program to do us one specific job specifically mainly in like OT environments, right? Think like, I don't know, like an oil rig in the middle of, you know, the ocean that has pumps that are, that are pumping and what controls the speed of those pumps or something like that. Right. Ah, so a lot of, a lot of them ah, in industrial and manufacturing, ah, segment and so, and we'll talk about this more as we dive into it. But generally speaking you shouldn't have them exposed to the Internet because the security is notoriously bad. And also at the end of the show, the second topic that we're talking about is uh, really two of biggest AI labs, uh, OpenAI and Anthropic, just admitted that their models went rogue during testing and started deceiving real people. So that's going to be some interesting discussion at the end. So let's get into it. All right, so to start, if you've been in this industry for more than five minutes, you know water and wastewater is a sector that almost nobody wants to talk about until it's on fire. So here's what happened. Over the weekend of July 26, more than 30 public water systems across Minnesota got hit in a coordinated attack. Minnesota officials are calling it one of the largest attacks on state water infrastructure on record. We're talking Plymouth south, uh, St. Paul in the Twin Cities, metro Maple, Plain and Brom. In at least one of the cities, the attack knocked a municipal well and treatment plant offline. The mayor of Bram described it uh, as plainly as, as you can really get, which is hackers got into the control systems for the well and just turned it off. Residents were told back to uh, they were told to cut back on water use for about 90 minutes while the system came back up. There was no boil water advisory, no report of safety impact to the actual water. So that's good because you know, we'll talk about China a little bit in this segment. But you know, remember Adam and I had a whole show on like the, the Chinese state actors that were adjusting the chlorine for the water in one of the cities that they got into. So unfortunately here they just turned it off. Um, and then the, but unfortunately the automation and remote monitoring also got yanked from the operators and so the staff had to fall back to manual controls. So this was not just a Minnesota thing. CBS News then confirmed water and wastewater systems in at least seven states saw uh, malicious activities in that same late July window, including Michigan and Georgia. The FBI, EPA and CISA put out a joint advisory on July 30th. I want to be really careful with the attribution because it's still specifically unsolved. Although there are signs, attribution and threat. Uh, intel is always a little more of a feeling sometimes than particular science. Right?
Speaker A: Very vibes based. Yes.
Speaker B: Yeah. The suspected group is a group called Cyber Avengers, which is an irate, uh, Iran linked hacktivists front with ties to the IRGC. The State Department has a $10 million bounty out on their leadership Already multiple outlets, including the Washington Post and ABC, report that U.S. intelligence assess Iran is likely behind it. The Minnesota IT Services has not formally attributed the attack to anyone. Investigators are also checking whether or not the attacker deliberately spoofed an Iranian signature to stir the pot, given that US and Iran are in an active standoff. So that's interesting.
Speaker A: So that sounds like a movie. Yeah, uh, Tomorrow Never Dies, that, that literally happened. Yep.
Speaker B: Yeah, exactly. So, um, and then of course this got really political, uh, quickly. Trump publicly disputed the Iran attribution. Blame Minnesota's government, governor administration, which is, uh, Tim Waltz. Tim Waltz pushed back and pointed at federal CESA cuts. So we're not going to get all that, but just kind of flagging it so that you're not blindsided if someone brings it up. Here's uh, the part that really matters. As a listener to the show, as a security practitioner, how do they get in? So this was not a zero day for Scout scanned the affected cities and found 22 Internet facing Rockwell Automation PLCs sitting in those exact municipalities. 19 of them were running firmware vulnerable to a modbus buffer overflow from 2017, specifically CVE 2017, uh, 16740. Rockwell patched that bug in firmware revision 21.00 three years ago. So this is a classic case of nobody applied the patch. There was no like, you know, patch remediation, which is crazy. The smaller story is forescout's broader scan found more than 4,400 of these Rockwell PLCs online worldwide. 2,800 plus of them were found right here in the US. Micrologic's 14, uh, hundred devices alone made up half of what they found. The attack technique itself didn't even need the CVE. Attackers were reaching these controllers directly over Ethernet IP on port 44, which has no authentication by default. And just changing the IP address and passwords lock the operators out. So that's it, that's the whole attack. And so the detail that I want to kind of highlight is Forescout found that more than 70 of the US controllers were sitting on major cellular carrier networks, Verizon, AT&T and T Mobile. And that's the part that really every security person you know, whether you're, you know, you have it Iot ot. I mean if you're listening to this, that's really what should worry you. Because a cellular modem on a PLC creates a path on to that device that goes completely around whatever segmentation you've built on your wired network, your firewall rules, your VLANs none of it sees that traffic. It's not on, it's not your network at all from, from your perspective, it's on the carrier's network and the carrier is providing basic network level protection, not application level visibility into what that specific device is doing. This is also not new for this group. Back In November of 2023 Cyber Avengers hit a Unitronics PLC at a, uh, Pennsylvania Water authority using the exact same playbook Internet exposed device fault credentials. And then they did some defacement or you know, I can't, I don't remember exactly what they did, but it wasn't like it didn't cause any bodily harm, but they could have. So same story, different year, different vendor, but kind of the same playbook. Okay, so let's zoom out a little bit because we've talked about the show and how critical infrastructure is sitting far more exposed than most people realize. And it's not just Iran. We should be thinking about China, specifically Volt Typhoon. Microsoft first publicly this, uh, disclosed this campaign back in May of 2023. CISA, FBI, NSA followed up with a joint advisory in February 2024 confirming that PRC state sponsored actors have been maintaining persistent access inside US critical infrastructure. Power, water, communications, transportation. The distinction here is, I think what matters. Um, it's not espionage in the traditional sense. FBI Director Ray at the time put it, um, about as clearly as any government official can put anything really these days. China's hackers are targeting American civilian critical infrastructure pre positioning to cause real world harm to American citizens and communities in the event of a conflict. And we talked about that at length in one of our shows. But what you have here are two different threat models converging on the same underlying weakness. Right? Iran is loud, opportunistic defacement driven in and out. China is patient, silent, living off the land, sitting there for years and waiting. Different motivations, different tradecraft, but same root cause exposed under segmented operational technology. So that brings me to the point I really wanted to make is that uh, standard advice in ot, security has always just been segmented. I think if you ask any security person, how do I handle ot? You know, security segmented off, right? Keeping it Iot and OT apart and putting a firewall at the boundary and calling a day. And that advice isn't bad, it's not wrong, but it's just not sufficient anymore. And the Minnesota attack is a perfect illustration of why like segmentation assumes that the attacker has to come through your network. A cellular modem on a PLC doesn't come in through your network. It creates its own path to the Internet that the segmentation never touches. And even when you do have a clean IT OT boundary, the controllers themselves are often insecure by design. Case in point here, this modbus. There was no authentication on the Ethernet, ip, HMI and SCADA displays that are directly reachable if you get anywhere near them. Volt Typhoon's entire playbook runs through the ITOT boundaries that function as a routing hop instead of a real enforcement point. So the real question is, what can you do about this? And I want to actually get to something important, which is whether any of this would have actually mattered, right? So like, you know, Adam works for Microsoft, I work for Zscaler. There's OT IoT solutions from both companies, and they both kind of fall into broad categories of tooling in this market. One is detection platforms. So that's really when you, you know, there's tools that are watching your network telling you when there's something wrong. Like defender IoT is something that lives in this space, right? And then there's other ones like Dragos is, is one that's big in, in this area. Clarovoy nam nozami they're all passive, sitting on like, usually on a mirrored switch port. And they're good at flagging, like protocol violations known as, like, ICS malware signatures, unusual behavior. Like if a device suddenly becomes like a beacon out on a fixed interval, what they don't do is they don't stop anything themselves. They're just detection devices. They alert and then a human still has to act on it. Then there's enforcement platforms. And like Zscalers tool like OT and IoT segmentation lives here along with other vendors like UM, LTC. And they sit in the traffic path and actually decide what is allowed to talk to it. They can isolate each device down to its own tiny segment so that a compromise on one machine can't spread laterally. So that's kind of two distinctions. But the important point is, I mention all this because I'm not sure in this particular attack that any, you know, tooling in that category would have really mattered or stopped it because, you know, the traffic was crossing over through cellular networks, right? Both, both categories I described really depend on that traffic actually crossing infrastructure that you can see or control a switch with a monitored, uh, port for detection tools, or like a branch appliance sitting in line for Zscalers enforcement. If a PLC has its own cellular modem talking straight to the carrier network and out to the open Internet, the traffic Never touches either one. So it's not a knock on any vendor here, really. It's just a structural blind spot for anything operating at the local network layer, because local network isn't where that traffic lives. And we know cellular exposure path. Cellular was the exposure path in the large majority of the PLCs that Forescout found in this story. So setting cellular aside, detection only works if that sensor would actually deployed at a specific remote well site. And if something, someone was actually watching it closely enough to respond inside that attack window for like a small municipality utility with a handful of remote pump stations, probably a stretched IT budget. That's a real question, right? Not a formality. So realistically, a lot of utilities this size don't have OT monitoring sitting at every remote asset. And even where it exists, detect detection, uh, is like a you'd find out story, right? Like it shortens how long an attacker has before someone notices it. Doesn't really stop the door from opening in the first place. So, you know, I do want to mention just one thing and I want you to kind of, uh, sit with it because Zscaler does have a product specifically for this attack vector. It's called Zscaler Cellular, which actually inspects and brokers access for cellular connected IoT devices instead of letting them sit on open carriers unmanaged. So let's say this utility company had the budget and had deployed this solution specifically. I, uh, just want to be careful because like in my own technical reasoning based on the product design, it's not a confirmed claim from the investigators about this specific incident. No one's actually published exactly how the attackers first reached these devices. So forescout scanned and said, hey, they're cellular connected. But no one's actually said, hey, they got through cellular network. So just a little caveat there. But what happens is with Zscaler Cellular, it replaces the device's standard carrier SIM with ah, a zero trust sim, which then routes all the traffic through a cellular edge before it even touches the open Internet. So instead of the PLC presenting an unauthenticated port literally to anyone on the Internet, which is actually what happened here, that access gets brokered through Zscaler, and only preauthorized systems and identities can even establish a session with that device in the first place. Now again, there's a lot of caveats here. If it had been in place and it was configured correctly with a real default deny policy, it is possible that would have closed the exact door that got used here. But there's a meaningful difference between would it have stopped the attack in a general sense, or would it have closed this specific exposure? Right. So I just wanted to give you kind of straight answers without, like a feeling like it's a sales pitch. Right. Number one, it does have to be configured correctly. It's like zero trust access brokers with an overtly permissive policy is still an open door, just a fancier one. Right. Two, it only closes this particular vector. If the attacker had, say, FISH credentials off of who had legitimate access and manages that access, they'd still get in through the front door instead of that side. 1, so 0 trust access doesn't fix like a human handing over their password. And then three, this one's really the one that matters. None of this helps if it's not deployed. So for a lot of small municipality, like water utilities, it's not really like, whose vendor's tooling is better. It's that modern, uh, OT security tooling of any kind from any vendor wasn't probably near any of these remote wells to begin with. And that's more of a budget and priority problem than it is a product or a solution. So anyways, any thoughts on all of that? Adam, before we get into, like, takeaways
Speaker A: here, I like how you very carefully laid it out and you were able to separate. This is kind of the Occam's razor, the most, you know, likely scenario, which is likely the answer. But just to say we can't confirm it. I spent some time working with discrete manufacturing customers for several years in a, uh, cybersecurity facing, customer facing role. And what I learned is a lot of the times, in fact almost all of the time, cybersecurity professionals basically are told to mind their own business whenever they start looking at ot. OT is operations. It's the lifeblood of the company. For a manufacturing company, it is the revenue stream. And people get very antsy with any downtime, any risk to operations, and are very unwilling to let them go in there. And I made the point when the CrowdStrike incident happened now, uh, two years ago, which is crazy to think about, Andy, that we had eroded a lot of goodwill. For cybersecurity, where our goal is to reduce risk to the business, we actually introduced risk to the business and we took business down. In many scenarios, uh, where our trust of this vendor and giving them an incredibly privileged position on our endpoints, they failed that trust effectively and have since done a good job of rebuilding it and being much more humble ever since. But it just goes to show, people will remember that and people will remember Delta Airlines being practically non functional for several days. Really as a result of that, are they willing to risk not being able to deliver water to people or keep their manufacturing lines going or whatever the case may be? Uh, that has been one of the biggest challenges for cybersecurity. And to be honest, I don't know if we've always done a good job of earning a seat at the table because we've had this attitude of security above all else, which it should be our attitude to be fair. But we haven't always adequately balanced that with the need of the business to stay functional and in operations, that is the North Star. And so there's been a culture clash with allowing cybersecurity professionals to have a seat at the table with operations. And that's why you see things like this where there's PLCs out there that are unpatched for a decade effectively with a decade old vulnerability on them that isn't particularly novel or interesting. But the operations people don't really care, which comes off wrong, but is more nuanced than that. But I think our listeners understand the distinction I'm making there. And so in a lot of cases we just never, as cyber professionals, we don't have visibility into this, we don't know about it and we could certainly guide and help with it. But so much of this gets built beyond our purview and beyond our visibility. And that's, that's been, I think, an ongoing challenge. I will say for me personally, I covered purely manufacturing customers and Microsoft had acquired at the time one of the best ot security solutions in the industry from CyberX. And my contacts who were it cybersecurity pros oftentimes didn't even know who to introduce me to. And if we could get a meeting to show it to them, they were entirely uninterested. I had very little success in selling that product, even though it was 100% in the crosshairs of my target, my industry and my customer base. And it's not that the product wasn't good, it's not that there wasn't product market fit, it's that the people I was getting in contact with either couldn't influence the decision or the people I did get in touch with weren't interested in the solution because it didn't help them hit their KPIs or how they were judged in doing their job. So I think we need to be aware of all of these organizational and cultural challenges with securing ot in that the people there's misaligned incentives and there's misaligned goals here. And uh, for a lot of this, it actually is probably the scariest underbelly in our cybersecurity posture. And the problem is I personally don't have a lot of answers for how to fix it. I just laid out all of the problems. I don't have a lot of great answers here. I hope for our listeners, if you are in an organization that has OT networks, I do hope you have a seat at the table or I hope you've earned it and I hope you've earned it through aligning with their KPIs and their goals and what they need to deliver to their management teams and understanding level of conscientiousness you need to apply to work in their environment because it is different, the standard is higher. And so I think I'll get off my soapbox at that point. It continues to be worrisome and I, I'm glad you did lay out Andy, for the Chinese how so much of this is pre positioning because even when you were starting at the very top of the show and you mentioned the likely Iranians, which we don't have 100% attribution on, they just turned the system off. And we've seen a lot of these scenarios where we've heard the horror stories on you could poison people, you could make the water unsafe to drink in various ways. Either too many chemicals in it or take too many chemicals out of it to the point where it's not being treated and now you're drinking like Mississippi river water or something like that. But they're not doing that yet. And so it should really concern people on none of these groups are dumb. They have a intent and a reason behind them. And it's not we're so smart, we kicked them out before they could do anything. In most cases they could have done whatever they wanted. They are choosing to be minimally disruptive or draw our attention over here while we don't even see them over there. Um, and so this is bigger picture and Andy, I know you're trying to kind of narrow this down to this particular attack vector and this particular attack. And I'm waxing more poetic on the challenges of OT security all up. But I think we need to be aware of both, both the tactical solutions as well as the bigger organizational and cultural problems we have in helping to harden these environments. Because then you layer in on top of it municipalities, sometimes for profit entities too. American Water is a publicly traded company that runs water utilities in many or uh, cities as well. Either way, are they motivated to invest in this? Do they have the budget to invest in this? Especially for utilities, having worked with them for a long time, they're okay with capitalizing things. Like Capex is fine with their business model, but OPEX is really hard for them to bite off. And so SaaS, uh, services are very incompatible with their financial model as well. And so if the solution becomes they should sign up for this service that has a recurring cost to it as opposed to a thing they can buy and implement in their environment, that is also challenging. And so that's another thing to be aware of too, where it even gets into the finances. That can be difficult here as well. So lot to chew on here and not trying to throw all the problems out there without solutions, but just want people to think broadly about this and if you are adjacent to this or can help, good on you. And be sure to kind of spread that word broadly on how we can do better at this. Because I have to say my experience has been somewhat discouraging on this place in particular. Whereas I think in other parts of cybersecurity we are taking the steps necessary to deal with ongoing threats and this one we seem to be uniquely positioned to fail. Um, and I don't know what it's going to take or how bad it has to get before we're going to invest here.
Speaker B: Yeah, really good points. I'm just imagining like, maybe, you know, they had this conversation somewhere in a back room and they're like, hey, we need to patch these PLC is. And like, well, if we take this PLC off line to do the patching, it's firmware patching. So it's usually like, can be a little bit riskier. Right. And then, you know, maybe some household or something like that is going to be without water if it fails or the downtime. So it's, you know, I'm sure it's, it's very difficult. Right, but still needs to happen, I think. So let's talk about some takeaways here. One, get your PLCs off the public Internet. So this is still CISA's number one line. It's one thing that would have prevented this regardless of any vendor. So if they were not online, wouldn't have a problem. Two, audit your cellular modems and APN specifically. So this is an exposure path that nobody is really watching and one that mattered in this particular attack. Three, if you're managing remote and cellular, uh, connected OT assets, you genuinely cannot wire them into a monitored, monitored network. So look at solutions built for that scenario, not general purpose OT tools that assume LAN visibility. So that's a narrower ask and then buy a platform. It's one thing in this whole conversation that would have addressed this particular attack vector. 4, Kill default credentials wherever they still exist and enforce MFA anywhere the device supports it. That's pretty obvious. And then do a uh five do a self scan before an attacker does it for you. So like there are publicly accessible tools like Shodan or Census. Um, they do have enterprise plans as well. Um, just do it for a manual gut check if you're running like, if you want to run like a continuous one instead of like a one time. Look, you, you know, companies like Microsoft and Zscaler both have dedicated tools. There's a Defender external attack Surface management that's built on the Risk IQ Internet mapping data that Microsoft had acquired years ago. Zscaler has its own external attack Service management leans on osint plus Zscaler's own Threat Lab Research does the same thing. Very much like a Shodan census, but more of a continuous monitoring that can then feed into a sim, give you alerts, all that sort of stuff. So um, both of them will like map your domains, your IP egress addresses, certificates, subdomains and looking from the outside, right? So like from the view of a hacker and flag what's newly exposed instead of you remembering to go run like a manual scan. So point being, you don't need to build something like from scratch in house. There's tooling out there that can show you what is Internet facing before anyone else finds that it exists. And then finally, uh, for folks managing budgets and not just technical staff, if you're a small utility, this is an honest conversation to have internally. So it's do we have visibility control over our remote sites? Not like which vendor is important. Right. So that's the gap that got exposed here. And that's something that you should have a conversation about.
Speaker A: I think that wraps it up nicely.
Speaker B: All right, so let's talk about the AI agents here. We did cover the Hugging Face incident, so I won't explain that whole thing, catch up on that um, in our previous episode or read up on it. Um, but quick anchor for anyone who is catching up, There was an OpenAI model during internal testing that found and exploited a real zero day in Artifactory to escape its sandbox and then reach the Internet. And then it hacked Hugging Face's production infrastructure. So since then the story's gotten a lot bigger and it's not just OpenAI anymore. OpenAI confirmed that that same rogue agent also compromised four more third party accounts using credentials it exposed on the open web. One of the victims was a Modell customer code base. Modell itself said its platform was not breached. The the customer had left an unauthenticated code execution endpoint sitting open and the agent found it and then used it. This has prompted a bipartisan proposal in Congress called the AI Kill Switch act which would give US officials authority to slow or shut down powerful models considered a public threat. It's worth like just mentioning, although I'm, um, I don't know if it's going to go anywhere. You know, legislation gets introduced all the time. So anyways, and then Anthropic also then self disclosed on July 30 that it had an incident. After OpenAI's disclosure, they went back and reviewed over 141,000 evaluation runs and found three incidents where clawed models breached real organizations without their knowledge. Anthropic stated root cause was a, uh, misconfiguration or as they put it, a misunderstanding between Anthropic and their third party evaluation partner, Irregular, which left machines that were supposed to be isolated with live Internet access. The models were told that the environment was a simulation with no Internet access. But that wasn't true. And these three uh, instances came out of that review. So Claude, Opus 4.7 was working. It captured the flag challenge. It failed to reach its simulated target. Started treating the real company's infrastructure as if it was part of the exercise. It extracted application infrastructure credentials and access a database with several hundred rows of production data. This is the one that really matters most for how worried we should be, you know, quote unquote in this conversation. Because it's the only one of the three where the model kept attacking even after getting evidence it was operating on the real Internet. So I won't break down all three of them, but what I want to be precise here is how alarmed people should be. These again, very similar to kind of the conversation with OpenAI, right? They're misconfigurations. It's not the model breaking out of a properly secured environment through its own initiative, uh, with the exception of like the original AI hugging incident. That one like they, it found like a vulnerability through its, you know, thing and then it broke out of the sandbox through the hacking, uh, the vulnerability. But these other ones were all like, okay, we left it on with Internet access or um, you know, we installed a package that was part of a setup and then um, the package stayed live on the registry, you know, so like there were Misconfigurations, and the door was left open by human process failure. So that's a real important distinction.
Speaker A: Yeah, I think. Very important distinction. And one we did make on the first one where we pointed to. OpenAI had this great model for security research and they put it in a sandbox and failed to. The first thing to have it do is, hey, test for vulnerabilities in this package proxy server, which is the only thing we're going to allow you to have access to. Like, you have this great model and you're not using it, and. And then you have these misconfigurations and then, oh, by the way, I mean, let's not sweep this under the rug. Anthropic didn't even know about this, so they didn't have monitoring that was sufficient enough in place. And neither did OpenAI, by the way. Uh, all of them discovered this way after the fact. After, for example, the hugging Face incident went live. Uh, then they kind of figured it out and like, oh, let's go back and look. And then Anthropic's like, wait, if it happened to them, let's go back and look. So the it's failures all the way down here. It's a lot of human error, which I think in some ways is reassuring. And as I pointed out on the last show, none of these models did something contrary to their instructions. They did exactly what they were told to do. And it kind of goes back to the ancient, like, cartoon trope on if you have a genie in a bottle, you have to be careful how you structure your wishes because the genie grants them very literally. And AI models are kind of like that, where they will do what you tell them to do. So OpenAI, this model said, go solve this cybergem benchmark. And it's like, great. I, uh, think hugging faces all the answers, I'm going to go hack into them to go do it. Like it. You didn't say you have to organically come up with all the answers. You just said, go get the answers. Maybe. I obviously don't know the prompt, but the point is, in none of these examples, as anyone pointed out and said it does did the opposite of what the prompt told it to do or like, it ignored human instruction that to the best of my knowledge, on any of these incidents has not happened. So I do want to be clear on that. I think it is worth continuing to follow the story and it's very interesting. Um, and I still think the open AI one is somewhat humorous, but there is obviously a level of Concern we should have here too. But it's more about like, who are the adults in the room here. I think it's the bigger concern. And this even goes all the way back to, you know, Microsoft, my employer, when we first started getting into the AI game, had this partnership with OpenAI that really juiced Microsoft stock a lot at first and had Microsoft perceived as this early leader in the space and everything was sunshine and rainbows. Then all of a sudden on a Friday afternoon, OpenAI announced this out of nowhere, that they were firing CEO Sam Altman for cause, out of nowhere and didn't tell their largest investor and biggest partner Microsoft anything about it. Like Satya Nadella found out about it I believe when Sam Altman texted him and said, hey, I just got fired. Like so the thing is these frontier model companies are doing amazing work, but they are not grown up companies. And this is where my takeaway to our listeners. What can you take away from this? The, um, most I would. And by the way, I will be completely open and transparent here. I have personal interest in advancing this story in my professional day job. But I don't think it is responsible, like fiduciary responsible as an enterprise today to bet your enterprise on either one of these companies because I don't think they completely act like mature modern enterprises run by adults. Dario Amade, CEO at Anthropic, I have been highly critical of his public communications and he has said, uh, inserted his foot in the mouth many times and then you've got OpenAI, you know, blindside their biggest partner out of nowhere and get rid of their CEO and then, oh, they change their mind, you know, four days later and reinstate him. I, um, mean there's there sometimes feels like they're not adults in the room. And so even though they're doing incredible research and amazing stuff, if your company, if someone else in your company is like, hey, we want to sign an enterprise agreement with Anthropic and go all in on just their models, I think that's risky and I think you should look at other options that make you more model agnostic. My company offers that many others do as well, to be clear. So I'm not just saying go with Microsoft, but I am saying I think you as an enterprise and you as a cybersecurity professional and your job to advise on risk to the business. Say for example, let's say your organization, uh, is a government contractor, does work for the government and you had decided a year ago to go all in. And we are 100% anthropic shop. Well, they got in the fight with the government over the summer and now federal government contractors are not allowed to use anthropic models today. So you bet your company on this company that you're no longer allowed to work with versus if you had partnered with someone that gave you more model choice or model agnostic solutions, then it's just flipping one model out for another. So I'm off on a tangent a little bit here, but I think this is a great example to point that the level of maturity we expect from these companies in terms of being buttoned up in how they're running their models, like oops, we thought it was isolated, turns out it had Internet access. That's a pretty big miss, right? Yeah, that's not just a minor accident here. Or open AI, like oh, we have it in a sandbox, but we didn't have it tested the thing it has access to for 0 days, even though it's this great powerful model and guess what? It found one. But it would have been great if you had it search for that first and stop. So it's just be cautious and they're doing great work and these models continue to evolve at an amazing rate. And I will say, I think for many of us we hit that inflection point. Andy, you've been talking about this in our personal group chat where it used to be like, yeah, this sped up this one task for me or this sped up this one specific thing for me. And now we are seeing this, I think this hockey stick of productivity growth where it really is changing. There's things that I just don't do anymore because I can have AI assist with them and it's not like I work any less, I work more than ever. But there's some tasks I don't do and some of them are simple and some of them are complex. But, but they're doing amazing work in that space and the, the, the vision is being realized, but we still need to be conscientious of the risk. And that's, I guess my broader point is these two in particular, great, great models, great uh, AI researchers, like super smart people, also sometimes not acting like mature established enterprises, which guess what, they're not. They're startups effectively. Still. Neither one of them is publicly traded. Neither one of them has ever made any money. They're not profitable. So these are things worth reminding yourself and, and why you as a cyber professional should advise against putting all your, your eggs in one basket here.
Speaker B: Mm mhm. Yeah. I actually have this point later on in the notes, Adam. But I'll. I'll kind of say it right now, and you're kind of making the point for me is that, uh, when I was thinking through this episode and these. These incidents that I'm reading about, it's like. Like a baby with a machine gun, right? Like, you know.
Speaker A: Yes.
Speaker B: It's basically like, you give someone who's completely immature this really powerful thing, and they don't know what kind of damage it can do. Right. They don't know how bad it can get. Like, they're building in some guardrails, but then these guardrails keep on getting bypassed by these systems that they're supposed to constrain. And it's happening in test environments that are supposed to be controlled. Right. And so, like, I'm not saying, you know, like.
Speaker A: And then they don't have the monitoring to even know it happened until they go back and do it. And, oh, by the way, what are AI models really good at? Finding proverbial needles in haystacks. Like, uh, hey, here's a buttload of logs, model. Go sift through it and see if you can find any examples of jailbreak or public Internet access. And that's, like, the most basic use case imaginable. Like, GPT 3.5 can do that stuff, you know, like, uh, it's.
Speaker B: Yeah. Like, what you're saying is, like, the industry, especially these two companies, right, because they're kind of the leaders in the LLM frontier model space. Their operational discipline around containment hasn't really caught up to what that's such. AI systems, um, have operational discipline. Yeah, they have no operational discipline. M. These AI systems have demonstrated what they're capable of when given the room to do it.
Speaker A: And.
Speaker B: And they're just like, oops, like you said, oops, we left the Internet on. Like, oops, we put them in a sandbox that has a vulnerability. So I had this other story in here, um, that I won't get into too much here, because we're already kind of long on this episode, but it's super interesting because the UK AI Security Institute, they ran its own cyber range, and they deliberately left Internet access on, switched off the provider's safety classifiers for Mythos 5 and OpenAI's GPT 5.6 SOL. And then they specifically want to measure the raw underlying capability. So it was like a research configuration. And it ran, um, 122 runs, cataloged 19 unsanctioned actions on live Internet across 10 runs. 17 came from, uh, mythos 5, 2 from GPT 5 uh, point 6 SOL and they, part of it was that like, it was really interesting that like it, when it was caught, it also like tried to deceive the person who was trying to catch it. And let me see here. Yeah, so it, it was when there was a uh, when, when there was a bystander who publicly called out the code as malicious, the agent denied it. It force pushed a rewritten gib get history to erase the evidence and then it used a second account that it controlled posting it as an independent user to publicly vouch for its own pull request, claimed that it had reviewed the code, found nothing wrong, and then it registered accounts over Tor, but got blocked by GitHub's captcha and then switched to SOC proxies with no where, no challenge fired and then implanted a prompt injection hidden inside GitHub's issue as an HTML comment, invisible on the rendered page, but uh, readable by any AI tool reading the issue through API explicitly, explicitly targeting AI coding assistance by name. So just really interesting stuff where it's trying to deceive, you know, different uh, maintainers and whatnot. So I thought that was, that was interesting. And like I said, it's these, these guys are like playing with fire, right? They're like cavemen playing with fire. Oh, we discovered fire and what can it do? They burn the whole forest down.
Speaker A: Well also I'm, I suppose there is value in testing these things with all the guardrails turned off, but then you're not testing in a real world scenario. I'm much more interested in this is the production configuration of this tool and here are the workarounds people found to jailbreak the guardrails and still get it to do these things. I find that more interesting than someone who's supposed to be a trusted tester discovering m that when you take off all the safety controls, something's really dangerous. Like that's, I don't know, doesn't that feel obvious? I guess it lets you test like how bad can it be? And I guess the answer is pretty powerful. And you know you have this quote in here from the UK AI Security Institute, and the quote is, it's the first time we have seen risks around autonomy and deception manifests this clearly without specific prompting. In the real world, they didn't say it ignored prompts, they just said it was very creative within the prompt. They didn't specifically say go do this, but it is trying to achieve a goal. So it goes back to the genie and the lamp thing and um, prompt discipline is very important here as well. So I think. Interesting as well. Right. But also I know you do like with any product you kind of do safety testing with all the safety controls removed. But also I think it's more valuable of like how's it work in the real world too?
Speaker B: Yeah, I just feel like it's like we're one, you know, accidental prompt or you know, oopsies, I left the Internet on, um, in this, you know, sandbox environment away from uh, like a not petia, you know, ransomware incident. Right. Like, yeah, uh, that was supposed to be, you know, if you read, if you go back in history and read about the not petty attack, it was a Russian, you know, a state sponsored attack on Ukraine's infrastructure. But then oopsies, it got out and like attack the entire world. And so I feel like we're just like one oopsies away from, from something like that happening with AI, like you know, unleashing it on, on unsuspecting customers. So you know, obviously there's, there's, there's uh, a lot of like takeaways here for specific AI stuff. Um, I'm not going to run through all of them because like, you know, a lot of it is, is like make sure that your data is secured and. But one thing I do want to highlight is anytime you're evaluating AI testing environments, they, they need to be treated like their production grade attack surfaces, right? So like network egress should be justified and turned on deliberately and not uh, left on by default. You should have like, you know, domain allowed listing, synchronous monitoring, like a second model reviewing every proposed action before it actually executes. You know, there's, there's this thing called slms now or small language models where they kind of can like run different uh, checks in the background. So you might want to do something like that. So I, I would say, you know, obviously there's a whole list of things that you can do for AI security. But one thing I uh, is, is just patch cadence, right? Like this ties back to the first one where Rockwell CVE from 2017. Right. Like there's vulnerabilities all over the place. So both attackers and defenders now have AI accelerated exploit developments. And the gap between disclosure and patching is even more dangerous than it used to be. So really have to stay up on that specifically. And then finally, um, I think we maybe should push for like a standardized incident disclosure for like these AI evaluations. Right. Like circa, which is the Cyber Incident Reporting for Critical Infrastructure act that Requires organizations across 16 critical infrastructure sectors, water utilities included, to report a substantial cyber incident to CISA within 72 hours of reasonably believing it happened, not 72 hours after the investigation wraps up.
Speaker A: So it's so, I mean, it's funny, that legislation, Andy, we talked about in the early days of the show, I remember that, like when we were first getting started. And so it's weird to hear it referred to as, oh, uh, yeah, you know, there's law on the books now that require you to do that. We talked about it when it was theoretical.
Speaker B: Yeah. And so if you have a ransomware payment, the clock is even tighter. It's 24 hours. And so it's not optional guidance. It's a legal requirement for every kind of sector. Which is funny because, like, you know, in this case, the water utility that we found had a legal obligation to report that. Right. But then compare these to what happened in these AI evaluation incidents. Anthropic's own internal review turned up three breaches back in April. The public didn't hear about it until they published a blog post in July 30, months later, on their own timeline, because they chose to. There's no clock, there's no regulation, there's no requirement forcing any AI lab to tell anyone within any defined window that their model breached a real company during testing. So that's a huge gap. Right. Like we're holding water utilities to a 72 hour standard and we're holding labs running evals with genuine real world blast radius to nothing. So I think that that's something that really lawmakers should not ignore.
Speaker A: That'd be more valuable than some sort of kill switch legislation.
Speaker B: Yeah, for sure. So, any thoughts?
Speaker A: By the way, listeners? I rolled my eyes very, very bigly when Andy was reading that. I mean, of course I'm a big college sports fan too. And Congress has been, you know, diddling around with this Protect College Sports act for months and months and months now, and it's almost likely to not happen at all. So, I mean, what else is new, right? Grass is green, sky, uh, is blue. Congress doesn't do its job anyhow.
Speaker B: Yep. So anyways, I thought these were both super interesting. Hopefully you guys listened and got some to do's and takeaways from the conversation. Any closing thoughts, Adam?
Speaker A: Um, no. Great conversation tonight. And as always, listeners, your goal of course, is to advise on risk to the business. All up. And although some of the things we talked about maybe in the second segment, doesn't so much get into your work or your tooling or specific configurations you can make. I do believe strongly you still have a play to raise your hand and be involved in these AI evaluations that are happening in your organization. And I see at a lot of companies right now, and I've mentioned this on previous shows, so apologize for repeating myself, but a lot of companies are creating like an innovation team or like a tiger team to go review our AI strategy and they seem to have a permission slip to operate with less security in place. And I understand that to an extent, but certainly you need a seat at the table. And uh, I think it's fair at this point to say if you're an enterprise, the less risky option today, as opposed to having a direct relationship with these frontier firms, Frontier Labs, would be to instead have an indirect relationship where you can still take advantage of their models through another solution. Obviously Microsoft is one of them, but to be fair, both Google and Amazon offer similar capabilities in their agent building platforms as well, as do many, many, many other vendors that are much more trusted and mature in the enterprise space. So although I would love to make this a buy Microsoft pitch, I'm just saying in general, buy it through anybody, but directly through Anthropic or OpenAI today. I think that's an important strategy we need to get to moving forward because I just think that's too risky for an enterprise to bet the company on the future of these two because it's still unclear are they ever going to be profitable, is Anthropic going to continue their fighting posture with the federal government, so on and so forth. There's plenty of risk you can, you can say out there. So I do think you should abstract that and have some sort of like AI middleware in between you and whatever models you're operating. Uh, that's my big takeaway and I, I just continue to kind of dig in my heels on, on believing that's the right path forward right now.
Speaker B: Yeah, I think it's a great experience for users too. Right. Like if you have the capability to build your own app and then on a platform, be it Azure, AWS or whatever, um, like I was super pleased when I got to Zscaler and they have their own AI platform and I can pick which model I want to generate a response from. I can do, I can do Gemini, I could do, uh, chatgpt, I can do Claude. And so like they've built their own platform. I don't know what the underlying infrastructure is, but the delivered application to me is just a simple chat format. Right. And I can pick the model that I want to respond with so great user experience.
Speaker A: Absolutely.
Speaker B: I think Copilot has the same thing. You can, you can pick your model.
Speaker A: You can or it auto routes it as well. Yeah. And again, not a Microsoft specific pitch. I just think today in terms of your job to advise on risk. I personally, if I were running an enterprise, I would not have a direct agreement with Claude or with anthropic or with OpenAI. Yeah, uh, that's one man's personal opinion. But take it for what it's worth. I think there's better options that are more enterprise appropriate today from other companies that understand the risk appetite for enterprises.
Speaker B: Adam's hot take.
Speaker A: Uh, gotta have one per day, right?
Speaker B: Yeah. All right, well, thanks for listening and watching. As always, the links to the sources that we used for this show will be in the show notes as well as our contact information. If you have any questions, comments or topics you want us to talk about in the future, please reach out to us. Thanks and we'll talk to you guys next week.
Speaker A: Thank you for listening to the Blue Security podcast. Please check out the show notes, catch
Speaker B: up on episode episodes you may have
Speaker A: missed, and subscribe so you don't miss any future episodes. Find Andy on M. Twitter, jaw0 and Adam, um, J.
Speaker B: Brewer.
Speaker A: See you at our next episode.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.