
The Industrial Security Podcast · 2025-12-13 · 44 min
Key moments - from our scoring
Substance score
42 / 100
Five dimensions, 20 points each
Recovery capabilities are often overlooked in industrial cybersecurity despite being a pillar of the NIST Cybersecurity Framework. This episode addresses that gap by exploring how modern backup and recovery solutions handle the unique challenges of OT environments. Stephen Nichols explains why traditional home-baked backup approaches - storing copies on local network-attached storage - fail when ransomware infiltrates a facility; infected systems encrypt both production and backup infrastructure simultaneously. The solution involves immutable cloud storage combined with architectural principles like the Purdue model: agents on control systems initiate lightweight, compressed backups that flow through DMZ-resident management servers to cloud repositories (Azure, AWS, or dedicated clouds like Acronis offers). Acronis has OEM partnerships with major automation vendors including Honeywell, Emerson, Rockwell, and Siemens, validating its agent software on DCS platforms like Delta V and Ovation. Key capabilities include universal restore (booting to dissimilar hardware by detecting chipset, network, and storage drivers) and compliance alignment, making this relevant for manufacturers, utilities, and critical infrastructure operators facing both regulatory requirements and genuine downtime cost pressures.
Ransomware that penetrates a facility encrypts both production systems and any online backup storage on the same network. Immutable cloud storage solves this by keeping backup copies offline and protected from encryption attacks.
The Acronis agent on each control system initiates communication outbound to a management server in the DMZ, which then copies data to cloud storage. This respects the principle that control systems should not receive inbound connections from less-trusted networks.
Universal restore allows backup images to boot on dissimilar hardware by automatically detecting and loading chipset, network, and storage drivers. This matters because hardware fails and replacement units may have slight variations in components.
Honeywell, Emerson (Delta V, Ovation), Rockwell Automation, and Siemens have OEM partnerships with Acronis, meaning they've tested and validated the solution on their DCS platforms.
The agent is lightweight and compresses data before transmission, meaning fewer packets are sent; it's designed for high-latency environments and can be configured to share the operational network or run on a separate network without demanding priority bandwidth.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains a handful of genuinely useful technical specifics - Purdue-model-aware agent architecture, the three-driver universal restore requirement, and malware-scanning backups before restoration - but these are buried in extended product-pitch monologues and platitudes about uptime. The ratio of novel content to padding is poor for a 44-minute episode.
the modern best practices is to have that that copy actually in an immutable storage in the cloud
we have the ability to scan that backup through the same engine that we would use to scan you know, actively on a system to look for things, meaning that we can mark a backup as safe to recover or infect it
The framing of cloud backup as increasing resilience rather than merely efficiency is a mildly fresh angle, and 'weaponizing non-IT people' is a memorable coinage, but the underlying content - 3-2-1 backups, immutable cloud storage, air-gap concerns - is standard industry thinking with no contrarian or first-principles arguments.
here's a cloud that increases your resiliency rather than simply increasing your efficiency
weaponizing non I people
Steven Nichols is a country manager and former solutions-engineering lead at a backup vendor, not an independent OT security practitioner or industrial operator who has done recovery at scale under duress. He speaks competently about his own product but brings no operator-side war stories or cross-vendor perspective.
My current role is as the country manager for a Chromis. But I've been doing technology and in all sorts of different places from professional services account management for about twenty years
I've been at Acronas for about six years. Prior to this current role that I started in January, I actually ran our solution engineering team for the Americas
A handful of concrete data points appear - recovery times of 5 - 15 minutes, 60,000 endpoints in 30 days, named OEM partners (Honeywell, Emerson Delta V/Ovation, Rockwell), CrowdStrike's 8.5M devices and $10B figure - but all customer case studies are unnamed and the numbers largely come from the host rather than the guest's direct experience.
they were back up and running with those devices in an average of five to fifteen minutes
we were able to get that onto about sixty thousand end points within thirty days
Andrew Ginter raises one genuinely sharp challenge - whether cloud-based recovery is itself a risk vector given the CrowdStrike precedent - and follows up by asking whether Acronis customers actually used the product during that incident. However, Ginter frequently answers his own questions with extended monologues, reducing the guest's airtime and preventing deeper probing.
if the cloud is a vector for such widespread destruction, potentially through malicious or honest means, then is cloud based recovery a good idea
Did you have customers in your knowledge who were hit by CrowdStrike and you know recovered this way
Computed from the transcript - who did the talking, and the words that came up most.
We've been hacked. Everything is down. Or more mundane - there was a power surge and 5% of our cyber gear is fried. How do we get back into operation fastest? Stephen Nichols of Acronis joins us to look at rapid recovery of OT systems - from the mundane to the arcane.
Transcribed and scored by The B2B Podcast Index.
1 - >
Speaker 1: And having some kind of you know, local copy is 2 - > better than not having any The modern best practices is 3 - > to have that that copy actually in an immutable storage 4 - > in the cloud. 5 - >
Speaker 2: Welcome listeners to the Industrial Security Podcast. My name is 6 - > Nate Nelson. I'm here with Andrew Ginter, the vice president 7 - > of Industrial Security at Waterfall Security Solutions, who's going to 8 - > introduce the subject and guest of our show today. Andrew, 9 - > how's it going. 10 - >
Speaker 3: I'm very well, Thank you, Nate. Our guest today is 11 - > Stephen Nichols. He is the country manager for Canada at Akronus, 12 - > and we're going to be talking about rapid recovery of 13 - > OT systems in different kinds of emergencies, including cyber attacks. 14 - >
Speaker 2: Then, without further ado, here's your conversation with Steven. 15 - >
Speaker 3: Hello, Stephen, and welcome to the podcast. Before we get started, 16 - > can I ask you to introduce yourself for our listeners 17 - > and to say a few words about the good work 18 - > that you're doing at acronis. 19 - >
Speaker 1: Absolutely my pleasure to be here. So my name is 20 - > Steven Nichols. My current role is as the country manager 21 - > for a Chromis. But I've been doing technology and in 22 - > all sorts of different places from professional services account management 23 - > for about twenty years, and in that time I've worked 24 - > at a number of different companies. Most recently, I've been 25 - > at Acronas for about six years. Prior to this current 26 - > role that I started in January, I actually ran our 27 - > solution engineering team for the Americas, so I've been deeply 28 - > involved in setting up pocs, working with our partners and 29 - > our customers to make sure that they both understand the 30 - > technology and get the most out of it. Acronus, who 31 - > have been around for about twenty two years, have a 32 - > solid background in backup and recovering of data as well 33 - > as cybersecurity, so we understand that space and we've definitely 34 - > got ways we can help you to solve the problem 35 - > of uptime and data availability. 36 - >
Speaker 3: Our topic is rapid recovery, which you know, to me 37 - > implies recovering from some kind of backup in of course 38 - > OT networks, you know, some of which are our critical infrastructure. 39 - > We if something goes wrong, something breaks, something gets hacked, 40 - > we need these things to come back quickly. So yes, 41 - > but you know what I'm used to seeing at I 42 - > don't know, power plants and such, is that you've got 43 - > the control network with you know, a bunch of equipment 44 - > on it. You've got a lot of the time a 45 - > parallel network. Most of your devices have two network interfaces, 46 - > one to the real time network and one to call 47 - > it a management network, and there might be a network 48 - > attached storage device on the management network that stores backup 49 - > files for all of your equipment. You know, you're solving 50 - > a problem. Is is this the problem you're solving? Why 51 - > is you know, a parallel network and a network is 52 - > to that storage not not enough? What is the problem 53 - > we're solving here? 54 - >
Speaker 1: The thing generally when we're talking to you know, in 55 - > that that space, when we're talking to customers, the real 56 - > thing that they're trying to solve is up time. So 57 - > there they downtime is incredibly expensive disruptive, especially you know 58 - > with in supply chain issues where you're talking about just 59 - > in time delivery. There isn't a lot of flexibility and 60 - > there's a whole lot of cost when you have downtime. 61 - > So making sure that there is a way if again, 62 - > whether it's hardware failure, whether it's a corruption of the data, 63 - > whether that you know, haven't forbid that should be some 64 - > kind of cyber attack, making sure that you have the 65 - > ability to recover from that quickly is super important. And 66 - > let me give you an example. There's an auto manufacturer 67 - > that we work with and if they have a problem 68 - > you know, with their you know, one of their systems, 69 - > the only solution that they had prior to working with 70 - > us was that somebody from Germany had to get on 71 - > a plane and bring a you know, a new device 72 - > and get that installed. So that's a huge amount of downtime. 73 - > So being able to have a very reliable recovery of 74 - > that data and not just being able to recover that 75 - > back to the same device, because a lot of times 76 - > it can be actual hardware failure. So it's about being 77 - > able to recover to dissimilar hardware and be able to 78 - > do that quickly and efficiently. And that means being able 79 - > to to back those control systems up. It means being 80 - > able to work within the network environment from a security 81 - > point of view, and it means being able to uh, 82 - > you know, for example, if it is just data corruption 83 - > is not physical hardware, being able to what we like 84 - > to refer to as weaponizing non I people. So if 85 - > it is just you know, on that individual endpoint, then 86 - > being able to do that recovery quickly and simplify that process, 87 - > so it doesn't you don't have that bottleneck of it. Tuh. 88 - > The other thing that we're we're seeing that is really 89 - > driving a lot of you know kind of that that 90 - > approach is around compliance. So in a lot of cases, 91 - > from a compliance point of view, making sure that there 92 - > are reliable backups of those systems and the ability to 93 - > do that in a secure way makes a huge difference 94 - > to checking those boxes from a compliance point of view. 95 - >
Speaker 3: So that all makes sense in the abstract, but you know, 96 - > the concrete example I had in mind again was you know, 97 - > ICEE sites roll their own they've got a storage server, 98 - > they keep the backups on it. But to your point, 99 - > you know, for reliability, if ransomware gets in there, your 100 - > backup server had better not be online or it's going 101 - > to get encrypted, just like everything else got encrypted. So 102 - > now we need offline copies, and you know, for compliance, 103 - > a lot of the compliance regimes say, well, your offline 104 - > copy has to be off site in case, I don't know, 105 - > there's a fire in the building and you lose all 106 - > your backups because that was the room that caught fire. 107 - > Is this is this what you're talking about? Is this 108 - > the problem we're solving here? 109 - >
Speaker 1: Yeah, absolutely, and you know, again having some kind of 110 - > you know, local copy is better than not having anything 111 - > one hundred percent. But you start to get into a 112 - > lot of complexity when you're trying to solve that that 113 - > problem with kind of you know, home baked solutions. So 114 - > you know, ultimately the what the modern best practice is 115 - > is to have that that copy actually in an immutable 116 - > storage in the cloud. So, you know, using the Purdue model, 117 - > having the the agent you know, move up through the network, 118 - > make be able to you know, securely initiate that communication, 119 - > copy that to to the to a storage or management 120 - > server on site that can be in a d M 121 - > z uh and then have that copied out to the cloud, 122 - > whether that's an uber public cloud like Azure or a WUS, 123 - > or whether that's to you know, a dedicated cloud solution 124 - > you know, like a Chronus offers. Being able to do 125 - > that and store that in a mutable way is really 126 - > where you're going to check those compliance boxes without adding 127 - > a large amount out of operational overhead on top by 128 - > using a solution this designed to solve this problem. 129 - >
Speaker 3: I know, this is one of the things a Cronus 130 - > does you solve this problem. Can you talk about your solution. 131 - > You know, what have you got? How does it work? Please? 132 - >
Speaker 1: Being able to solve for that problem really requires a 133 - > few things, and one of them is ease of use. 134 - > It's about being able to back up and restore that 135 - > data without a lot of overhead or complication. And in 136 - > a lot of cases that restorees even to dissimilar hardware 137 - > if you have physical failure, Even if you know it's 138 - > the same model of hardware that you're getting from the 139 - > same manufacturer, there might be a slightly different chipset, a 140 - > slightly different thing like that. So being able to restore 141 - > that dissimilar hardware. One of the things we have is 142 - > something called universal restore. It means as long as we 143 - > can load three drivers, we can make bootable. So as 144 - > long as we've got the chipset driver, the network driver, 145 - > and the storage driver, and we have technology in place 146 - > to be able to detect and load those drivers automatically, 147 - > and even if we can't, we have ways you can 148 - > manually inject those meaning that the ability to restore and 149 - > get that system back up is paramount and we know that, 150 - > so we really focus a lot on being able to 151 - > do that. And the other thing that I think is 152 - > really important about how we solve that problem is easy 153 - > of use, so you know, just a handful of clicks. 154 - > Being able to deploy, you know, configure the settings for 155 - > a backup or being able to recover that backup means 156 - > that you're you don't have to learn a complicated interface 157 - > to be able to get things to work. 158 - >
Speaker 3: Can we go back to the basics. Don't know how 159 - > your system works? You've talked about cloud, you know, how 160 - > do you get data out to the cloud? We talked 161 - > about Purdue model. How does that work? 162 - >
Speaker 1: Our solution is sort of divided up into three pieces. 163 - > So the first piece is an agent. So that agent 164 - > gets installed onto the control system. It is you know, Windows, Linux, 165 - > broad support for any of those legacy systems. But with 166 - > that agent, once it is installed, that's what initiates the communication, 167 - > so whether that's going out to what we call the 168 - > management service. The second piece that's installed on premise, and 169 - > that is the piece where you're going to set up 170 - > all of the settings, where you're going to do all 171 - > of the configuration, and the agent is going to reach 172 - > out to that management server to say, hey, what what 173 - > are the settings? What do I you know, how often 174 - > do I need to back up? If you need to 175 - > initiate a recovery that can be done from the device 176 - > or from the management server. And basically the agent is 177 - > the one that's going to be initiating. That's how we 178 - > respect the Purdue model. That management server, that second piece 179 - > lives in the DMZ, so it's the part that can 180 - > be connected both to the secure network and to the Internet, 181 - > meaning that it can then copy those that data out 182 - > to the cloud or go and retrieve it from the 183 - > cloud and bring it back on premise. And the advantage 184 - > of that is it can you know, you have a 185 - > way to be able to securely move that data. But 186 - > it's the agent, it's doing the heavy lifting, and that's 187 - > what's installed on the actual control system itself. 188 - >
Speaker 3: So Nate, let me jump in for a second. I mean, 189 - > this episode is a little bit unusual in that I 190 - > can only remember one other episode in all of our 191 - > over one hundred episodes where we've talked about the recovery 192 - > pillar of the this cybersecurity framework. You know, the this 193 - > framework is, of course, hundreds of experts got together and said, 194 - > here's what a complete cybersecurity program looks like. It has 195 - > six pillars. In the modern instantiation of the of the 196 - > framework govern identify, protect, detect, respond, and recover six pillars. 197 - > And I only know if I only recall one other 198 - > episode where we talked about recover, and that was the 199 - > Salvador Tech episode. So here we are talking about the 200 - > recovery pillar, you know, and it's it's important. I mean, 201 - > in especially safety critical industrial stuff. Nobody wants the system compromised. 202 - > Nobody wants you know, heavy industry, nobody wants, well forget 203 - > heavy industry, their factory, anything to go down. And you know, 204 - > if it goes down for any reason, we want to 205 - > be able. We have to be able to bring it 206 - > back or you know, we've done the business that is 207 - > serviced not just go down because of cyber attacks. It 208 - > might just go down because equipment burns out. This happens. 209 - > You've got to be able to recover from these these outages. 210 - > And you know what is there? We're talking about backup 211 - > and recovery. You know, it's a pillar. Is there? Is 212 - > there anything else? The only other thing I know of 213 - > in that pillar is rebuilding from known good original media 214 - > as an alternative instead of backing up and restoring, rebuild 215 - > a machine from scratch, which probably takes a lot longer 216 - > than recovering from backups, So yeah, most people back their 217 - > stuff up, and you know, might also have recover from 218 - > from known good media as sort of a second level 219 - > of recovery option to be used in I don't know, emergencies. 220 - >
Speaker 1: For whatever reason, do you. 221 - >
Speaker 2: Think it's inherently just less complex than those other five pillars, 222 - > such that fewer people would either be involved in that 223 - > space as specialized in that space as vendors or want 224 - > to talk about it on podcasts or is it very fruitful? 225 - > And maybe we, just as podcasts hosts, have not been 226 - > doing a good job of finding these folks. 227 - >
Speaker 1: Well, I don't know. 228 - >
Speaker 3: I haven't actually gone out to see who else is 229 - > in the space. I know a Cronus is one of 230 - > the major players. I see them everywhere. But I think, 231 - > and I don't have the niss CSF open in front 232 - > of me, but from what I recall, there are far 233 - > fewer things to do requirements in the recovery pillar than 234 - > in many of the other pillars. So it is sort 235 - > of because there's fewer requirements, it's arguably a little bit easier. 236 - > It's you know, there's not as much to do, so yeah, 237 - > I think that's part of it, But it is important 238 - > and so yeah, maybe we have been remiss. I'm happy 239 - > that we have Steven as a guest here. This is 240 - > an OT system. You know, sometimes they're old and slow, 241 - > you know, sometimes they're modern and fast. Is the is 242 - > the backup data passing across the same you know, competing 243 - > for bandwidth on the same network that is doing the 244 - > real time I don't know, train control or something horrible 245 - > like this. Is it on a separate network and you 246 - > know these agents? Do you get pushback from the vendor saying, no, 247 - > you're not installing your agents on my stuff. You've invalidated 248 - > and support agreement. Blah blah blah. You know, can you 249 - > talk about the problem of doing this on an OT network. 250 - >
Speaker 1: First of all, the agent is very lightweight. It does 251 - > a little bit of the work, so it will actually 252 - > compress that data prior to it being sent, meaning we're 253 - > actually sending fewer packets. It is designed to work in 254 - > a in a in a high latency environment because the 255 - > which means that it doesn't demand a lot of priority 256 - > on network traffic. To be able to do that, it 257 - > can be configured in a way to work over that 258 - > same operational network, or you can configure a separate network 259 - > that it's the flexibility of the model I think that 260 - > is really important. And as far as you know those environments, 261 - > we actually have really tight relationships with a number of 262 - > ot vendors. So I talk about Honeywell, e v R, Emerson, 263 - > Delta V and Ovation Rockwell. You know, they have OEMed 264 - > our product, they've tested it, they've validated it on their systems, 265 - > and they resell our solution, which means that they're certified 266 - > for those systems. And because of those tight relationships, you 267 - > can be confident as a customer that you're getting a 268 - > solution that is going to be reliable and not interfere 269 - > with the primer use of that device, because one hundred 270 - > percent the backups are important. But if it's not delivering 271 - > the service that it needs, if you're not you know, 272 - > moving people or energy or water around, you're ultimately, you know, 273 - > making more problems than you're solving. 274 - >
Speaker 3: Well, you talked about low bandwidth, you talked about not 275 - > competing with the control system. Is it possible to throttle 276 - > how much of the network you use for the backup function. 277 - >
Speaker 1: One of the things about the protection plan or the 278 - > backup plan that we have the configuration settings in those options. 279 - > We have the ability to set up limits on the 280 - > amount of bandits, so you can say a certain percentage, 281 - > you can say some certain number of kilobits per second, 282 - > and you can also have that change at different times 283 - > of the day, different days of the week. So I mean, 284 - > if you're running twenty four to seven, you can have it, 285 - > you know, consistent across the board. If you have you know, 286 - > more priority times that you know you need that operational 287 - > information to take priority, and you have you know other 288 - > times where you can have the backup take priority. Of 289 - > both of those things, we can absolutely accommodate. It's about flexibility. 290 - > It's about making sure that you have all of the 291 - > capabilities that you need without interfering with the primary requirement 292 - > of the network. 293 - >
Speaker 3: We've been talking sort of again in the abstract systems agents. 294 - > You generally can't take a Windows agent and install it 295 - > on a PLC, and you might have a challenge installing 296 - > it on an old XP system that we still see 297 - > running in the plant. Sometimes. Can you talk about, you know, 298 - > what kinds of platforms you support. 299 - >
Speaker 1: Let's start with PSC. So in most cases that information, 300 - > those configurations are actually being backed up or protect did 301 - > or exported to a control system somewhere, and that's usually 302 - > being done by the proprietary software within the within that vendor, 303 - > and once it's on the control system, we have the 304 - > ability to back that up. Additionally, we have those deep 305 - > OEM relationships in a lot of cases, what is actually 306 - > being done to back those up directly is the Acronus agent. 307 - > And when we talk about legacy systems, we've been we've 308 - > been in this data protection, in this this backup and 309 - > recovery business for twenty plus years, and we know that 310 - > a lot of these systems absolutely rely on legacy operating systems. 311 - > So we have maintained support for even out of supported systems, 312 - > back to things like XP you know, Windows Server two thousand, 313 - > you know, lots of those those systems, And the reason 314 - > that we do that is because we know that those 315 - > those critical systems can't easily be updated. 316 - >
Speaker 3: We've talked a lot though about sort of the process 317 - > of backing things up. The agent's back things up. You know, 318 - > the PLCs have have got relationships with you folks. They 319 - > might have you know, your software embedded in them. That's 320 - > that's all good. Can you talk about the other side 321 - > of the coin. There's I don't know, a cyber attack 322 - > or you know just plain old electrical fault in you know, 323 - > part of a high power process and you know a 324 - > half dozen you know, a couple of Windows devices, a 325 - > couple of PLCs fry. So we scramble to find replacement hardware. 326 - > We get the replacement hardware there, and and then what 327 - > how do we restore? 328 - >
Speaker 1: Yeah, I think that's a great question because you know, 329 - > I've often said that, you know, backups are wonderful, but 330 - > it's really recovery. It's the ability to get things back 331 - > up that that matters. And so if you kind of 332 - > touched on a couple of scenarios there, and we have 333 - > a couple of different ways to address that. So uh, 334 - > let let's start with the last one you mentioned where 335 - > it is now new hardware. So I talked a little 336 - > bit earlier about our universal restore, So really it doesn't 337 - > have to be exactly the same hardware, and I think 338 - > that's that's important. But in addition to that, just generally, 339 - > how do you how do you go about get recovering? 340 - > And when it's a physical machine, the we have the 341 - > ability to have bootable media, so you can you can 342 - > go to that management server, you can create a bootable 343 - > a USB or or CD, you know, optical drive version 344 - > that you can go boot on that machine. It will 345 - > give you all you know, very very quickly. It installs 346 - > a lightweight version of an operating system and uh it's 347 - > it will be registered to the management server and be 348 - > able to pull down and do that recovery, whether that's 349 - > you know, information you have stored locally or even if 350 - > you have information that is stored in the cloud. So 351 - > it makes the process of being able to do that 352 - > just a few clicks on the physical hardware itself. The 353 - > other option, let's say that you don't have a you know, 354 - > damaged hardware, but it is in fact, you know, some 355 - > kind of cyber attack or you know. Another example could 356 - > be the challenge with crowd strike a while ago, where 357 - > you know, you have some kind of just you can't 358 - > boot the system. It's physically the physical system is fine, 359 - > but somehow there's been some corruption or some damage. In 360 - > that case, we actually have the ability right on the 361 - > machine what we call one click recovery. Technically it's three 362 - > or four clicks, but it gives you the ability to 363 - > on boot hold down the F eleven key. It will 364 - > boot into that recovery manager that can be on the 365 - > machine and then you can, you know, simply choose to 366 - > restore from a backup a week or a couple of 367 - > weeks ago before the problem occurred where that ransomware happened, 368 - > and that gives you the ability again very very simply 369 - > without needing a lot of time to configure and understand. 370 - > It's just a matter of a couple of clicks. 371 - >
Speaker 2: Steven just mentioned there the crowd strike incident, Andrew. At 372 - > the time we're recording this, I assume that we're going 373 - > to release it some weeks later. But at the time 374 - > of recording, we are almost one exact year to the day, 375 - > just two days off of the anniversary of when a 376 - > faulty configuration update caused crowdstrikes flagship program to bug out 377 - > and do blue screens of error in what was I 378 - > think eight point five million devices. This was probably the 379 - > most highly publicized it or cyber related event in history, 380 - > affecting hospitals, shops, airlines, flights were grounded across America. I 381 - > think that the number that I found on the web 382 - > was that it caused somewhere in the region of ten 383 - > billion dollars in damages all told by the end, a 384 - > number which is probably relatively rough, and also reminds me 385 - > of the last time such an event occurred, or at 386 - > least that I can remember, which would have been an 387 - > actual cyber attack, namely not Petia, which, while it caused 388 - > fewer public and very visible errors to the general public, 389 - > managed to cause hundreds of millions of dollars in damages 390 - > for very important supply chain companies or not. So this was, 391 - > if anything, ever was to be a case study in 392 - > what it looks like to do recovery. It was the 393 - > CrowdStrike instant last. 394 - >
Speaker 3: Year, and I agree with that absolutely, and the crowdsyke 395 - > incident was not a cyber attack, but you know, it 396 - > does illustrate what some people what I have worried about 397 - > for some years on the side of cloud connectivity for 398 - > control systems. If you know, an honest mistake in software 399 - > development could or testing or whatever the process was, could 400 - > cause you know, eight point five million PCs to go down, 401 - > and you know, industrial consequences in terms of grounding air 402 - > flights and shutting down other processes. If that's possible through 403 - > human error, it's arguably possible through a deliberate attack as well. 404 - > And you know, so crowd Psyche was not an example 405 - > of a liberate attack, but it's an example of the 406 - > scale of what's possible. Kassea was actually a deliberate attack. 407 - > This was on the IT side, not the OT side, 408 - > but it was ransomware that was inserted into a cassea 409 - > software update server and infected I think, what was it 410 - > eight hundred or one thousand businesses all at once, sort 411 - > of on the scale of not Petya, but you know, 412 - > ten years later using different technology. So you know, I 413 - > worry that it is possible to use these these cloud 414 - > based systems to reach back and you know, kill a 415 - > lot of industrial stuff that might otherwise seem very heavily defended. 416 - >
Speaker 2: And I think that this kind of risk has come 417 - > up on our podcast before because sometimes we talked to 418 - > vendors who provide cloud based security services, for example, security 419 - > operation centers through the cloud. In this case, the question 420 - > that naturally comes to my mind is if the cloud 421 - > is a vector for such widespread destruction, potentially through malicious 422 - > or honest means, then is cloud based recovery a good idea. 423 - >
Speaker 3: That's an important question. But you know, here we're talking 424 - > about a backup system. We have to be careful that 425 - > we're not throwing out the baby with the bathwater. You know, 426 - > look at the world. Most you know, the vast majority 427 - > of industrial sites are already connected to the cloud. One 428 - > more connection is not changing your threat profile materially, and this, 429 - > you know, their corners connection is a connection to a 430 - > backup service, something that is increasing your resilience. That's increasing 431 - > your you know, the strength of your security program. This, 432 - > you know, backing up and recovery is part of the CSF. 433 - > You don't have a complete security program until you've got 434 - > this capability. And you know, unlike the other cloud connections 435 - > at most industrial sites, here's a cloud that increases your 436 - > resiliency rather than simply increasing your efficiency, which you know, 437 - > predictive maintenance and other sort of conventional cloud connections do. 438 - > So you know, that's part of the problem. Part of it, 439 - > you know as well, is that you know, we need 440 - > we need a way to recover and there are measures 441 - > we can take. I mean, I talk about, you know, 442 - > plugging my own book network engineering in my latest book 443 - > Engineering Grade OT Security, you know, free copies of which 444 - > are still available from Waterfall. Check out the website or 445 - > send me email. There's a chapter on network engineering. You 446 - > can talk about extra inspection, you can talk about you know, 447 - > high lockdown, separate paths to the Internet. You can talk 448 - > about uni directional gateways when you deem the threat of 449 - > compromise over that information flows as a credible threat. So 450 - > you know, yes, in theory it's a risk, but what 451 - > you have to look at is in practice, what's the 452 - > benefit here? And you know, if I recall you know, 453 - > Steven said later in the interview, they have the option 454 - > of on premise backups as well. Now you lose the 455 - > off premise capability, but yeah, there's a lot of options, 456 - > and you know, you need to look at the big picture, 457 - > not just oh, look there's a cloud connection. You know, 458 - > we're doomed. That's it's not that simple. So you mentioned 459 - > crowd strike. I mean I find that fascinating. It was 460 - > a lot of hosts that went down with CrowdStrike. Is 461 - > this sort of hypothetical or is it real? Did you 462 - > have customers in your knowledge who were hit by CrowdStrike 463 - > and you know recovered this way? You know, you've you've 464 - > been deploying this, you were you were a part to 465 - > the services team at a Cronus for a long time. 466 - > Tell us, you know, what's the experience of using this, Like. 467 - >
Speaker 1: Yeah, I've got to the crowd strike example is you know, 468 - > I brought it up because it is real. We actually 469 - > do have a number of customers. One in particular I 470 - > can think of that had about two hundred devices and 471 - > they they reached out to us to say, hey, can 472 - > is this something that a coronas can help with? And 473 - > we were able to, you know, walk them through the 474 - > very simple process. And they had about thirty of those 475 - > machines in particular that were remote where they didn't have 476 - > IT people who could easily access those machines and their 477 - > process of recovering those booting into safe mode, being able 478 - > to you know, go in and make registry changes. Although 479 - > that takes a lot of IT experience, It takes someone 480 - > who kind of knows what they're doing. You don't want 481 - > just your average person and doing that. But for those 482 - > remote machines, what they were able to do was walk 483 - > a non technical person. Again, this is something we kind 484 - > of call weaponizing the non IT employees and the ability 485 - > to walk them through a very simple process that took 486 - > two or three clicks and was able to fully recover, 487 - > meaning that they were back up and running with those 488 - > devices in an average of five to fifteen minutes. As 489 - > opposed to having someone have maybe have to travel two 490 - > or three hours or maybe even longer than that to 491 - > be able to recover. And that makes a huge difference. Again, 492 - > it's you know, going back to what problem we're solving, 493 - > it's uptime. It's the ability to make sure those systems 494 - > are online and doing what they need to do. 495 - >
Speaker 3: Just going in and deploying this in an OT network. 496 - > I mean sometimes it's built in the vendor. As you 497 - > point out, you've got relationships with the vendors. When you 498 - > get the control system, it's there. You know, if I 499 - > want to apply something like this after the fact because 500 - > you know my vendor didn't support it or whatnot, you know, 501 - > what what's that feel like? 502 - >
Speaker 1: So we really try to make it a flexible model 503 - > to deploy. So you know, if you've got a small environment, 504 - > you know, maybe it's a highly secure environment, you can 505 - > actually deploy and register those agents just by you know, 506 - > through sneakernet right walking around with a USB and doing that. 507 - > But that doesn't really scale, right, So I think back 508 - > to you know, we have a large logistics company that 509 - > we work with in North America and they needed just 510 - > again a broad range of devices, so they were able 511 - > to work with scripting and group policy to be able 512 - > to deploy that out, and we were able to get 513 - > that onto about sixty thousand end points within thirty days, 514 - > and that really made a huge difference to their You know, 515 - > choosing us as a as a provider is that because 516 - > the effort it takes to actually get this deployed in large, 517 - > complex environments can be a huge deciding factor. 518 - >
Speaker 3: Most of the discussion we had so far, you know, 519 - > applies sort of universally. It applies if you know there's 520 - > an electrical fault and I have does machines fry? It 521 - > applies if there's a software fault, and you know eight 522 - > percent of crowdstrikes machines are blue screen. But we also 523 - > worry about cyber attacks and recovering after cyber attacks. Is 524 - > there anything we need to do differently or everything we 525 - > need to think about differently in the world of cyber 526 - > versus sort of normal failures. 527 - >
Speaker 1: Yeah, for sure. And part of that is that sometimes 528 - > just recovering doesn't eliminate the problem. Right, If you're just 529 - > recovering the ransomware back on, it's really not going to 530 - > give you the level of protection that you need. So 531 - > although the process of the backup and the recovery aren't 532 - > dissimilar in those cases. One of the huge things is 533 - > a lot of those control systems. Historically the approach to 534 - > cybersecurity has been isolation, right, it's been the air gap network. 535 - > So the reality is there are fewer and fewer truly 536 - > air gap devices today. People want to be able to 537 - > manage them remotely. People, you know, there's lots of those 538 - > types of things, which increases the vulnerability of those two 539 - > cyber attack but putting you know, a full suite of 540 - > protection on that device. A lot of the times it's 541 - > older operating systems that don't support the modern anti virus 542 - > and anti ransomware solutions. And secondly, there usually isn't a 543 - > lot of resource capacity there to be able to do that. 544 - > So how do you solve for that? And one of 545 - > the things that we can do is once we've taken 546 - > the backup, we have the ability to scan that backup 547 - > through the same engine that we would use to scan 548 - > you know, actively on a system to look for things, 549 - > meaning that we can mark a backup as safe to 550 - > recover or infect it. So if you did have a 551 - > ransomware attack that that that you know had some some 552 - > case where things were affected and you needed to be 553 - > able to recover from that. You can with a high 554 - > degree of confidence know that you can choose a backup 555 - > that isn't just going to reinfect those systems. And the 556 - > other piece that we have and this is this is 557 - > something kind of the first thing we did in the 558 - > in the world of security is something we call active protection. 559 - > What active protection is, it's a behavior based engine that 560 - > runs at the agent level and it detects processes that 561 - > are suspicious, particularly when files start to become encrypted. So 562 - > what it can do to text that knows that that's 563 - > not a normal behavior can stop the process and then 564 - > revert those files from system cash. This means that again 565 - > you're you're aware that there's a problem, you get on 566 - > alert that that has happened, but it means that it 567 - > should help to prevent that that downtime and then you 568 - > can go and do a recovery from a backup which 569 - > has been scanned and marked is safe. 570 - >
Speaker 3: And we've been talking about OT but I understand you 571 - > folks do this kind of thing for I T as well, 572 - > So you know, in a sense you're exposed to the 573 - > the the sort of the I T O T traditional 574 - > sort of conflicts or debates or you know, you're you're 575 - > you're living in the in the world of both I 576 - > T and O T. How's that working. 577 - >
Speaker 1: You know a lot of people that I talked to, 578 - > you know, bring up this idea of O T and 579 - > IT convergence, and when they do, you know, the first 580 - > things that come to mind are, you know, being able 581 - > to manage O T devices remotely. It's it's about having 582 - > you know, a network topology that will support the convergence. 583 - > But you know, in our case, it really it's about 584 - > being able to have a single solution. It's about you know, 585 - > getting away from siloed tools and being able to have 586 - > a single skill set. U have expertise, have a single 587 - > vendor that you can work with that can support those 588 - > So you know, we have and all the things we've 589 - > been talking around the the OT about you know, backup 590 - > and recovery, being able to recover to dissimilar hardware, about 591 - > you know, enabling the end user even to do some 592 - > of some of those tasks, being able to centralize the storage, 593 - > being able to work with remote workloads as well as 594 - > on premise workloads. You know, it's it's about that bringing 595 - > all of that together and being you know, it's the 596 - > same software, the same vendor, the same management server can 597 - > you know, very very similar interface. You don't have to 598 - > go and use a different tool and therefore have to 599 - > become experts in more and more pieces of software. It 600 - > really lends itself well to bringing efficiency and operational excellence 601 - > to both the IT and to the OT. 602 - >
Speaker 3: What strikes me about this, mate, is that we're talking 603 - > about using the same backup technology suite for OT and 604 - > IT both. You know, I'm reminded that this a lot 605 - > of people nowadays, when they hear the phrase iotic convergence, 606 - > they think connecting those networks up, you know, the cloud scenario. 607 - > We talked about the original vision back in two thousand 608 - > and five when the Gardner Group coined the phrase, you know, 609 - > IoT convergence and coined the phrase operational technology. The original 610 - > vision is that teams would come together. Why do we 611 - > have two sort of centers of expertise in the company, 612 - > one for SQL server for use on the OT side 613 - > and one for SQL server for use on the IT side. 614 - > It makes no sense combine these teams. It's the same knowledge, 615 - > you know. Why are we using a different relational database 616 - > on the OT side than on the IT side. When 617 - > you know, we have an application that needs a relational database, 618 - > it can use any one of them. On the OT side, 619 - > we use SQL server. On the IT side, we use Oracle. 620 - > But on the IT side, long ago we bought an 621 - > enterprise license. We can deploy as many Oracles as we 622 - > want free of charge. Why would we keep buying the 623 - > same solution from a different vendor. Combine these, uh, you know, 624 - > increase our leverage with the vendor, reduce the amount of 625 - > training required. This was the original vision for i OT convergence. 626 - > And you know we we see that here in the 627 - > a Chronus Solutions saying you can use the same technology 628 - > across the board. You know, it was only pushing a 629 - > decade later that people started talking about ITOT integration in 630 - > terms of connecting the networks. And of course everyone almost 631 - > everyone connects the networks today. But the original vision is 632 - > this vision, which is reduce the complexity company wide. Well, 633 - > this has been great, Thank you Steven. Before I let 634 - > you go, can you sum up for our listeners? What 635 - > should we take away from this episode? 636 - >
Speaker 1: Well? Thanks, it's been a great conversation. So you know, 637 - > if I was looking for what I'd like people to 638 - > take away, it's that we can provide a reliable, secure 639 - > way to back up and more importantly, recover their OT 640 - > environment as well as their IT environment in a way 641 - > that is going to increase up time and help those 642 - > sleep at night, so things are properly protected. And we 643 - > can do all of that in a way that is 644 - > closely aligned with their vendors. With the OEM relationships that 645 - > we have, it's easy to deploy and easy to manage. Ultimately, 646 - > if you're looking for more information, you can certainly go 647 - > to a coronas dot com. But if you want to 648 - > start a conversation, please connect with me on LinkedIn and 649 - > I'll get you connected with one of our engineers and 650 - > we can do a deep dive. 651 - >
Speaker 2: Andrew, that appears to do it for your conversation with 652 - > Stephen Nichols about recovery. You know, not the maybe sexiest 653 - > topic we've talked about on the show, but equally important 654 - > to anything else we've discussed, and important that some folks 655 - > are focused in this area. 656 - >
Speaker 3: Absolutely, it's you know, it's recovery. Backups are underappreciated. There 657 - > sort of happened silently behind the scenes until you need them, 658 - > and then it's a mad panicle. And we've covered that somehow, 659 - > haven't we And yeah, you know, here's a way to 660 - > cover it. The new buzzword in OT security is resilience, 661 - > meaning if you get hacked, when you get hacked one 662 - > of these days, minimize the impact of the attack. And 663 - > one of the ways to minimize it is rapid recovery, 664 - > you know, And there's operational benefits. I mean, on we 665 - > deploy equipment industrial settings that often that's expected to last 666 - > well over a decade, which is sort of beyond the 667 - > lifespan of you know, a lot of it. Equipment stuff 668 - > wears out, you know, the ability to replace with slightly 669 - > new or slightly different hardware and recover and keep going. 670 - > That's an important operational benefit forget cyber attacks. And you 671 - > know what I didn't know about, you know, was the 672 - > clever bit of of offline anti virus scanning built into 673 - > the solution, saying you know, both anti virus scanning and hey, 674 - > you're not supposed to be encrypting those files. Why is 675 - > there new copies of these showing up? You know, ransomware detection. 676 - > You know, these are our lovely augments to sort of 677 - > a like you said, a mundane backup capability. 678 - >
Speaker 2: Well, thanks Stephen, for being on the podcast with us. 679 - > And Andrew is always thanks for speaking with me. 680 - >
Speaker 3: It's always a pleasure. Thank you man. 681 - >
Speaker 2: This has been the Industrial security podcast from Waterfall. Thanks 682 - > to everyone out there listening.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.