
Effective Engineering Manager · 2025-06-09 · 1h 19m
Jeremy Franzen draws on three decades of international operations experience - spanning healthcare, Bitcoin, finance, and retail - to address a persistent friction point in software delivery: the handoff between engineering and operations. Rather than the traditional 'throw over the wall' model where engineers ship code and ops maintains it in isolation, Franzen advocates embedding ops personnel directly within engineering teams (a practice aligned with modern SRE methodology) to surface problems early and build shared understanding before code reaches production. At TiVo, he built a 24-hour Network Operations Control Center staffed with engineers, backed by detailed playbooks and runbooks that enabled incident response without constantly waking engineers at 3am. He emphasizes that both teams must align on SMART goals - specific, measurable, achievable, relevant, and time-bound - rather than letting engineering optimize for throughput while ops optimizes for availability. His work at Shutterfly, managing 400 petabytes across 15 million customers with a lean ops team, demonstrates how this collaboration enables massive scale. Critically, he stresses automation as force-multiplier: the philosophy is simple - if you do something manually twice, automate it the first time - because machines detect and resolve incidents orders of magnitude faster than humans.
Embed operations members directly in engineering teams during development, not just at handoff. This allows ops to understand code design decisions, provide early feedback on operational constraints like database locking or memory management, and prepare production environments ahead of time.
Playbooks describe what the code does and why; runbooks provide specific step-by-step instructions for responding to known failures with escalation paths to engineers. Better documentation means fewer 3am calls because ops can resolve issues independently first.
SMART goals are Specific, Measurable, Achievable, Relevant, and Time-bound. Engineering and ops must define them jointly because they naturally optimize for different things (engineering for throughput, ops for availability), and misaligned goals create conflicts when problems occur.
If you do something manually twice, automate it the first time. Machines detect and fix problems vastly faster than humans can wake up and respond, making automation the most cost-effective way to reduce incidents and on-call burden.
Use non-confrontational language ('what can we do together' rather than placing blame), assume good intent from both sides, and share specific details about operational pain points so engineers understand the downstream impact of their code.
Computed from the transcript - who did the talking, and the words that came up most.
We are featuring a guest Jeremy Franzen, an Operations expert with over 30 years of experience running Ops at public companies and startups. Jeremy shares his expertise in building strong collaboration between software engineering and operations teams through open communication, shared goals, feedback loops, joint automation and monitoring, shared security practices, knowledge sharing, and coordinated approaches to incident and change management, all tracked by common performance metrics. In the end, Jeremy provides a checklist that our listeners can start using today to build and run quality production environments together with their Operations teams.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Uh, welcome, uh, to the Effective Engineering Manager podcast today. Today we have a great guest, uh, Jeremy Franzen. He's, um, a friend and a colleague. We work together at two companies at uh, uh, Trade Beam and Shutterfly. And um, uh, Jeremy, uh, um, is an, um, operations leader with, uh, 30 years of experience building international teams and uh, supporting engineering teams of all sizes. Jeremy, uh, welcome.
Speaker B: Thanks, Lava. Appreciate you inviting me on the podcast, man.
Speaker A: Yeah, yeah, Good to see you. And, uh, what would you like to talk about today?
Speaker B: Um, a lot of what engineering does and what they think about and how they interact with the guys that run the platforms, um, is really one on a kind of surface today. Right. Because what the engineers do, how they think about what they do, how they secure their code, how they document their code, and how operations takes it, deploys it, runs it, triage it, um, makes it stay alive. That, that's a lot of how the teams interact. Right. And how engineering managers can help support operations better and how operations can better support engineers as well.
Speaker A: Yeah, I think it's an amazing topic because I've seen when you and I worked together, I think it was awesome. But I also seen, you know, situations when it wasn't great. You know, sort of like a throw over the wall, uh, type ah, of an operation. And it wasn't fun. So. Well, maybe you could talk about yourself. I mean, today we are here, now, um, you know, you've been doing it for 30 plus years. How did you, how did you come to this moment? So, tell, um, me everything.
Speaker B: When I was young, I got, uh, a bug and decided to go in the Marine Corps. And the Marines decided they wanted to put me in technology. And I ended up doing this. And I'm glad I did what I did. I've been around the world, more than 40 countries. Um, I built teams that span the globe. I've worked at a bunch of Silicon Valley startups and some big companies in Silicon Valley, um, from healthcare to Bitcoin to finance to retail, uh, you name it. And everywhere I've been involved, um, the focus has been on operational excellence, usually with a focus on a foundation from itil, which is an international standard for operations. And I've had great fun and great success doing what I do.
Speaker A: Nice. Yeah, and I remember that I, um, mean when we met first time, I think, what, almost like 23, 24 years ago, right. I think you were maybe my first, you were my first, I mean, my first real ops, you know, uh, uh, uh, people and um, uh, really, uh, you run a tight ship. You Know, everything was smooth, uh, well, when we didn't mess it up. And, uh, um, that was good. Um, so, uh, and that's good stuff. So, um. Well, maybe we should, uh, dig in,
Speaker C: uh,
Speaker A: tell us everything. I mean, I'm here to learn, and I hope our listeners will learn, too.
Speaker B: Absolutely.
Speaker A: What's on your mind? What does it take?
Speaker B: One of the first things you brought up was when we didn't mess it up. Right. And a lot of people tend to think that, um, when engineers write code and they hand it off and they walk away, it's just going to work. And operations guys, uh, have this habit of saying no. Always the first thing they say is, no, we need this thing. No, I need to buy this. No. Right. And it's a really bad habit, um, because engineers sometimes make mistakes and ops guys sometimes make mistakes. Right. I mean, both teams can, you know, both teams can make a mistake, but if you're not tight with the other teams, if you haven't taken the time to understand who they are and what they do and how they. They think, how they work. Right. Knowing what they do, okay, that's important. But not as important as knowing who they are. Right. So if your leaders are writing code that goes into production, and they have a team of people that are focused on automation or SLAs or KPIs, or, you know, high availability, all of the fun stuff that operations guys have to worry about, um, they need to understand the mindset of the ops guys. Right? The mindset is something's going to break. Something's always going to break. And no matter how good your code is, your code can't make memory not fail or disks not fail or networks not hiccup or. There's so many variables out there that can be a problem. So when the ops guys are all grumpy and always saying no, I recommend to them to start saying, not yet, or let's talk about when we can, versus always no. And on the engineering side, I recommended. And when I was at TiVo, I did this. I embedded my operations guys in different engineering teams. So one of my ops guys would sit in the product team, and one would sit the engineering, and one would sit in the hardware, and they would bring back issues that surfaced from those teams back to us. So we could understand what problems they're having, where they're headed, and we could give them advice as to, you might think about doing it this way, or we might understand the new path they're going down so we can prepare the environments ahead of time. To make it easier for them to be successful. Um, if we keep things in silos and engineers just work with engineers and when they're done, they throw their code over the wall and QA picks it up and they test it and they say, okay, this could. And then they throw it to the wall to Operations to deploy it. We, whether it's with Terraform or Kubernetes or whatever automation tools are out there, um, and then the ops guys have to get it up, get it running, get it monitored, get all the different key performance indicators, um, that engineering things are important and that operations things are important into a dashboard to trend and trace and figure out what's going on. When there is a problem, Operations is going to turn to QA and say why didn't you see this? And then QA is going to turn to engineering and say, how did you let this happen? And that's not productive, right? So if we get ahead of the curve before they put the code into production, before they even put the code into qa, if we can short circuit that and start being there during the creation of the code. If they have scrum teams embed a member in the scrum team, they may not be very productive or they may be able to give ideas or nudge people in the right way to make them think about. If you have like on a database, you lock a table before you update a record, that may prevent other things from happening and slow down the system or could create a race condition. There's a myriad of things that different viewpoints help tease out of a design.
Speaker A: You know what, it's pretty cool what you're saying because I think what you are describing is, um, at this point, I think it's a well known model which is called site reliability Engineering or sre. Right. Where teams, uh, ops teams are embedding with engineering teams and engineering teams, uh, if it's, you know, if they're open to it, uh, sitting on the calls and then wearing pagers maybe half of a day a week. Right. And that creates this natural communication channel and uh, understanding channel. Right. The challenges that each team is facing and how to approach them together rather than uh, throw the ball. I love it. It's good stuff.
Speaker B: Well, I mean um, at TiVo we had, you know, data centers around the world. We were supporting with great ops teams and when there was a problem, we would try to solve it first. How would we try to solve a problem?
Speaker C: Right.
Speaker B: Which is the biggest thing is how are we going to understand what these guys put into our production Environment, how are we going to support it? How are we going to maintain it? So I built a network operations control center, a noc, and the NOC hired engineers for it. They were all solid. It was 24 hours a day, lots of cool graphs up on the walls, and they had access to everything they needed to touch to try to fix a problem. But there's so much code and there's so many variables. How do they know what to look at? Well, it comes back to documentation, right? We would have playbooks and runbooks. Playbooks would describe the code, what it's supposed to do, why it was doing what it did. Right. And some of the technologies used in creating whatever the environment was. But a runbook was a very specific set of instructions that said, when this code is running, if you see this kind of failure, do this thing on the command line or restart the system or whatever the action was. And if it doesn't work, escalate to these people, the guys who wrote the code, right? And you may call her up and say, hey, it's 3:00' clock in the morning. But right now it's 11:00' clock in the morning in England. And their systems are all down because something in this code you wrote isn't working right. Something's not there. Can you jump on a call and help us fix it?
Speaker A: Right.
Speaker B: The better that documentation was, the less often engineers got phone calls at 3 o' clock in the morning, which is awesome for them because no one likes getting up at 3am to try to fix a problem at work. It just, you know, no one really has that desire. So the teams started getting a whole lot better at figuring out how to do what to do and how to write documentation that was actionable for the operations guys to get their job done. Right. So if you think about it, the better the documentation, the better it creates an incident management protocol that lets us manage an incident without having to pull in engineers unless it was absolutely critical. Right. We also have root cause analysis meetings after failures like that. So we would have an RCA and the meeting would be the next day or the same day if possible. And we would pull in the engineering leaders, we'd pull in the engineers who wrote the code, the engineers operationally who helped resolve the problem. And we'd also bring in QA because we need them to understand what the failure was so they can test for it in the future and prevent it from occurring again. Right. Streamline the process constantly, revet every and tighten up that code so when that problem would occur again, there's error checking that would cause it to fix itself or alerts that would give us a heads up that something is going to fail.
Speaker A: Yeah, good stuff. And I agree if you think of it right, imagine there's no docs at all. Yeah. How are you going to run the
Speaker C: system if there are no dogs?
Speaker A: What is supposed to do? How do we know it's working? And uh, yeah. And um, how is it working? And I uh, think uh, this um, I think startups early stage is sometimes maybe until there are no customers it's maybe acceptable. But the moment the first customer shows up things are going to go, things are going to go break and then um, you need to know uh, how it works, how do you know it works and what it takes to keep it working and um, how do you know it's not working and what it takes to go bring it back. Right. So good stuff. Uh, I love it. And so we talked about um, um, and you brought up before, um, need for the shared goals and objectives. Right. Can you dive a bit deeper in this?
Speaker C: Sure.
Speaker B: So if you think about the latest thing is Scrum, it was not the latest thing. It's been around for a long time but it's been adopted a little more widely now. Um, some of the Scrum methodologies think about um, smart goals, right? And you try to make sure that these goals have a very specific um, purpose. Right. You want everybody to buy into what this goal is and define them together. Right. Because engineering's goal might be to process 10,000 transactions per minute and operations goals might be completely different. They may be focused on high availability, they might be forced on disaster recovery, they might be focused on archiving or backing up data and they're not really focused on performance. Making sure the teams work hand in hand to define the goals. Make sure that each goal you define is measurable. You should be able to measure just about anything. Make sure it is actually something that can be done, it's achievable. Make sure that it's actually relevant. Right. And um, I've seen from time to time where people come up with this cool new thing they heard about and let's do that. And it's not really relevant to what we want to do from a business perspective. Try to tie it back to revenue if you can. Right. Either revenue, customer performance, customer satisfaction, something that is tied to that makes it very relevant to the company. Also make sure there is no, we'll get to it when we get to it kind of thing. Make it boxed, put it in a box, time, box, this thing segment out everything that needs, all the dependencies around this goal and say this needs to be done by this time. Right. Because if you put limits on how long it'll take to build it, people start assigning, uh, urgency to the task and try achieving that task in a certain amount of time. Right. Once it's measurable, achievable, relevant, and time bound, once it gets to in the. From the engineering team to the QA team, QA guys know what they're measuring, what it's supposed to do, why it's relevant, and when operations gets it. Since it's been defined well, it should have been documented well. And we should have whatever key performance indicators are necessary to measure it and maintain it. Yeah, it makes things a lot faster, a lot smoother. Especially as you go into production and you start going to, like, multinational, multi terabyte or petabyte systems.
Speaker A: Yeah, yeah. I remember at, uh, Shutterfly, I think. How much did we have?
Speaker B: Like, 400 petabytes.
Speaker A: Yeah, 400 petabytes. That was insane. I remember we just. First time I heard it, oh, my God. Yeah, that is pretty insane. And we had what, like, uh, 12, um, 15 million, uh, um, customers. Right.
Speaker B: Uh, just insane.
Speaker A: Yeah, that was insane. I think even by modern measures, I think that was a pretty insane scale.
Speaker B: Uh, yeah, it was. And if you think about the operations team that ran, that was really pretty small. So we had optimized ourselves internally to figure out the absolute best way to get what we needed done, done at minimal cost and maximum productivity. All right. We had an amazing guy. Mike Coogler was in charge of that storage. Um, and dude was absolute rock star at what he did. But he couldn't have done it alone. Right. He needed to work a lot with the engineers to make sure they understood what limitations he had to work with, which mean they had to stay inside of the boxes he drew, and they did, and they were able to make it work flawlessly. It was amazing.
Speaker A: Yeah, flawlessly. Um, I'm not sure if you remember. Um, um, um, there was a great partner I had in, um, uh, Shutterfly. Uh, Murthy Adari.
Speaker B: Yeah.
Speaker A: And, um, yeah, he's been great. And we worked together, um, um, um, on the migration to the cloud. That was really a serious task because we had to run the whole, uh, system, um, as it is, and at the same time moving to, uh, the cloud. And I remember that Murthy and I, we came up with this idea that we would, instead of doing, like, a gigantic cutoff, we would move it, like, piece by piece and at some point both, uh, in house, collocated plus um, cloud would work together. Um, and just an example how important it is to have uh, to have this open. Um, like you said, having a shared objective. Right. We, we work together. It's not like you do your thing, I do my thing. And um, and flawless. I think I remember that, um, um, when I came to Shutterfly, I said we are going to build a flawless system. We are not going to build a perfect system.
Speaker C: But from the customer point of view it should be flawless.
Speaker B: Yes.
Speaker A: Uh, yeah, it was uh, quite a challenge. So. Yeah, especially, yeah, I mean especially this
Speaker B: stuff we interact with isn't flawless. Most of the stuff we use, whether it's your banking app, whether it's your Slack, whether it's your web browser, whatever else in the back of it, behind the scenes stuff that users don't see, there'll be a lot of panic going on. There may be a lot of people fighting a fire to try to keep the user out of the picture. Right, exactly. So flawless is to the perception of the user and the customer.
Speaker A: The moment you say perfect. Right. You have to, you know, you need
Speaker C: a 10x budget to do to get the same thing you get out of flawless.
Speaker A: Right. It's okay to be imperfect, but it's not okay not to be flawless. Yeah, yeah, good. Uh, stuff. So, and uh, uh, we talked about um, so you mentioned the goals. How do we. And it sounds like the um, bidirectional communication between teams within engineering and ops is super critical for the company to be successful. How do you approve? What is your guidance on maintaining those uh, feedback loops?
Speaker B: Well, feedback loops are huge. Right. And uh, there's a huge number of lessons and learning and all sorts of intelligence, um, emotional intelligence training around feedback loops. Right. Uh, the best thing you can do when you see something, say something. It's a really simple concept and let people know. Don't ever be confrontational. Right. Engineers don't write stuff trying to make operations guys mad. It just doesn't happen. That's not in their mindset. They're not staying up. And I thinking of how can I make these guys get up at 3am and sweat for a couple of hours while they get stuck working again.
Speaker A: Yeah.
Speaker C: Uh, you know, but sometimes I think from the ops point point of view it might feel like this.
Speaker B: Yeah, well that's exactly perspective. When you're up at 3 o' clock in the morning three days in a row trying to, trying to fix this, you're always thinking to yourself, uh, how can we make this not happen anymore? This is ridiculous. Right? So when you provide the feedback loop to engineering, say, hey guys, I experienced this, I saw this and this is what I did to resolve this. Or we saw this and this is what we did to resolve this. What can we do together to make sure that we get this thing to perform better, to stay alive longer, to fix this memory leak, to fix this bug, whatever the case may be. How can we better align our deployment strategies and our operational strategies so that I'm not getting up at 3 o' clock in the morning because a couple of days like this I'm going to get grumpy and then everyone's going to be walking on eggshells. We don't want that. Right? So feedback. If you don't give them feedback, if you don't let them know what pain you're in, they don't know you're in pain.
Speaker A: Yeah. You know what I like? I like when you say, uh, how can, what can we do together? Right? Because if, uh, if, if you continue this, you know, over the fence approach, uh, when someone comes out and says, what can you do to make it better? Right? You, you, you, you, you take, you put work on someone else's plate and you know, and keeping your plate clean, right? And I think this, uh, what can we do together is just super critical. And the, the ops and the engineering teams, I really believe they should be sitting in the same real or um, virtual cubicle areas. Right? You should be able to, you know, stick your head over, you know, the, the cubicle wall or if you don't have walls, just, you know, turn your head and have a conversation, A uh, polite conversation. I think this whole polite thing is very important. Like emotional intelligence is important because
Speaker C: it's
Speaker A: easy to get upset if you keep waking up at 3am or if someone keeps finding bugs in production in the code you wrote. But I think this whole thing, it's important to know that we are in this together and people around you and with you are in the same boat. Right? Um, that's critical.
Speaker B: You know, and a lot of, a lot of the stuff that, over my years that I have found, um, that has happened, um, automation tooling, like CI, CD pipelines, um, getting, you know, SRE involved, like methodologies involved, um, helps like continually improve your product. Right. The more automated you can make things and the better the tooling and the testing are. The, Then it. Manual errors get dropped drastically.
Speaker A: Exactly.
Speaker B: I mean, it's really, really important to think about how we can get humans out of the picture because most of the engineers that I have worked with have an amazing linear chain of thought and they know that when X is done, Y will follow. When Y is done, Z will follow. When humans get involved and X is done, sometimes they'll do Y and sometimes they'll do A, or they might fat finger a command and accidentally delete something. Right. The less humans have to touch the box, the better off we are. Right. Ideally, operations guys want to be the guys that go in and unplug and replug in a system, uh, pull out the box that burned out a CPU and put in a new one, fix the network port that went bad on the switch and they're not worried about the code. Right? Yeah, uh, uh, that's ideally what they wanted.
Speaker A: That's, yeah, that's the holy grail. And if you look, you know, if you look, I'm 100, I'm, I'm so pro automation it's not even funny. In fact, I'm a bit on the extreme side. I, my approach is not approach. My philosophy is that if it's not automated, it doesn't exist. Maybe you tell me that it exists in your head, right? But from where I stand, it doesn't
Speaker C: exist because I cannot know what's inside your head.
Speaker A: But, but, but the big guys like Facebooks and Google's and Netflixes of the world, uh, it, it is like that and I've seen it at Facebook, right? You just, you know, there's a guy on the data center, he sees red, red, red, um, you know, red LED blinking. You know, he just walks up with a card, pull out, pulls out the box, put, put, put the box in, it's done. Yeah, that's, that's, that's all this is, this is all human involvement. And I think we're all joking like
Speaker C: well, maybe we'll just replace those with the robot.
Speaker B: Well, and there are companies right now that actually do have robots, um, out of San Francisco that are working with Nvidia today that are assembling computers, including putting the cpu, the memory, putting in the box, bolting it down, getting it all ready for the data center. And soon there'll be robots that'll actually be able to rack it, stack it, plug it in, cable it. You'll still need for now, at least for the next couple of years, humans to make sure the architecture is right, the network is right, security is right. But yeah man, I mean not too distant future, right? Engineers are going to be less hands on the box and a lot more logical. And a lot more architectural. Um,
Speaker A: let me ask you on the automation a bit. Uh, it's sort of like a double click. So what is your personal uh, approach? How do you approach automation in practical terms? Like imagine the. Let's say people and money is not an object. What, what is your approach? Okay, people and money is always, always,
Speaker B: it's always going to be, it's gonna
Speaker C: be okay, let's, let's assume it is an object.
Speaker B: But if you have to do something manually twice, you should have automated it once. Second time you're doing it, it should be with automation. Right, exactly. Once you see a problem, automate it. Because if a machine, a computer can catch the problem, it can fix it much faster than it takes a human to wake up, get out of bed, come to cognition, log into whatever, Zoom, um, call or WebEx or Google Meet or whatever, get online, get debriefed on what's going on, jump into the system, investigate it and then fix the problem. Right? That's going to take time. And time is, is your enemy, right? I can make more machines, I can make, I can write more code, I can make more money. I can't make time. So automation saves time. The less downtime you have, the more money you make as a company. So that is the main objective of operations guys, is to make it foolproof, make it highly available, make it globally available, make it incredibly performant and uh, always up, always there, always working and secure. Right? Um, from an automation perspective, if you can't automate it, you probably doing it wrong.
Speaker A: Yeah, yeah, it's a good, a good point. What my personal observation, and I've seen it, I mean I don't have to like go too far in my career, sort of repeatable theme. The. So I was always able to bring uh, the engineering and the operations teams together. Right. We would always go through the uh, the incidents together. We will do RCA together, we would do postmortems together. Right? And when I show up on the, on, you know, in the company, they don't have it. And then I establish the process, we teach the people, you know, it all works. Everyone's happy and it's good, it's good, good to have this whole, you know, we're in the same boat, like sometimes literally, almost literally in the same boat until like um, uh, one incident at uh, sleep number, I think I was up until like 7am from like 11pm like it's botched deployment. Um, but here's the thing. Teams learn to build um, runbooks.
Speaker B: Yes.
Speaker A: Right. And they get Everyone understands, everyone gets good at it and they get updated. But the moment you say, hey guys, we have to automate, right? Because, okay, we've done it like three or four times, right? And everyone is looking for faults. And, you know, it takes hours and hours. And um. And if you think, if you think of it, right, the one hour of E Commerce Company. I mean, if you remember Shutterfly, right. We are, we are talking about like million, million and a half dollars, right? For this money you can, you can
Speaker C: buy a new team.
Speaker B: Yeah.
Speaker C: Buy all the tools you need and automate everything. So.
Speaker A: But the argument always was, well, it's always, it's a priority against the product development. It's a priority of the roadmap. It's a priority of this and this against this and that. And I've never been able to convince the whole thing, hey, guys, this is a priority too, right? The automation, like you do it twice, third time you must automate.
Speaker C: It's hard.
Speaker A: How do you make it happen?
Speaker B: Sell that, right? So the way I've been successful with a lot of stuff, um, in, in the civilian world, in corporate worlds, is to break it down cost wise, right? Shutterfly is a great example. They made a billion dollars in a year, right? Just round number. They crossed a billion dollar threshold. There was a party, everybody's excited by it. And then I sat down and I ran the numbers. I said, okay, let's take a billion dollars, right? A billion dollars is a lot of money. So if we look at a billion dollars, we say, okay, there's 1 billion. And we divide that by 365 days. And then we divide that by 24 hours and the 60 minutes and the 60 seconds, then you're looking at literally $31.70 per second that company's making. Okay? And it takes my guy an hour to fix a problem versus a computer fixing it in a second. That one second the computer took to fix the problem cost my company 31 cents. But if it took my guy 10 seconds, that'd be $3 and 10 cents. A hundred seconds would be 31 bucks, etc. Etc. Right. Running the numbers just that simply and showing them to finance because they've got a big stick in the company and say, guys, look, It'll cost us $25,000 because of this incident we just had where the site was unavailable for two hours.
Speaker A: Yeah.
Speaker B: If we would had automation, that would have cost us probably A$50. So our bottom line would show a drastic difference based on the automation efforts versus manual intervention.
Speaker A: So you're Saying that it's the best is just take the numbers essentially go to uh, planning, uh, the people responding, uh, responsible for the budgeting and for the planning of, uh, allocation of people and money. Show them with the money. What is it going to cost to have a notage, right. And what is going to, and how. What is going to have to do it once automatically and uh, not ever to have that outage again. Right.
Speaker B: I mean at Shutterfly, when they made a billion dollars in a year, one hour of outage and looking at $114,000 loss for the company.
Speaker A: You know what? I think your math M is wrong actually because um, most like most e commerce companies like Shutterfly, they make all the money like 80, 85% in Q4. Yeah, it's, it's not even, it's not 12 months.
Speaker B: Oh yeah. So I mean I flatline it out across the year just to, you know, because it's easier not to do that. But if we looked at it over a three month window instead of a 12 month window, the $114,000 an hour would become closer to three quarters of a million dollars for an hour downtime in the fourth quarter. That's huge, right? That's absolutely ridiculous money. If we can't automate stuff, if we can't get things always up, always on in a solid, documented, secure way. And another thing I would say is security. Right? Security. Infosec guys love to be hated.
Speaker A: I don't want to work in security.
Speaker C: I'd rather have someone who wants to work in security.
Speaker B: Yeah, Somebody who wants to take that load. Um, when you're writing code, there's some basic tools people should use, right? Because operations guys at the end of the day are required to secure the environment. So the infosec guys give them input on what they should do. They'll do penetration testing, they'll do um, application evaluations. They'll look at the supply chain for your uh, applications to make sure they don't using libraries that are corrupt or compromised. Um, but engineers can take simple early steps as they're writing code. Look at the OWASP top 10. Those are the 10 most and they're updated quarterly. The 10 most common faults that engineers step around to get their job done faster, right? They'll write, they'll write a secure key in a config file in plain text or usernames and passwords in plain text in a config file. Well that's cool. It makes their job easy, fast and simple. But it's a huge security vulnerability, right? Some of those things, uh, they're required. But how do you know, right? How do you know as an engineer what you need to do? Well, knowledge training, right? If you're not sure, ask, uh, a lot of engineers, operational engineers as well, don't like to ask questions because they don't want to seem like they don't know, right? They don't want to seem ignorant. But I've yet to meet anyone who knows everything. No one can. So if you're not constantly training, you're not constantly learning, you're falling behind. That's engineers, that's operations, that's infosec, that's every team, right? Always learn, always train, right? Figure out what the best practices are in Operations, in DevOps, in SRE, in InfoSec, in QA Engineering. I mean constantly training each other, cross train each other as well. So when operations guys are talking about how they're going to create this virtual ip, so this kind of load balancing is going to be round robin or weighted or least used or whatever the methodology might be. Teach that to the engineers so they can write their code with that as part of their thought process because they may write their code differently if they think it's going to be round robin. Every request is going to be a one time request, in and out, so they'll have to have their own key. There'll be no session state on the user base or it's going to be assigned to a server so you can keep sessioning the client. You don't have to worry about a unique key on a per transaction layer. There's a whole ton of stuff that you change your mindset about as you write code. The more knowledge you have of how it's going to be run. You follow me on that?
Speaker A: Absolutely. And I think you're absolutely right because the security, it's a tough topic because uh, the malicious agents are always like a half of a step ahead otherwise they wouldn't be able would just solve all the problems and you know, they would go away. But um, but I think you have a very good point that we engineering managers must continue educate their teams and self educate both on how to keep the system secure and sort of like really like take responsibility, not expect the security guys to help, uh, their job. They will help but I think it's a joint effort because plenty of things can be done very much upfront like scanning the code base for like you mentioned the keys in the configuration files or uh, known vulnerabilities and they're like static code analysis tools. You Just pull, plug it in into the CI CD system and just not let it through. Um, but I think educating the engineering both for what it means, what vulnerability is and how to prevent and how to stay on top of it. It has to be like uh, a, uh, P1. Right. Because who cares if your system is fast and reliable, if someone hacks you and then it leaks into the public and public markets, you know, the damage is going to be, can be in tens of millions and even billions of dollars. Right. There's no, it shouldn't be optional. And I think it's like it's oftentimes just skipped somehow. Right. Or ignored.
Speaker B: Yeah.
Speaker A: And it's unacceptable.
Speaker B: Uh, your reputational damage, especially as a small startup right up.
Speaker A: It's going to be a death.
Speaker B: Yeah, it'd be death for a startup. Disney gets hacked, they probably get hacked more often than we know, but no one really pays attention because their brand is so powerful.
Speaker A: Yeah, right.
Speaker B: So the damage to them is probably less than the damage to a small startup in the Silicon Valley trying to get their reputation built. And when they get hacked, it's much bigger news because there's much more to lose.
Speaker A: Right, right.
Speaker B: You know?
Speaker A: Yeah. Right. Yeah. It's not optional, man. I mean I've usually when I come, um, I take on a new job, first person I find is the security people. And my question is, how can I help?
Speaker C: I don't want to deal with this. Let's solve it before it becomes a problem.
Speaker B: Yeah, let's get that um, back right now. Right away. So one of the things about security, this leads into the next one is logging, Right. Something's going to happen. Whether it's going to be an operational failure, an engineering failure, security failure, whatever, things break. That's just the way the world. Unfortunately, no matter what we do, not everything works all the time. Ah, well, how do you triage that? How do you learn from it? If you're not logging what your code is doing, we don't know what broke. We may be able to trace it down to what system broke, but not why the system broke or what broke in the system or security may not be able to look at those log files for forensically and say this bad guy from this place came into our system this way to compromise us. So logging and of all things in the world, monitoring, um, become really critical as you run your code. Right. There can be too many logs, but at the same time there can be way too little logs. So there's going to be uh, a tuning effort that, that's going to be required over the lifetime of any software engineering project that logging can either be dialed up or dialed down based on what team needs what. Right? So I would recommend that people make that a topic to talk about with the ops guys and the security guys. Hey, look, folks, this is our code. This is what these log files mean, and put it in the runbook. When this code kicks out, this message in the log file, this is what the problem is, and this is how to solve the problem. And if this doesn't work, escalate to one of these three people. And the reason you put three people in an escalation list, one person might be on vacation, one person might be sick. So you need a third person just in case.
Speaker A: Yeah, good point. Um, and, uh, so what is your take? I mean, I've been there and I've been in the meetings where, you know, the logging systems would, you know, eat more money than the company makes. So. No, I'm serious. Like, I'm not going to give the names of the, you know, companies that are amazing companies, but the sort of like a. I, uh, call it runaway logging. So how do we go from. Okay, how about we overlog? Personally, I think overlogging at the beginning is important because then you cannot narrow it down.
Speaker B: Right.
Speaker A: But at what point we have to start paying attention to narrowing the scope of what we log and what is not beneficial. What's your take?
Speaker B: Uh, that's what I was saying. There can be too many logs. Right. When you look at tools like Splunk or other monitoring tools that ingest logs to help you trend and track data, they're really cool tools. They charge per quantity of logs. So when you're first starting out and you're logging everything at debug level, you're pulling out gigabytes and gigabytes of logs, and they're making a ton of money, even though you're getting minimal value. But, um, conversely, once your systems are up and live and replicated, running around the world in multiple locations, in multiple clouds, and you've tightened up the logging to be just around performance and availability, and then your logs are now a lot smaller because you had the time to chop into there and figure out what's useful and what's just noise. But you actually trade the cost of massive logs versus massive systems, and those costs still stay high. I mean, there are startups out there, like Cribble, um, who are actually running their business on how do I dedupe log files? So you only get stuff that's new. Cool. But at the same time, maybe we should write our, uh, code in such a way that logs in such a way that we don't need to have massive log files. But it does capture the critical information which ties into. What are the performance indicators you're looking for when you sit down and define. These are the key performance indicators we're looking for. Whether it's response time, transaction time, whatever memory pool size, then logging that. So you can print that.
Speaker A: Okay, so you're saying that define what is meaningful across, um, business ops and engineering. Log that.
Speaker B: Yes.
Speaker A: And then if you want, you add more. Right. Because. Yeah, I like it. I think it's a great idea. It's a very easy to follow guidance, I think.
Speaker C: Yeah.
Speaker A: People usually like, you know, log everything and then, you know, then, then you try to. Okay, what is important?
Speaker C: Who the hell knows? I mean, you have like five exabytes of logs.
Speaker B: Yeah. And then you know who's going to take the time to go through five exabytes of log. Right. Humans don't have that kind of time. So you start tools and the tools cost money and they cost processing time and processing power. Like, oh, I'll grab for that line item. Now if you've got five terabytes of logs and you're grabbing for one line item, it's going to take a couple of weeks. It's not going to work.
Speaker A: Yeah, I know, I know. We've done it. I mean, the, so my, my pet peeve always been with logs. I mean, sooner or later it always comes, always comes. Someone looks at the, the, the, the, the expense. Um, and like, why is that this
Speaker C: particular logging system costs like $10 million.
Speaker B: Right.
Speaker A: And then, and then you get another, usually you get another, hey, make it happen. Right. Make it disappear or like cut in half. First thing I personally have done, you log into the system. No matter what. Those vendors from open source, commercial, hyper commercial deduplication, uh, what's simple is just, you just scan. You don't even need tools. You scan like two, three pages of logs and use your brain to see if a particular line repeats without changing. Right. So lines repeating without changing. Right. It's a good signal. They put some, some log somewhere in the loop and it's just spouting, you know, terabytes of data, um, not, not creating any value. Right. So, and if I've heard, hey, we use it to track if the system is alive, man. If you need to output, you know, a terabyte of data, maybe you should create a counter.
Speaker B: Yeah.
Speaker A: And read counters instead of, you know, Putting, you know, putting, you know, 2, 300 bytes of strings, you know, every microsecond.
Speaker B: Exactly.
Speaker A: I think just applying minimal smarts, you don't even need a. I call it ni.
Speaker C: Natural intelligence.
Speaker B: Yeah. I mean, things like that. Right. Just common sense goes a long, long way and it's a very uncommon attribute, unfortunately.
Speaker A: Yeah. Well, I think this is where this sort of, um, you mentioned continuous education, continuous improvement, Right. If the team didn't learn anything in a year, it means that something is seriously wrong.
Speaker B: Yeah, absolutely.
Speaker A: Cool. What else?
Speaker B: Oh, man. Um, change management. Let's talk about that for a second. When you have production systems that are up and running and you need to make a change to it, right. Um, there is a sign of maturity in the company when you're a small startup. Um, you make a change right now, you just jump on and you change it. Right. And as you grow and more teams get put in place, you start limiting access to least used privilege or least needed privilege on any system. And once you start taking people's rights away, you have to put process in place to fix that. Because if engineers are writing the code and there's a problem, they can jump on and fix it on production. Cool. Once you start having customers and once you start having revenue tied to that, you need to start backing away from having the wild, wild west of just jumping in there fixed. Right. You've got to back off and say, okay, we need to start scheduling these things and then we need to figure out the best time to do these things. Right. If your code is running around the world, where's the biggest user base at? Ah. And when do we least impact the biggest user base? Or is there a way we can upgrade code that supports your Europe during their downtime, supports the US during its downtime, then supports Asia during their downtime, so it's least impactful to those users when you're upgrading the code. Right. There's a whole bunch of things like that. Right. But if you don't have effective change management, um, you're going to have more disruptions and have a lot harder time trying to figure out the way we made this change was a problem, the code we deployed was a problem, or we're not sure what happened, we're not sure what changed or why. If you start tracking it and trending it and keeping basically a change management system in place, and they're free ones out there, there's open source stuff out there that say, okay, this is the rollout strategy for this code. We're going to do a rolling Upgrade. So any server that has an IP address that the last digit starts with a 1 or 2M or a 3, we're going to upgrade those.
Speaker A: Or you can hash it.
Speaker B: Yeah, you hash it and then you, you plan that rollout strategy or you monitor as the stuff's being upgraded. Can you do an A B test where a small subset of users are assigned to the new code base and then watch it perform, watch it behave, see how users interact with it and see what their performance is like to make sure the upgrade worked well and then roll it out the rest of the way. Right. There's a bunch of different strategies you can use, but what I'm saying is if you don't have that, you will need it. Right. If you don't know what changed and something does break, it's going to be a lot more difficult to figure out why it broke.
Speaker A: Yeah. So do you think, I mean, just going back to this, uh, what you said about automation or the importance of automation, do you think it's possible to inv. I mean, do you think it's worth investing in automating change management where essentially getting closer and closer to this ideal of CD where you just push your code and it's either landed production and it's completely guaranteed to work, or there are so many checks along the way, they just guarantee that it's going to be caught and, uh, you know, sent back to you and reverted. What. What's your take?
Speaker B: I mean, is it worth it?
Speaker A: Or like, is it just. It should be some sort of a balance.
Speaker B: Well, the bigger you get, I mean, like, from what I understand in Facebook, you can probably write code, commit the code, it'll automatically deploy and you can keep an eye on it. And if it doesn't work, you can roll it back dynamically as well, which is really cool because they've automated a lot of stuff, but they're huge, they're massive. Right. When you look at smaller companies that don't have billions of dollars of revenue and, you know, an unlimited supply chain of engineers and operations guys, well, then you don't probably have that level of automation. Investing in it is probably good. As you get bigger, you probably need. There needs to be a, like a dance, a balance. Right. The bigger you get, the more automation you should have. Right. But if you're a smaller company, change management might just be, hey, we're deploying code. Jump on the zoom call or get on the slack channel or whatever as we deploy the code so we know what's going in, what changed and why Deployment's done. Yay. Let's go grab a beer. But as you get bigger, right, you start automating the entire supply chain and the change process. So there's documentation around what was written, what was changed, what the expected result's gonna be, and what to look for during the deployment to make sure the deployment went well, how to test it, how to do an end to end test and a smoke test to make sure that things work the way they're intended to work. And then once everybody's happy with that change, yay, once again, let's go get a bear.
Speaker A: Okay.
Speaker B: The smaller you are, it's not that it's not important, but it's probably not going to be as formal and automated until you get bigger.
Speaker A: Okay, so you're saying that if you are small, it's okay to sometimes like, you know, maybe not the whole thing is not fully automated and once deployment is done, sort of keep an eye on things for a bit, making sure that things are green. Right. Uh, and uh, because the, the impact is low, the damage to the customers is low. I mean the, the, what is it called? The blast radius is not very big because you become bigger and bigger. And uh, the, the blast radius begins to get measured in, you know, millions of users and you know, millions of dollars or even even tens of thousands. I don't know. What, what do you think? What do you think? They um, update the monetary blast radius when you can. You totally know that you need to start automating at what point? Thousand dollars?
Speaker C: Ten thousand? A million?
Speaker B: It's down to the company, man. I mean, like I'm at a small startup right now and the blast radius for us is really small because to us, $100,000 is an awful lot of money. And if you look at Facebook, $100,000 they don't, that's not even going to impact their P and L, no one's going to notice. Right. So their blast radius needs to be much bigger for it to become something critical. Um, as for us, we have a smaller blast radius that is very critical. So it's going to be unique per customer and unique per company. But I would say what I, what I tend to do is if there's a blast, if something bad happens and it's going to cost my company X. What I want that coming out of my pocket in my paycheck because I messed up,
Speaker C: that's not going to last long.
Speaker B: Exactly.
Speaker C: Right.
Speaker B: So if you use that mindset, right, it's, you can be ah, hard, but if you use that Mindset as just a stopgap to say, okay, you know what? This cost my company 50 bucks because I screwed up, I'll, um, be willing to pay that $50 into the slush fund for the next company party.
Speaker A: Right.
Speaker B: So I'm good with that.
Speaker A: What do you, I mean, I mean obviously automation means code. Right. And um, as, I mean if you have uh, ops and uh, engineering separate, I mean engineers can write code and ops just use their operational aptitudes, systems management aptitudes, an SRE environment. Everyone codes.
Speaker C: Right.
Speaker A: But I was just thinking that what is your take if, because some, or like the classical SRE culture was invented or rather documented by Google in their SRE book, what is your take on getting, uh, engineers, uh, participating in uh, wearing a pager, let's say 10% of their working time, so essentially being on call, even though you're like a UI coder, but getting them to be fully operational responsibility, of course, you know, being handheld by the real operations team. What do you think that sort of would bring the message of importance of automation, um, much stronger and faster?
Speaker B: Yeah. Well, it's so embedding people in other teams. Right. Like I talked about at the beginning, if, if the engineers are also responsible to help resolve problems, then their mindset starts to shift pretty quickly to making sure the problems don't pop up. And if they're going to help write the automation to make operations guys lives easier, it's also going to make their lives easier. So rotating people embedded into the ops team or operations guys embedded in the engineering teams, either way, you start pushing each other's methodologies back and forth and you start educating each other on why you did what you did the way you did what you did and what we need from you to make sure that we don't have to call you at 3 o' clock in the morning. Because that's just not fun. Right. I think is the more the engineers get involved and they don't have to be there all the time. Right. That's why 10% of your working week or working month, you have a pager or you get on the on call list, um, they will quickly figure out that, you know, what if we took an extra 15 minutes or an extra 2 hours per sprint or whatever the case may be and focus on what ways could this. Let's red team, blue team this. What ways could this break? What ways could I break? Because never say never.
Speaker A: Yeah.
Speaker B: Things will and can break.
Speaker A: Yes. And I agree. I mean, it doesn't matter. Uh, I don't know why? It's sort of low, this universe. Or maybe just entropy.
Speaker C: Things will go. Things. Things are going to.
Speaker A: Things are going to break. And uh, it's amazing, right? I mean, and what's, what's also amazing, making things work.
Speaker C: You like maybe two or three. Two or three ways to do that, right? But if you think of, you know, m. How many ways of things can break, it's like practically like an unlimited number.
Speaker A: Uh, so which sort of, kind of, sort of. I'm going to ask when things break, right? What is your guidance for engineering managers and ops? What is the best ways to manage incidents to minimize downtime? Um, before, during, after, what are your, uh, what is, what does your experience tell?
Speaker B: So what, what I've seen work the best in the past, and I've had to deal with this most of my career, is actually having a defined incident management process, right. So if you develop a complete, comprehensive, this is the plan that's going to involve everybody, like all the teams, Right. Unfortunately, QA usually doesn't get pulled into incident management. They test, but they're not the ones who wrote the code, not the ones who write, who run. They don't run it, they just test it. They didn't write it, they just test it. So QA is usually left out of incident management until the end for postmortems. But during an actual incident, you need to outline an actual process that defines roles. You should define who has what responsibilities during an incident, what kind of communication strategies you're going to use. Are there going to be phone calls? There's going to be text messages, pagers? Um, is going to be slack? Is it going to emails, Whatever. Whatever it is. And then making sure that people understand. There is an incident manager, someone who runs the incident, whose only job is to make sure the right communication goes out, the right actions are taken, and it's documented what happens during an incident. So you can do a post incident review, a root cause analysis meeting, right? To figure out lessons learned, right? This is what we saw happen, this is what we did and the time it took us to do it. And then this is the lessons that were learned so that we can incorporate that into the next cycle. So when you guys are building the next version of X, they understand what happened before and how to make sure it doesn't happen a second time. M. Right. Because the more engineers understand what the opsky people are going through, right. And the weirdness that they have to experience all the time, the more empathy they'll build and the more likely they are to Collaborate to say, hey, I'm proud of the code that I write. I wrote this. People are using this around the world and it's making their life better, or it's giving them entertainment, or it's giving them education or whatever it is about the code that they wrote. Nobody writes code to do nothing. Right. They're actually writing code for a business usually, or they're writing code for a function or a purpose. They expect people to use that code and they should be proud of the work they did. Right. And so if they're proud of the work they did and they have operations guys that are suffering because it's not working as well as it should, they should have enough empathy to work with the ops guys to come up and be collaborative on how to resolve these problems long term to make their code better, faster, stronger.
Speaker A: Yeah, yeah, 100%. And um, I agree, 100%. I think just a plan, a documented plan, can be in Microsoft Word, can be in to me personally, ideally somewhere in confluence, referenceable, um, step by step, who is the captain for the deployment or who's the running captain for the incident? Uh, key people, key communication channels, actions. Right. I think it's going to go a long way. I mean, if we just, if you just. If a team goes from not having a plan to having a plan, it's going to just cut down the downtime because instead of people, instead of running around with like chickens with heads cut off, they're going to be just doing what's known. Right. Already. Yeah.
Speaker B: Uh, something to point to, something you can look at. And look, a plan is just something to change once created. That's an old adage. Right. Once you have a process in place, you look at the process and you evaluate that process as well. Right. So the incident management process, you're doing a review of the incident. Right. You're doing a postmortem. Part of that. Postmortem should be, are there ways we can improve the process as well? Because there might be a way to streamline things to make it faster or make it less cumbersome or, or to include more people or exclude people that shouldn't have been there. Right. And what are your escalation policies? It should be defined in your management that you have an escalation. So if you have to call this engineer and they don't get out of bed, you call engineer number two and they don't answer and you call their boss. Then their boss gets out of bed and he calls everybody because.
Speaker A: Yeah, yeah, ah, yeah.
Speaker B: Like getting out of Bed, whatever the case may be. Right. But they're.
Speaker A: It usually explodes right from you know, no one to everyone.
Speaker B: Right.
Speaker C: With maybe with one or two people knowing what's actually going on.
Speaker B: Right. The rest of people are just being there, grumpy because whoever was supposed to answer didn't. But yeah, I mean I love it
Speaker C: like you know, middle of the incident, what is the status?
Speaker B: There's been many a time when I've been the incident leader, instant manager of an incident going on. And my entire job I had two phones. One was on the incident and one was talking to the executives, telling them what the current.
Speaker A: Exactly. And, and you know what, it's a good point because I think it's actually a trick, like an operational trick. Um, some people don't know where there's like a implicit agreement between the ops leader and the execs. Guys. You don't go into slack channel.
Speaker B: No, you don't go.
Speaker A: You go to me, right?
Speaker C: Because.
Speaker A: And it's sort of like there has to be. Doesn't have to be by the way, incident captain or an incident manager. Right. But it has to be dedicated communications
Speaker C: officer of the incident. Who is, who is responsible.
Speaker A: Um, uh, answering questions of the, of the leadership. Right. And, and they have rights, right? I mean what kind of a boss you are.
Speaker C: If the production is down, you know, tens of millions of people cannot use the system and they're just happily sleeping in your bed, right.
Speaker A: You, you will, you will engage, right?
Speaker B: Yeah.
Speaker A: And then get the. Is it, is the engagement going to be beneficial? Um, I, um, mean, except maybe yelling,
Speaker C: you know, work harder. Right. But that's. I think this is the uh, chief communication social.
Speaker B: That's one of the operations guys job is to deflect.
Speaker A: Yeah, exactly.
Speaker B: Deflect. Non essential communication from the incident that's occurring right now. People in the incident can get it resolved and you don't have to worry about it. Right. You have to worry about everybody interrupting you. Give them time to think, uh, time to execute, get it done, get it fixed in the minimum time possible. So the operations lead, whoever's the, the incident manager has to be able to bounce between incident team that's working on resolving it and then giving updates to everybody else around. This is the current status, this is what we've learned. The next update's going to be in 5, 15, 30 minutes, whatever the case may be. Then they can go back to the incident team and be like, all right, what are you guys doing? How can I help?
Speaker A: Yeah, yeah, you know what? I think this whole sort of like uh, having those two layers of communications is critical. Again, the worst thing that can happen to the team is a destruction is a big boss yelling work harder. And what I think giving. Going back to your idea. Not ideas like a. I think it's a, like a very, very established common sense. But best common sense, best practice or having an incident, um, management plan outlining two communication channels. Right. This is the communication channel. High intensity, higher rate, super technical. And the second channel is for you know, regular, steady, high value, low technicality management channel. Right. I think uh, no matter what you use text, slack, whatever teams, I think it's beneficial because early in my career I've done mistake of.
Speaker C: Hey, it's an open communication channel. No it's not.
Speaker B: Yeah.
Speaker C: Take the microphone.
Speaker B: Yeah, well that you. When that does happen, then you'll have you know, people who are non technical jumping in and asking questions, which is not bad. We want to educate people. Just not the right time or place,
Speaker A: not the right time. And you know what? I think another, another uh, I'll call it as it is. I think another danger of uh, higher level of remote leaders, like remote from the problem leaders being exposed to the low level, messy, fast moving, complex environment of uh, incident resolution. I've seen it too many times. They just make a uh, they come to conclusions. Their job is to make, to come to conclusions fast. Right. And they're, that's why they're on those levels. They're like, this is a freaking mess. These are the people who have no
Speaker C: idea what they're doing. Right. This is really. I really have to m. Make sure that everyone, all those idiots, um, taken care of.
Speaker A: I think this is the even worse situation. It's sort of like you don't want a patient to wake up on surgeon's
Speaker C: uh, table in the middle of the surgery.
Speaker B: Right?
Speaker C: Yeah.
Speaker B: Waking up during surgery is never a good time.
Speaker C: Yeah. Then the patient is going to be, hey, doctor, what are you doing? Are you sure that's the right thing you're pulling? Yeah. Uh, not a good idea.
Speaker A: Yeah, good stuff. So what else? Um, so, um, metrics and performance. You mentioned that you wanted to cover it as well. What is your take on that?
Speaker B: We've talked about that lightly a couple of times, but it's something we need to focus on as you deploy code. As code is running different people in the executive team and in the leadership teams in engineering, QA operations, all the different. They want to see different metrics. You need to be able to establish common things that you will measure performance indicators like how often you deploy, how often you change, um, how much time it took to deploy or what the meantime to recovery was. The mean time between failure change, failure rate, all these kinds of things help people understand what they built and how it's running. Right? And building a cool dashboard with high level stuff. For the executive team, huge benefit because they can look at a dashboard and point at a metric and be like, well, we're doing X number of transactions per hour right now. And I feel good about that because. Or there's a performance degradation. What's going on with this? How can we solve this problem? Right? That may get you the money you need to fix that one piece of code. You have to fix or upgrade that network. You need to upgrade or change the methodologies you're using. There's a lot of different stuff around metrics and performance, right? If you don't have measurements on what you're running, how do you know what it's doing? You've got to be able to measure it. You've got to be able to see what it's doing and trim. That I found over time, right? With trending performance, I can look at how things performed this time last week and look at production today. And if they're not quite the same, uh, that's close. Okay, cool. If it's more than 5 or 10% different, did we have a large uptake in customers? Was there a decrease in performance? What's the variable that's making my metrics change? Because I can look at the past to help predict what today should be. And if today's too much of a difference than the past, there might be a problem that we haven't seen yet that's starting to bubble up. So the more you monitor and the more definition you have around what this key performance indicator is, what this metric should be, then you'll be able to start getting ahead of problems. So if you deploy a change and you trim that for a couple of days and you go back to the engineering team, hey guys, look. Before we deployed this change, it took us half a second to perform this transaction. Now it's taking us a second and a half to perform this transaction. It's not noticeable to the customer too much. They're not really saying anything about it, but I can see it. So we may want to take a look at the code we changed last time to see if there's ways to optimize this before we push again to fix the problem before it becomes a real issue. Right? Things like that. The more monitoring you get the better the performance caters are. You're going to have a better understanding across the board of what's available, what the performance is and where we can improve. And if everybody's in the same page of music, you get a much better symphony.
Speaker A: So let me ask you this, who defines, uh, what is a working application in production? Is it ops, is it engineering, is it qa, is it product?
Speaker C: Uh,
Speaker A: whose plate is it on?
Speaker B: So I think it's a joint effort, right? So operations guys are going to look at different metrics. Different metrics. They should all have the availability to see all the metrics. But people will look at different things based on their specialty. Like marketing guys are going to look at site performance, total number of views, they're going to look at, you know, how many people have seen what we do and how well our message is getting out in the world. So they want to look at some things, operations guys want to look at the low level stuff. They want to see how many transactions are occurring, what kind of memory is being used, how much disk is being used, what the network speed is, how many retrans there are. Engineers are going to be more interested in how their code is working, right? What their heap size is, what memory utilization there is, how long it took a transaction to execute, right. What kind of libraries are being called, when and how long it takes those libraries to either compile, be brought in or be used. So there's going to be different views to the same system based on the role of the person looking at the system. Right. Um, the people who define that are usually the people that are closest to it. Um, but you need to, the ops guy may say I want to see these five things and the engineers may say I want to see these other five things and the executive team wants to see these three things or these four things. And marketing wants to see something and product wants to see something completely different but cool once it's defined, right. And you can come to an understanding of what you need, why it's important, and you know what, we should bubble up, um, the higher the level you are, the less detail there usually is, right? So the lower on the food chain you get to the more specific detail on a specific metric you're looking for. But the guys at the top of the food chain really aren't really interested in the number of milliseconds it took for this packet to go between point to point, right. They want to see transactions, not packet layer. It's going to be based on role, responsibility. But no matter who wants the metric metrics are never bad. Um, it can become expensive if you trim these metrics and then log everything. It goes back circulars back to that. We're keeping log files of everything and now all of a sudden our costs explode. There's some cases to keep metrics that you can archive for a long time so you can have past performance to try to predict future behavior. But you don't want to keep all of the detailed metrics for everything forever. That just becomes too much.
Speaker C: Uh,
Speaker A: okay, so basically you're saying that we should um, work together across the orgs to identify maybe top four or five things which are really important to them and maybe implement those first, start tracking them, start aggregating them maybe for the historical reasons. And um, once that nailed, um, uh, see if anything else is needed.
Speaker B: Absolutely.
Speaker A: M. Good stuff.
Speaker B: I know a lot of good ops guys that want to monitor and measure everything and they usually do because they're geeks that way. But uh, at least start with stuff that people perceive value with. Right. The executive is going to see a different value proposition than the engineers will or the operations guys will or the product guys will. But whatever they see as value, trend it, monitor it, get dashboards up for it.
Speaker A: You know, don't you think we should first listen to the people with the
Speaker C: money or people who make money and you know, our paychecks?
Speaker A: Yeah, the customers. Yeah.
Speaker B: You keep them happy, you keep getting a paycheck and that's a good thing.
Speaker A: Yeah. I think my new knee jerk reaction has always been, okay, who's the person
Speaker C: who's writing the check, who's writing the paycheck? And ask them, hey, what do you want?
Speaker A: Uh, but I think business, uh, if you let ours, I mean if you let us gigs to, you know, define the metrics, they're going to be a lot of them, they're going to be a lot of fun and um, um, but I think the, in the end of the day it's about the value creating for the customer and they got to be the metrics that show that. Right. Um, um, what those metrics are, I think it's really. But I think it's important for engineering managers to be aware that they need to think of it and they need to go and ask those questions. Right. What is it a meaningful behavior of the system to you? Uh, I don't know. Whom would you go? But at least cfo, uh, uh, maybe chief, uh, product Officer, if there is one, or um, a business owner. Right. I've been surprised before. Right. When you go and ask. Hey, if you ask a business owner, uh, or the big boss what is important to you, you may be surprised. And I'm saying that you have to go to the big boss. Because if the big boss is not happy, how's the company going to be happy?
Speaker C: Are you going to be happy?
Speaker A: Right. Uh, so I've seen, um, senior, uh, CTO asking for, um, uploads, um, per second. Like for example, Shutterflyer. Uh, we had a great cto Satish. Um, um, I won't think of it, but then when we started talking, it turns out if users cannot upload or customers cannot upload a photo, right, you cannot make a product, period.
Speaker C: It's not even transactions, it's uploads.
Speaker A: And the more you upload, the more products you want to make. Essentially, if you don't go and ask, you won't know. And, uh, somewhere in the middle, no matter Ops or engineers, it's very hard to invent a metric without asking, uh, or it's going to be a geeky metric, which is important.
Speaker B: Well, some metrics that are geeky are important, but usually the metrics that the business wants to see eventually tie back to money in some way.
Speaker A: Yeah, exactly. There are proxy metrics. For example, uploads can be proxied by, uh, Iops, um, in the storage. If iops goes from, um, 15 million
Speaker C: per second to 5, that tells you something.
Speaker A: Yeah. Well, uh, Jeremy, uh, it's been an awesome conversation as always. And, um, so traditionally I'm gonna ask, uh, ask you, uh, for a checklist for our listeners, for the engineering managers who are listening to us. Can, uh, you give a checklist maybe three to five points? Um, what are the things that engineering managers can start doing tomorrow to live, uh, happy, productive lives together with operations teams? Sure.
Speaker B: Okay. Things to take away from this that you can keep top of mind. Right. The first thing is people, right? Everybody you work with is a person. Take the time to know the person. Go out, grab coffee, get a drink, raise a glass, eat some food, whatever. Learn not to what they do, that's important. But learn who they are, learn about them, learn about their kids. You're going to be much more successful building bridges with other teams. If you know the people more than just the role, that's important. That's number one. Number two, don't forget metrics. We just finished on that. It's a great point. The more people see what you do, the more value is created in their mind about your role. So if you want people to understand what you do, let them see it. Um, embed your teams, get people cross communicating, cross trained. It's really important. The more people understand each other, they understand what everybody else does. The more productive you're going to be and the less failure is going to occur. Um, constant learning, always learn. You're never educated enough. I do it on a constant basis. I'm always trying to keep up to speed on what's occurring in the world today, what technology is coming out, what I should be on top of. Because I'm sure that the day that I stop doing that is the day that someone's going to ask me to do X and I don't know X. Um, security is not something to forget about. Keeping the back of your mind all the time. Always understand that there are different criticalities to security incidents. Respond to the ones that are high criticality urgently and try to prioritize other stuff as you can. Finally have fun. Work can be stressful. Believe me, I have been in pressure cooker situations. At the end of the day, we're all people. Just take the time, have fun and just enjoy, uh, the ride.
Speaker A: Well, Jeremy, uh, thank you, uh, much appreciate it. It was super, uh, insightful. Um, I'm not surprised. And um, I would like to encourage our listeners to share this episode if you find it useful, uh, or insightful. And we're always open, uh, to feedback and uh, um, it informs our future episodes and we always want to hear from you. And you can reach out to us@, uh, effectiveam.com on the web, contact, uh, effectiveam.com or the email or you can find us on LinkedIn. And uh, that was good stuff with Jeremy Franzen, the expert in Ops. And um, more good stuff is coming. Thank you very much. Thank you, Jeremy.
Speaker B: Thank you, Slava. I'll talk to you later.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.