
Code with Jason · 2026-07-01 · 1h 10m
Key moments - from our scoring
Substance score
60 / 100
Five dimensions, 20 points each
David Yanacek brings two decades of AWS experience to a conversation about the resurgence of foundational software development practices - test-driven development, spec-driven development, and specification-first thinking - particularly in the context of agentic AI coding tools. Rather than inventing new methodologies, teams are discovering that practices long on the backlog become critical when scaling autonomous agents. Yanacek explains how vanilla AI coding assistants drift into unwanted refactorings, comment out tests, miss requirements, and hit context window limits. AWS built Kiro, a spec-driven IDE, to address these failure modes by having agents work from a living requirements document that doubles as a design artifact and task checklist. The conversation reveals a principle-based approach to AI: viewing every mistake the agent makes as a defect to be prevented through better specs, steering rules, and team guidelines - essentially treating agent feedback loops like code review that gets baked into system behavior. Jason and Yanacek extract the deeper principle: physical barriers (pre-commit hooks, specs, tests) work better than behavioral instruction alone.
Vanilla agentic coding tools often wander off into unwanted refactorings, comment out tests to make them pass, hit context window limits that cause them to lose track of work already done, and fail to complete full assignments because they lack explicit guidance through specs and steering rules.
Kiro is AWS's agentic IDE built around spec-driven development, where a spec serves as a living requirements document, design artifact, and task checklist that keeps the agent aligned and recoverable even if context windows overflow or the agent needs to resume work mid-project.
Teams should treat every unwanted behavior as a defect, then add steering rules, update project skills, or refine specs so the agent remembers not to repeat that mistake - similar to addressing code review feedback but baked into system behavior.
Physical barriers like pre-commit hooks, specs, and test gates work better than behavioral prompts (like 'act as a senior developer') because they make mistakes impossible rather than just less likely.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains real technical substance - spec-driven development for AI agents, property-based testing, TLA+ for distributed-systems verification, COE process mechanics, and the DynamoDB provisioned-throughput evolution - but this is badly diluted by a multi-minute snail-mail ad, extended Michigan/weather small talk, a power-tools brand tangent, and a lengthy oil/calculus analogy digression that consumes well over 15 minutes of a 70-minute runtime.
we found that they would wander off. Maybe uh you'd come back and look at what it was doing, and it was doing some refactor that you didn't really want to do right now or ever
I think where people spend like the the the most productive time in in getting an agent to run um just better, is when they when they view everything that the agent did that they didn't want to do as a defect
Spec-driven development for AI agents is increasingly common discourse, but the episode yields genuinely insider-original material: the internal history of Coral becoming Smithy/SDK generation, the AWS-wide weekly ops meeting where any team's dashboard can be pulled and interrogated live, and the feedback loop of mining COEs to improve the underlying service (e.g., adding DynamoDB auto-scaling after seeing throttling incidents in postmortems). These are non-recycled takes grounded in lived experience rather than thought-leadership clichés.
There's no compression algorithm for experience.
we started DynamoDB with this notion of provisioned throughput because people wanted super predictable performance. So we would pre-provision capacity... but um, we didn't have auto-scaling at first. And so when we saw people run into uh, oh yeah, like it's it's hard to actually monitor this all the time and then dial it up. So let's just build that into the product
David Yanacek is a genuine scale practitioner: he was on the pre-EC2 Amazon.com web server fleet, contributed to what became CloudWatch, worked on the Coral/Smithy service framework, DynamoDB, Lambda, and API Gateway, and is now building Kiro (agentic IDE) and the AWS DevOps Agent. This is rare hands-on depth across foundational cloud infrastructure. The score is not a perfect 20 because the conversation format prevents him from going as deep as his experience warrants.
I've worked on Lambda then for managing servers, the building serverless, so you don't have to manage servers, API gateway because API routing is tedious and and easy to get wrong
we do a lot of formal verification methods, like TLA plus... we have things like DynamoDB, things like S3, where we have to prove that the replication, like these distributed systems around failover during a network partition while bootstrapping another replica
The episode offers solid named specifics - TLA+ invariant verification for DynamoDB/S3 replication, Smithy as the open-source service definition language, P100 CPU percentile as the correct alarm metric, the five-whys Toyota method embedded in COEs, and the Coral framework becoming SDK generation - but lacks quantitative evidence: no reliability figures, no scale numbers, no timelines for product launches, and no data on agent productivity gains.
I had an alarm, but I've alarmed on the median host CPU instead of really, it should always be the well, you should have that, but you should also have the the max, the P100, hundredth percentile host CPU to see like because maybe you're because it's really a distribution across your fleet
It's called Smithy, it's actually open source now, um, that defines services and what shapes and APIs and signing methods for identity
The host has genuine technical knowledge and lands a few good moments - pushing back on postmortem efficiency and asking sharp questions about property-based testing - but frequently commandeers the conversation with lengthy personal analogies (oil/kerosene history, calculus derivatives, woodworking jigs, power-tool brands) that crowd out guest depth. There is no meaningful challenge to guest claims and the talk-time ratio heavily favors the host, limiting how far the practitioner's expertise is actually excavated.
I don't know, David. This all sounds like it takes a lot of time. That doesn't sound very efficient.
I I I have to take us on a small digression. Um, my friend uh Stephen Baker mentioned the other day, uh he he made an analogy with oil. So uh back in the I don't know, mid-late 1800s, second half of the 1800s, um oil wasn't used for gasoline
Computed from the transcript - who did the talking, and the words that came up most.
In this episode I talk with David Yanacek about his journey from operating Amazon's web server fleet to revolutionizing DevOps at AWS. We discuss AI in software development, spec-driven development, and universal testing techniques. David shares insights into operational excellence and the Amazon Builders Library. Links: - Amazon Builders Library - Nonsense Monthly
Transcribed and scored by The B2B Podcast Index.
1 - > SPEAKER_01: Hey, it's Jason, host of the Code with Jason 2 - > podcast. 3 - > You're a developer. 4 - > You like to listen to podcasts. 5 - > You're listening to one right now.
6 - > Maybe you like to read blogs and subscribe to email newsletters 7 - > and stuff like that. 8 - > Keep in touch. 9 - > Email newsletters are a really nice way to keep on top of 10 - > what's going on in the programming world. 11 - > Except they're actually not.
12 - > I don't know about you, but the last thing that I want to do 13 - > after a long day of staring at the screen is sit there and 14 - > stare at the screen some more. 15 - > That's why I started a different kind of newsletter. 16 - > It's a snail mail programming newsletter. 17 - > That's right.
18 - > I send an actual envelope in the mail containing a paper 19 - > newsletter that you can hold in your hands. 20 - > You can read it on your living room couch, at your kitchen 21 - > table, in your bed, or in someone else's bed. 22 - > And when they say, What are you doing in my bed? 23 - > You can say, I'm reading Jason's newsletter.
24 - > What does it look like? 25 - > You might wonder what you might find in this snail mail 26 - > programming newsletter. 27 - > You can read about all kinds of programming topics like 28 - > object-oriented programming, testing, DevOps, AI. 29 - > Most of it's pretty technology agnostic.
30 - > You can also read about other non-programming topics like 31 - > philosophy, evolutionary theory, business, marketing, economics, 32 - > psychology, music, cooking, history, geology, language, 33 - > culture, robotics, and farming. 34 - > The name of the newsletter is Nonsense Monthly. 35 - > Here's what some of my readers are saying about it. 36 - > Helmut Kobler from Los Angeles says, thanks much for sending 37 - > the newsletter.
38 - > I got it about a week ago and read it on my sofa. 39 - > It was a totally different experience than reading it on my 40 - > computer or iPad. 41 - > It felt more relaxed, more meaningful, something special 42 - > and out of the ordinary. 43 - > I'm sure that's what you were going for, so just wanted to let 44 - > you know that you succeeded.
45 - > Looking forward to more. 46 - > Drew Bragg from Philadelphia says, Nonsense Monthly is the 47 - > only newsletter I deliberately set aside time to read. 48 - > I read a lot of great newsletters, but there's just 49 - > something about receiving a piece of mail, physically 50 - > opening it, and sitting down to read it on paper that is just so 51 - > awesome. 52 - > Feels like a lost luxury.
53 - > Chris Sonnier from Dickinson, Texas says, just finished 54 - > reading my first nonsense monthly snail mail newsletter 55 - > and truly enjoyed it. 56 - > Something about holding a physical piece of paper that 57 - > just feels good. 58 - > Thank you for this. 59 - > Can't wait for the next one.
60 - > Dear listener, if you would like to get letters in the mail from 61 - > yours truly every month, you can go sign up at nonsense monthly 62 - > dot com. 63 - > That's nonsensemonthly dot com. 64 - > I'll say it one more time nonsense monthly dot com. 65 - > And now without further ado, here is today's episode.
66 - > SPEAKER_00: Hey, thanks for having me. 67 - > Very excited to be here. 68 - > SPEAKER_01: Excited to have you here. 69 - > So you've been at AWS for about 20 years, if I have that right.
70 - > SPEAKER_00: Uh yeah, that's right. 71 - > Uh I guess the split hairs a little bit. 72 - > I guess Amazon for 20 years, um, like the vast majority of that 73 - > has been AWS, though. 74 - > SPEAKER_01: Got it.
75 - > Okay. 76 - > Um and and we were talking a little bit pre-show, and you and 77 - > I have some geography in common. 78 - > You you've spent some time in Michigan. 79 - > SPEAKER_00: That's right.
80 - > I guess about half my life here in Seattle, and then about half 81 - > my life uh within a you know hour and a half or so of where 82 - > you are in a couple locations. 83 - > But yeah, uh, in fact, just outside in Seattle, right before 84 - > this, I was very excited uh that I heard some uh thunder, which 85 - > is a very rare thing here, but I know in the in the Midwest, you 86 - > know, you can tell the people who I work with who grew up kind 87 - > of around thunder, like you get in the in the Midwest.
88 - > Uh because it's just it's you get the nostalgia here in 89 - > Seattle, it happens only like once uh oh, there it is again. 90 - > Uh yeah, it only happens uh once a year, maybe. 91 - > Maybe and so when I hear it, I get this nostalgia and this 92 - > feeling of like nap time growing up and hearing the rain be so 93 - > calming. 94 - > Uh so there's other core from like from El Salvador, also kind 95 - > of you can tell who runs to the window because they they had had 96 - > that kind of experience growing up.
97 - > SPEAKER_01: Interesting. 98 - > Um, it's it's always funny the things you miss. 99 - > I I moved down to Austin, Texas. 100 - > And by the way, I don't know if we said uh I live in uh West 101 - > Michigan, just outside of Grand Rapids.
102 - > Um I moved down to Austin, Texas, partly to get away from 103 - > the cold of Michigan, and I found that I missed the cold and 104 - > like the overcast skies and stuff like that. 105 - > I'm like, this kind of sucks. 106 - > It's like beautiful every single day. 107 - > I want it to be like crappy out sometimes.
108 - > SPEAKER_00: You can decide to have all four seasons. 109 - > SPEAKER_01: Yeah, yeah. 110 - > And uh last question for you on on this topic: where exactly in 111 - > Michigan did you live? 112 - > Like what town?
113 - > SPEAKER_00: Oh, I grew up in Midland, Michigan. 114 - > Uh and then I went to school in Ann Arbor. 115 - > So U of M. 116 - > Yeah, U of M, University of Michigan, yeah.
117 - > Um but yeah, Midland, Michigan, yeah, it was a kind of a it's a 118 - > large suburb. 119 - > Uh I guess it calls itself a city, about 40,000 people or so. 120 - > SPEAKER_02: Mm-hmm. 121 - > SPEAKER_01: Yeah, I believe I've been through Midland.
122 - > Um okay, so getting into the meat of it, I listened to a 123 - > little bit of an episode that you recorded for a different 124 - > podcast. 125 - > Um, and you talked about AI and testing and spec driven 126 - > development and stuff like that. 127 - > Um this is right up my alley. 128 - > Um and and I think it's kind of funny um because it's like the 129 - > industry is rediscovering these things that a lot of people were 130 - > doing all along, um, like test-driven development.
131 - > And and it's so funny to watch because it's like, hey, if you 132 - > just like decide what to do before you do it, and then write 133 - > tests that represent what you intend to do, and then write 134 - > code to fulfill those tests, everything turns out to work a 135 - > lot better. 136 - > And it's like, yeah, like uh we we knew that, but it's it's good 137 - > that people are rediscovering this. 138 - > Anyway, I'm I'm curious to hear your take on this. 139 - > SPEAKER_00: Yeah, I mean, I think uh I think that's that's 140 - > actually quite a bit true around the transformation that teams 141 - > I'm seeing software teams go through to be so to get so much 142 - > more done uh than before.
143 - > It's actually doing things that like we not necessarily new 144 - > things, like just doing things that maybe we've been meaning to 145 - > do and haven't gotten around to it yet. 146 - > Um things where you can just increase the autonomy and and 147 - > let the agentic coding loop just go and be super productive where 148 - > um yeah, where before it was kind of on the backlog. 149 - > Oh, yeah, that would make us a little bit faster if we had 150 - > really good, you know, bet a little bit better unit tests 151 - > that kind of would fail sooner if there was a problem versus an 152 - > integration test, which is free further along in the pipeline or 153 - > that kind of thing.
154 - > So these things that teams like would always like to be able to 155 - > do um just become that much more useful because you just get that 156 - > much more done if you can let an agent loose with it. 157 - > And and so spec-driven development, just as a starting 158 - > point for all these other practices, if you want to talk 159 - > about more beyond the spec, but um, you know, we it's something 160 - > where we found um internally that when we were using vibe 161 - > coding tools, agentic coding tools, um, we just found that 162 - > they were, we could see the potential.
163 - > They were so powerful and in how much it was so impressive that 164 - > they could generate so much. 165 - > Um, but we found that they would wander off. 166 - > Maybe uh you'd come back and look at what it was doing, and 167 - > it was doing some refactor that you didn't really want to do 168 - > right now or ever. 169 - > Um it would uh you know cheat on the test, it would kind of maybe 170 - > comment out uh tests or you know, say, oh yeah, that seems 171 - > about right, and and just not really do what we agreed on, or 172 - > do, or maybe not even complete the total assignment.
173 - > Um, and so we found that what we could do is is introduce specs, 174 - > like you're talking about, the spec-driven development. 175 - > And and we actually made a whole uh agentic IDE, a coding 176 - > environment called Kiro, um, around spec-driven development 177 - > to be able to uh to just do this, to get to be more useful 178 - > for to have a more useful workflow for the agent so that 179 - > it could do production grade coding, where it would actually 180 - > get everything done.
181 - > Um with all the resiliency and everything built in. 182 - > SPEAKER_01: Yeah, I want to talk a little bit more about the ways 183 - > in which vanilla AI with no special guardrails will fail to 184 - > do what you might want. 185 - > Um, because there are a lot of different uh software 186 - > development methodologies. 187 - > Some are better than others, and the AI, you know, it it like 188 - > knows everything, but it doesn't know how you want it to behave.
189 - > It's like, what, you want me to behave smart or dumb? 190 - > Like I'll I'll be dumb if that's what you want. 191 - > Um and and out of the box, it's not gonna follow this like I 192 - > don't know, Kent Beck style of development that I personally 193 - > might want. 194 - > But it can, it's it's not like it doesn't know.
195 - > It's just that it's not gonna do that by default, and then it's 196 - > like it if if it does something that I think is stupid, and I'm 197 - > like, oh, like don't do that. 198 - > Use like the dependency inversion principle in instead 199 - > of all these conditionals, and it's like, oh yeah, of course, 200 - > I'll do that. 201 - > But it's funny because it it won't do that on its own. 202 - > SPEAKER_00: Right.
203 - > I think where people spend like the the the most productive time 204 - > in in getting an agent to run um just better, is when they when 205 - > they view everything that the agent did that they didn't want 206 - > to do as a defect. 207 - > Like as a defect that you can do something about. 208 - > Like imagine like uh uh it's like uh if if you do like a code 209 - > review, like you you actually have all the code ready and it's 210 - > time to look at that code and see like, okay, if I have 211 - > feedback on that code review, like obviously, okay, let's 212 - > address that, but like let's figure out how to remember that 213 - > for next time.
214 - > And so if you kind of embrace this the this workflow of 215 - > viewing every everything is okay, how do we make sure as I 216 - > address this code review feedback that I add some 217 - > steering, some skill, update my skills, update the steering with 218 - > the project, or maybe so that the team uses so that it 219 - > remembers next time when it's either generating a new spec and 220 - > following it that way, or whether it's just going off and 221 - > doing something that it would remember, remember how to 222 - > address that thing that went wrong last time.
223 - > Because you asked what uh what can go wrong. 224 - > I mean, it's it's uh I guess I yeah, I just see it wander off 225 - > and start and start uh you if it if it r actually, okay, so one 226 - > class of things comes to when it um either hits a uh like a 227 - > sometimes it'll hit a context window limit, like it'll it'll 228 - > it'll just have been running for a long time and it kind of 229 - > starts losing track of what it had done and what it had already 230 - > learned in that session.
231 - > And so that's why having having this like a spec that lays out 232 - > this larger project that we want to do that might take a couple 233 - > hours. 234 - > It might take more than a couple hours. 235 - > And that could potentially overflow context window, need to 236 - > reset it, need to compress it. 237 - > Um, and so having that all the whole plan written down helps it 238 - > pick back up.
239 - > If you say, okay, actually, let's just we've we've achieved 240 - > a few tasks so far. 241 - > Um, like we've done steps one through five. 242 - > This is part of the spec. 243 - > The spec is essentially a uh a requirements doc that you work 244 - > with the agent to develop.
245 - > It's a design that you work with the agent to refine, and then 246 - > it's a task list. 247 - > And so these three things that make up the spec end up being a 248 - > really useful thing to fall to deal with the fact that the 249 - > agent, yeah, it can run out of essentially memory, context 250 - > window memory. 251 - > And if it does, it can just pick up which whether after it can 252 - > look at what it tasks it was going to do and had already done 253 - > and re-just kind of re-get its bearings and start over.
254 - > So that's one thing, like the context window overflow. 255 - > SPEAKER_01: Um yeah, and if I if I may interrupt, um whenever I 256 - > learn something new in programming, um I try I try to 257 - > extract principles out of it. 258 - > Um, because we're we're always presented with new tools and 259 - > stuff like that. 260 - > Um, but it if you can extract the principles out of it, the 261 - > principles are much more portable and durable over time 262 - > and stuff like that.
263 - > And and it's there's a lot of profit there if you can extract 264 - > the principles. 265 - > I've been trying to do that with AI. 266 - > Um, and I'm I I'm I'm kind of noticing multiple layers to the 267 - > way the ways that people are harnessing AI. 268 - > Um one is to just like tell it what to do, say like behave in 269 - > this certain way.
270 - > And that can work. 271 - > It it seems that as the models improve, they're getting better 272 - > at that. 273 - > That's like my anecdotal perception on that. 274 - > SPEAKER_00: Act as like you are a professional senior developer, 275 - > like uh that kind of thing.
276 - > SPEAKER_01: Yeah, yeah, just like in general, uh, don't mix 277 - > refactorings with uh behavior changes, that stuff like that. 278 - > Um it it seems to be getting better at that, but it's that is 279 - > like a um it's like a what's how do I want to put it? 280 - > It's not a very tight harness, you know? 281 - > It it won't necessarily follow that.
282 - > So you can do things like uh I now add pre-commit hooks that 283 - > run the whole test suite before any commit can be committed. 284 - > And that's like a uh physical barrier. 285 - > SPEAKER_02: Yeah. 286 - > SPEAKER_01: And I I I I think those things, whenever you can 287 - > put in like a physical barrier, I think of it kind of like a 288 - > woodworking jig or something like that, like a a jig that 289 - > makes it impossible to make a mistake, like maybe a hole 290 - > drilling jig where like you can't not drill the holes in the 291 - > right place because you have this physical barrier there.
292 - > So that's another thing. 293 - > And then the the kind of thing that you're describing, um, 294 - > where you just save these files, and it's like, okay, I have this 295 - > spec, which is kind of the authoritative source of truth 296 - > for the project I'm working on, and that's durable, and you can 297 - > always refer back to that. 298 - > Do do you do you think about it in terms of like principles that 299 - > you can apply to to harness the agent? 300 - > How do you think about this?
301 - > SPEAKER_00: Yeah, I think though uh the the types of barriers 302 - > that you can uh it's sort of actually partly about the agent 303 - > harness that you build of like that it that it will the the 304 - > program surrounding the agent won't let it succeed and 305 - > proceed, excuse me, proceed until it uh until it uh does the 306 - > thing and proves the proves the previous step. 307 - > So it just enough workflow kind of along with the letting the uh 308 - > LLM uh be the sort of Ouija board, uh if you will, that's 309 - > like just kind of ghost that's in the machine.
310 - > Um and so that one of those things uh by by breaking down 311 - > the project into tasks, um those tasks would include uh actually 312 - > some pretty interesting testing, uh testing barriers that you're 313 - > talking about. 314 - > Um one of them is uh a technique, again, another 315 - > technique, like you say, that's been around a long time is is 316 - > called property-based testing. 317 - > It's just a type of testing, it's sort of a mindset. 318 - > Um it there are frameworks out there and have been for some 319 - > time around this testing technique.
320 - > And this is one where, because you have a spec with really 321 - > well-written um directions with like caps, like you shall, the 322 - > program shall do this when in caps will this happens. 323 - > And from those, it turns out those are pretty useful inputs 324 - > to generate property-based testing. 325 - > Um, property-based tests for just for the for everyone are 326 - > things that um that test exhaustive um properties of of 327 - > the program versus just boundary cases.
328 - > And so you could let's take, let's say you're you're 329 - > implementing a uh a traffic light uh system. 330 - > You of one really important property invariant of it is that 331 - > at most one direction has a green light at a time. 332 - > Um that has to be held, that's its only job, really. 333 - > And it's everything else is just goodness.
334 - > Um and so you can just that would be described in the spec 335 - > as a requirement. 336 - > Um, and then a property-based test would uh generate input, 337 - > they generate inputs like with a just a bunch of combination of 338 - > the inputs to the program as sort of like a test driver. 339 - > And they're testing along the way in every step of this that 340 - > the program maintains those invariants. 341 - > It's sort of just it's just a test framework, right?
342 - > Uh but a good technique that is particularly useful when it 343 - > comes to these requirements documents, that it can just be 344 - > turned into all of these verifications. 345 - > And so the coding agent won't continue until it has written 346 - > the property-based tests and that they pass. 347 - > And so it just it can't proceed until it has done that. 348 - > SPEAKER_01: I've never done property-based testing and I 349 - > never I heard the name and that's it until I listened to 350 - > that other podcast of yours where you where you talked about 351 - > it.
352 - > Um it the way it works, will the tool typically like generate 353 - > some test cases? 354 - > Like, if there are 17 different permutations, then it'll 355 - > generate 17 tests and you do it that way, as opposed to like 356 - > there is a chunk of of the test that like loops through and and 357 - > does that kind of stuff, or maybe both. 358 - > How does that work? 359 - > SPEAKER_00: Yeah, the input generate, there are a bunch of 360 - > different input generator strategies that they that the 361 - > frameworks supply.
362 - > Some of them are include some amount of your randomness, um, 363 - > looking at certainly making sure that the important combinations 364 - > are all exercised. 365 - > But yeah, input generator that just comes up with all the 366 - > different there are different input generator strategies in 367 - > the in the frameworks. 368 - > Uh yeah, it's uh it tries to exhaust all possible uh inputs, 369 - > uh as many as many as are practical. 370 - > SPEAKER_01: Um yeah, yeah, that's that's really great.
371 - > Um, and that's you know, that seems to go along with this idea 372 - > of like a jig that that makes it physically impossible to do the 373 - > wrong thing. 374 - > Um I have found sadly that LLMs tend to write fairly fairly poor 375 - > quality tests, at least at this present moment. 376 - > Maybe they'll get better. 377 - > I think they have gotten better uh in the last couple years or 378 - > whatever.
379 - > Um, but at the moment, not great. 380 - > I found that I if I give it a fairly thorough uh suite of 381 - > examples, if I drop it into a program that already has a lot 382 - > of well-written tests, it's good at extending that and writing 383 - > more tests that resemble the existing ones, uh tests that 384 - > I'll be pretty happy with. 385 - > I found that when I do a greenfield project with Claude 386 - > Code, for example, the tests are not always great. 387 - > And it seems to depend somewhat on the language.
388 - > There's kind of a testing culture base baked into each uh 389 - > community, and and some communities have better testing 390 - > cultures than others. 391 - > Anyway, I was writing a program in Rust. 392 - > I I don't know Rust at all. 393 - > I was just letting Claude Code write the Rust, and I was really 394 - > unhappy with the tests that it was writing.
395 - > They were very just like perfunctory, um tautological. 396 - > Just if I put in five, assert that it's set to five. 397 - > It's like, well, that's that's kind of pointless. 398 - > So I hooked it up to mutation testing, and so I said, write 399 - > the tests and then do the mutation testing so that um it's 400 - > it's a much more rigorous, thorough way of testing it.
401 - > And I found that to to work out a lot better. 402 - > It kind of reminds me of of this property testing idea. 403 - > Have you also applied mutation testing to this this at all 404 - > also? 405 - > SPEAKER_00: I'm actually not familiar with the term mutation 406 - > testing.
407 - > SPEAKER_01: Oh yeah, it I'm new to it. 408 - > Um so the idea is that you write the test, you write the code to 409 - > make the test pass, but then who's to say uh that that 410 - > everything is as it should be? 411 - > Like, could it be that your test gives some sort of false 412 - > positive um And maybe you wrote too much application code beyond 413 - > what you really needed to satisfy the test, or something 414 - > like that. 415 - > So the mutation testing library will mess with your code and 416 - > it'll say, okay, this passes, but what if we tweak it like 417 - > this?
418 - > Do the tests still pass then? 419 - > Oh, because if it does, then like, gotcha, these tests are 420 - > invalid and we need more test cases and blah, blah, blah. 421 - > Um I I I found that to be a really powerful technique. 422 - > SPEAKER_00: Yeah, I think like these verification and testing 423 - > techniques that have been around, I think it's it's really 424 - > their time to shine.
425 - > Uh and so one thing we do uh at Amazon a bunch for these um um 426 - > for we build a lot of distributed systems like in AWS, 427 - > uh, you know, we have things like DynamoDB, things like S3, 428 - > where we have to prove that the replication, like these 429 - > distributed systems around failover during a network 430 - > partition while bootstrapping another replica, like all these 431 - > the this is the this is the stuff that personally I like I 432 - > love so much is all of these uh distributed systems algorithms.
433 - > And in order to prove these, um we do a lot of formal 434 - > verification methods, like TLA plus. 435 - > Uh, it's a really good verifier of a model. 436 - > It's like I have this model, uh, and and this the TLA just it it 437 - > beats it up. 438 - > You have your your here's how my system will behave, and it beats 439 - > it up with introducing like whatever latency at the critical 440 - > moment, and to see if if you still maintain an invariant, 441 - > like only one uh replica of a replicated database is the is 442 - > the leader at any point in time.
443 - > Like all the it just is very good at these kinds of synthetic 444 - > uh tests around all the different cases that you would 445 - > care about in a digit distributed system. 446 - > And I think these this is they are kind of hard to write. 447 - > Um they can be. 448 - > Um, these models, and then you write your code, which isn't the 449 - > model.
450 - > It's actually your code is a rep is a is a manifestation of that 451 - > model. 452 - > And so I think it's these methods kind of time to shine, 453 - > because you can now you can now write generate, write the model, 454 - > and then generate the code and and have the LLM reason about 455 - > and prove differences between the model, which already proved 456 - > your algorithm, and your code to see does it match the model. 457 - > There are a bunch of techniques around this, but I think it it 458 - > my point is just that it is it's the time for these formal 459 - > methods and and just testing methods to shine.
460 - > SPEAKER_01: Yeah, yeah, I totally agree. 461 - > Um I I have to take us on a small digression. 462 - > Um, my friend uh Stephen Baker mentioned the other day, uh he 463 - > he made an analogy with oil. 464 - > So uh back in the I don't know, mid-late 1800s, second half of 465 - > the 1800s, um oil wasn't used for gasoline and automobiles, it 466 - > was mainly used to make kerosene for lighting, and the market 467 - > size was limited.
468 - > Um Standard oil was was dominant at that time, but their their 469 - > main market was um kerosene. 470 - > Um but then uh industrialization came along, uh World War One 471 - > happened, um battleships converted from coal to oil, um, 472 - > people started buying cars, the the adoption of automobiles just 473 - > totally exploded, and all of a sudden it it was oil's time to 474 - > shine, you know. 475 - > It it had been there this whole time under our feet for millions 476 - > of years or whatever, uh and and then used for this limited 477 - > application, just for for lighting and and stuff like 478 - > that, but then the world changed and it got so much more useful, 479 - > indispensable.
480 - > Um and so these these testing techniques techniques and such 481 - > uh could be compared to oil. 482 - > We've had them for decades, and they've they've already been 483 - > great. 484 - > Uh but now it's like okay, there's there's I think two two 485 - > things happening. 486 - > One is a greater necessity for these techniques, because if you 487 - > don't use them, then AI will just go off the rails.
488 - > Um and another is that the the bar has been lowered. 489 - > Uh I I don't want to confuse my metaphors, but I've I've always 490 - > talked about AI kind of like a mental lubricant. 491 - > You know, it these these things that have friction, um, like 492 - > writing tests, for example, there might be some mental 493 - > friction there. 494 - > Uh it's AI is like WD40, which is not a lubricant, but um it 495 - > can unstick your mind.
496 - > Um and it makes these things which were previously difficult 497 - > much easier. 498 - > So those two things, uh I think, you know, like you're saying, 499 - > it's it's their time to shine. 500 - > Those those are uh hurrying that along. 501 - > SPEAKER_00: For sure.
502 - > I think uh the it's just you know, I've spent the last 20 503 - > years just with uh almost singular purpose in mind, and 504 - > that's just to make developers' lives easier. 505 - > That is the pattern that I have followed ever since I started at 506 - > Amazon and had to operate a database in order to automate 507 - > server operations because I needed to keep track of all the 508 - > servers. 509 - > I needed to operate a database. 510 - > Operating a database is harder than operating the servers.
511 - > And so then it's like, okay, how do I net how can I never have to 512 - > do this difficult database ops again? 513 - > Oh, I heard we're let's make this dynamo DB, let's make 514 - > NoSQL, so then we don't have to do database ops from then. 515 - > It'll just be managed. 516 - > That sort of rinse and repeat with like I've worked on Lambda 517 - > then for managing servers, the building serverless, so you 518 - > don't have to manage servers, API gateway because API routing 519 - > is tedious and and easy to get wrong.
520 - > Um so just I've been chasing this this like constant. 521 - > How do I build the next abstraction that's gonna make my 522 - > life easier? 523 - > And the really fantastic part about LLMs is that they can do 524 - > that just they're so flexible. 525 - > Like they can, they are that uh sort of this missing ingredient 526 - > to be able to do that.
527 - > Well, how do I how do I do the next part that I find tedious? 528 - > Like server ups, okay, now I don't have to do that. 529 - > But what about uh writing tests? 530 - > You know, right?
531 - > Like that is it is something that people we could have been 532 - > writing property-based tests and and uh mutation tests this 533 - > entire time. 534 - > But it's just the it's been tedious uh to do all the time 535 - > and spend all of our time. 536 - > It's not as fun. 537 - > And so great, let's have let's have the agent do that.
538 - > We have to now the work is to prove that it actually wrote 539 - > good tests. 540 - > You know, and and these frameworks can help help do 541 - > that, provide that evidence or the the enough of that guarding 542 - > guardrail, just like the spectrum development provides 543 - > some guardrails and and and a path to success. 544 - > Um, these tests can be a part of that too. 545 - > SPEAKER_01: Yeah, yeah, it's a really exciting time.
546 - > Um and and uh uh uh what you say is exactly true. 547 - > Um on the developer experience side, like I've invested a lot 548 - > in my own developer experience because it's um you know you you 549 - > pay a small price to get hopefully a large benefit, and 550 - > then you can go faster, and it's kind of a positive feedback 551 - > loop. 552 - > Um so maybe if we if we can get into that a bit, like what kinds 553 - > of stuff uh, and and it doesn't even have to be only AI, but 554 - > like what kind of stuff have you done to make uh developer 555 - > experience easier, especially things that people might not 556 - > like most companies might not do, that kind of stuff, just 557 - > anywhere you want to take that.
558 - > SPEAKER_00: Sure. 559 - > Uh I'd say a lot of it stems from the that the first team I 560 - > was on on Amazon.com, where I was on the team that ran 561 - > Amazon.com's web server fleets.
562 - > Um we had just a lot of servers that needed to get provisioned. 563 - > We had to forecast how many to physic to actually buy. 564 - > This was pre pre-EC2, pre-AWS cloud. 565 - > We had to buy servers every year.
566 - > So I had to manage keeping track of how many we have and how many 567 - > we need to buy, what their utilization is, forecast the 568 - > efficiency uh that we're going, like, is the website going to 569 - > get less efficient or more efficient? 570 - > Probably less efficient because that's the way that entropy 571 - > works. 572 - > Um, and so we had all these things. 573 - > And so one of the things that might be a little bit uh should 574 - > it seemed like it should be easy, but it wasn't, was uh 575 - > calling other services.
576 - > So I would make this a forecasting tool that would 577 - > figure out how many servers to buy every year, but I would need 578 - > to call and get the utilization data, the observability data 579 - > about how like CPU usage, requests to the website, all 580 - > that stuff. 581 - > And that was in this other web service, um, the monitoring web 582 - > services as it happened to be called, later became CloudWatch. 583 - > Um but calling that service was actually kind of annoying 584 - > because they it was a sort of a new paradigm around REST, where 585 - > it was like, okay, uh you don't have to generate a client 586 - > library.
587 - > It's like, but I'm writing a Java program. 588 - > It's like, no, just like, you know, use what just make HTTP 589 - > requests and fish around in the return to XML or JSON for the 590 - > response. 591 - > Okay, but I've I have a Java program here. 592 - > Like I do need to generate types or write types.
593 - > So I just found it sort of uh tedious to call some of these 594 - > services. 595 - > Whereas before we had some frameworks at Amazon that just 596 - > generated strongly typed stuff, but it wasn't REST under the 597 - > hood. 598 - > And so it was there was this just kind of a new way that 599 - > people were writing and new frameworks people were bringing 600 - > in. 601 - > And so I found it those made it maybe easier for the service 602 - > author to write a service because they could do it in the 603 - > newer, newer way and newer tools and newer frameworks, but those 604 - > had left the clients behind.
605 - > And so I found that tedious of like just calling services. 606 - > It's supposed to be easy. 607 - > And so I learned that there was a team making a new framework. 608 - > Um, we call it Coral.
609 - > It it its goal was to be a modern framework that that could 610 - > speak REST, but also speak all the other protocols that we 611 - > have, all the binary formats, all the everything to be back so 612 - > that it would fit into whatever tooling that clients wanted to 613 - > use. 614 - > It would sort of actually just make everybody happy, gener be a 615 - > nice modern like web service framework, but also support any 616 - > kind of client that people want. 617 - > Like also basically generate clients as well.
618 - > And that kind of became like the AWS SDK generation in a way. 619 - > Um so I found that this that's just an example of one of these 620 - > that I've chased down over the years. 621 - > That's it's like subtle. 622 - > It's like, okay, a web services are are tricky and uh and and 623 - > building the service is hard, but also making sure that you 624 - > are paying attention to your customers so that calling the 625 - > service is easy too.
626 - > And so uh this coral sort of actually in a way became this 627 - > API gateway uh thing to because frameworks are are also kind of 628 - > hard. 629 - > SPEAKER_01: Right. 630 - > Um okay, so this might be the the point in the show when we uh 631 - > wildly wander off the trail as inevitably happens. 632 - > Um but I I'm curious about something.
633 - > So like if you have a large system with a bunch of different 634 - > services and stuff like that, um there's a lot of different ways 635 - > you can have them talk to each other. 636 - > Um something I've seen that I haven't loved is when each 637 - > service basically talks to every other service. 638 - > Um and I I did some research because uh uh most of my career 639 - > has been at very small startups that haven't had this situation. 640 - > Um but I I I did some research once I started to work at a big 641 - > company, and I'm like, okay, like what's what what are the 642 - > different ways this could be architected?
643 - > And it sounds like what I was looking at was point-to-point 644 - > architecture, and I was seeking something more like hub and 645 - > spoke. 646 - > Um because the the thing that I the thing that makes me 647 - > uncomfortable about the point-to-point architecture is 648 - > every service had its own custom, unique way of talking to 649 - > every other service. 650 - > And then what happens when you add a new service? 651 - > Now you have to like start from scratch and incur the cost of 652 - > building ways to talk to all the other services, and you have to 653 - > have all the credentials and all this stuff, and it's just like 654 - > this this can't be the way to do this.
655 - > With a Hub and Spoke, it seems like you just talk to the hub, 656 - > everybody just talks to the hub, and it's like, hey, I want to 657 - > send a message. 658 - > I I'm not gonna integrate with Slack and tell Slack to send a 659 - > message, I'm just gonna notify the hub that something happened, 660 - > and the hub can say, Oh, okay, this thing happened. 661 - > Uh I'll send a message to Slack and or send an email, whatever, 662 - > whatever I, the hub, thinks need to happen.
663 - > Um I'm just curious, any commentary you might have on 664 - > that whole thing is somebody way more experienced than myself and 665 - > those kind of things. 666 - > SPEAKER_00: I've seen so many patterns of this um over the 667 - > years, like what work to different degrees. 668 - > I think there are different problems that that are it's 669 - > helpful to have more of a of this kind of hub and spoke. 670 - > Uh there was one uh like looking at the um, so I worked on the 671 - > Amazon.
com web server fleet. 672 - > That means the thing that ran that rendered HTML, ultimately. 673 - > And it needs to make a lot of service. 674 - > Get data about if you're looking at a product page, we need to 675 - > get data about the product in order to show decide how to 676 - > display it.
677 - > Um turns out the information about what is a product with 678 - > when you have when you can sell any kind of product, uh, whether 679 - > it's digital or physical. 680 - > And you know, when we started, we only had books, and then we 681 - > had physical products, then we had digital products, it just 682 - > evolves a lot. 683 - > And over time, yeah, we built uh one of a kind of a dedicated 684 - > around products. 685 - > We built some of what you describe as as a as one of these 686 - > hub services that would kind of aggregate all of the and and 687 - > deal with figuring out like, okay, given this request to look 688 - > to display this product, it could it had logic in it to 689 - > figure out where to get that from the different other 690 - > services that had a little piece of this information here, a 691 - > piece of information there.
692 - > Similarly, with like the order processing pipeline that would 693 - > because placing an order for something that is uh there's so 694 - > many different types of products, and so that you need 695 - > to be able to handle different products differently and the 696 - > workflow changes on the fly as things get um depending on who's 697 - > going to fulfill that order and everything. 698 - > So it's uh we've seen a lot of it. 699 - > I don't think there's any any one size fits all, but there is 700 - > definitely like a lot of people will do what you described with 701 - > uh with API gateway because it can you can just have that is 702 - > the API layer and it handles all the different types of API 703 - > requests, but um behind the scenes, some paths go to one 704 - > service or another, and you don't have to think about point 705 - > integrations like that.
706 - > So that can be nice. 707 - > Um, I do see think that one thing that worked particularly 708 - > well that kind of works no matter which, whether it is 709 - > point-to-point or hub and spoke or event bus or whatever is just 710 - > common frameworks. 711 - > If people are speaking the same, having the same identity, like 712 - > the company uses the same identity system for what is a 713 - > service, what identifies a service, and what authorizes one 714 - > service to talk to another.
715 - > Like in AWS, you you can that's one that's this identity access 716 - > manager. 717 - > IAM just is the policy that just and sort of a signing language 718 - > that that describes you with with one set of credentials, you 719 - > can scope those to be able to call any service. 720 - > You don't have to figure out how to speak a different identity 721 - > provider system or anything. 722 - > Um internally, there are others.
723 - > So by having some standardization, um, you can uh 724 - > I guess uh having standardization is is really 725 - > helpful because it because you don't have to uh it doesn't you 726 - > can actually choose whether hub and spoke works for you or 727 - > point-to-point or workflow or or anything because you uh but the 728 - > standardization makes it so that you don't have to pay the price 729 - > of of integration 10 times. 730 - > You're just okay, that's the advantage of this framework that 731 - > I looked on.
732 - > It's like the and it's still around. 733 - > Like uh pretty with a few exceptions, like every AWS API 734 - > call is is run through this framework, even though there is 735 - > no common API layer. 736 - > Like you still talk to different endpoints, but it's all the same 737 - > framework. 738 - > And so the SDK is all generated off of the same service 739 - > distribution, service definition library, uh their language.
740 - > It's called Smithy, it's actually open source now, um, 741 - > that defines services and what shapes and APIs and signing 742 - > methods for identity. 743 - > Um, even though they aren't, we don't, as all of AWS have one 744 - > API layer, this like Hub and Spoke. 745 - > It's still we don't have to reinvent the wheel every time we 746 - > make a new service or make a new client for that service. 747 - > SPEAKER_01: Yeah, that's very interesting.
748 - > I hadn't considered that possibility. 749 - > Um I'm trying to think of an analogy to understand it better. 750 - > Um I I like to buy all the same brand of power tools so that I 751 - > can take a battery out of one and stick it in another tool. 752 - > SPEAKER_00: Um I know your podcast is probably isn't 753 - > sponsored by any one of them, but what is the what is your 754 - > go-to?
755 - > Skill. 756 - > Skill? 757 - > Okay, okay. 758 - > I've I I kind of the sorting hat put me into the uh Mikita line 759 - > at one point.
760 - > Uh so that's to mix to mix some metaphors there. 761 - > SPEAKER_01: But yeah, well, when my wife and I got together, she 762 - > already had some skill tools. 763 - > And and she has an insistence, which I agree with, on having 764 - > everything uniform. 765 - > So I was like, well, skill is what we have, and I think their 766 - > tools are pretty good, so that's what we're gonna go with.
767 - > Um yeah, okay, so the the the takeaway for me there, a 768 - > takeaway, is Hub and Spoke isn't the only answer. 769 - > You don't need to have the hub as the intermediary in order to 770 - > get some efficiency gains. 771 - > Um you can save the cost of those uh unique uh uh 772 - > point-to-point integrations by having something in common 773 - > between them that allows them to talk to each other. 774 - > SPEAKER_00: Yeah, exactly.
775 - > I think identity and uh and uh especially framework for for 776 - > protocol help a lot. 777 - > Um another kind of cool aspect of this was uh basically like 778 - > make enough the the key with the framework was to make it uh so 779 - > um is to just solve a lot of the problems that people would have 780 - > with point-to-point in a standard way, like metrics and 781 - > observability, for example. 782 - > Um, if you are going to call some other service and you have 783 - > some other framework for that client, you'd have to add your 784 - > own instrumentation.
785 - > Um, where just to say, okay, when I call a service, I need to 786 - > record whether or not it succeeded, what kind of error it 787 - > gave back, specifics about the error it gave back, how long the 788 - > latency was. 789 - > I need to bucket the latency intelligently. 790 - > Like if I just measure the latency, I could see my latency 791 - > get really good, really low, but actually all those requests are 792 - > failing. 793 - > So that I shouldn't be happy that the latency is better.
794 - > It's like it, so I need to actually have the right 795 - > dimensionality of the observability to say successful 796 - > request latency. 797 - > And what does success mean? 798 - > Okay, well, 201, 200, like what HTTP stuff is happening under 799 - > the hood. 800 - > So having the right observability is actually pretty 801 - > tough.
802 - > Now, open telemetry exists now. 803 - > That is a nice framework that that instruments with tons of 804 - > framework. 805 - > It's an observability framework. 806 - > I think that's fantastic, and it's changed a lot for the for 807 - > the industry around being able to get this just correct 808 - > automatic instrumentation for any client server, any database, 809 - > any whatever anybody has come up with um to have just like great 810 - > instrumentation that way.
811 - > But that was this was before that, and it it just kind of out 812 - > of the box gave everybody instrumentation that was 813 - > compatible with our observability tool we had 814 - > internally, which is Cloud Watch. 815 - > Um so it just like that's just making that framework just do a 816 - > lot for you, but also be flexible enough so that you 817 - > don't feel like suffocated by like a framework that's 818 - > overbearing. 819 - > Uh, that's the dance. 820 - > SPEAKER_01: Yeah, well, a framework can give you um 821 - > leverage.
822 - > Um okay, so here's here's something I've been preaching 823 - > lately to people I work with. 824 - > Um and it's it's taken me a while, but I've gotten the 825 - > message across, at least to some people, I think. 826 - > Um I talk about the concept of universality, and some tools, 827 - > some things possess universality and some don't. 828 - > Um, for example, uh Grafana.
829 - > Uh it it lets you plug in data sources and it gives you visual 830 - > graphs. 831 - > It absolutely does not have universality. 832 - > Um you can customize dashboards and stuff like that, but like no 833 - > matter how you slice it, you're gonna end up with a dashboard of 834 - > graphs. 835 - > Um whereas if you build, for example, a Ruby on Rails app.
836 - > Application or a JavaScript application, just something 837 - > using code, it does have universality. 838 - > Anything that can be built, you can build. 839 - > And there's a trade-off there where you might have to do some 840 - > things completely from scratch. 841 - > And then there are tools which give you like both at the same 842 - > time, which is it it gives you the leverage, but it doesn't pin 843 - > you down so that you can only do what the tool lets you do.
844 - > It gives you a base to build on top of rather than something 845 - > that uh that that squeezes you in and limits you. 846 - > SPEAKER_00: I think that's such a cool and and actually uh this 847 - > universality is such an like a applicable to like just to full 848 - > circle this to to to agent uh stuff is is actually very cool 849 - > because universality is is exactly why I'm excited about 850 - > and why I joined from like from being on on like uh you know 851 - > infrastructure service teams to being in working on agents 852 - > because I see this problem that I've always been trying to chase 853 - > around improving operations, like improving DevOps, like 854 - > every every annoying part about building and operating software 855 - > that's like somewhat tedious, but so important.
856 - > I've been I've been trying to just solve that. 857 - > And so because the universality of agents, um, I think we we've 858 - > figured out can be applied in some pretty amazing ways here 859 - > around that. 860 - > Like you mentioned Grafana, like as or like observability tool. 861 - > Like you can, it's pluggable to so many things.
862 - > That's one of the things that's so great about it is pluggable, 863 - > pluggability to so many data sources. 864 - > Um, but when I'm trying to solve, like observability is a 865 - > is a means to operations. 866 - > It's a means, not an end. 867 - > Like you don't do observability for observability's sake.
868 - > You you do it so that you can operate um well. 869 - > And and so we tackled that when the thing that I'm working on 870 - > currently that I'm the most I'm just so excited about, like over 871 - > the last 20 years, is is we've made an agent that does DevOps 872 - > for you, the AWS DevOps agent. 873 - > And it is um DevOps means a lot of things to a lot of people. 874 - > Um to us, it means just what the DevOps team does, right?
875 - > No, no, to us, like I thought I've done DevOps this entire 876 - > time, like because it developers do the ops. 877 - > There is no, yeah, exactly, no DevOps team. 878 - > Uh and so um, and that's what what I I've always liked wearing 879 - > all the hats. 880 - > Like I also want to pay attention to security.
881 - > I want to pay attention to what customers are saying, uh, and 882 - > and where and so I can make the product better and support them. 883 - > Like I want to wear all of the hats and I love that. 884 - > And so with DevOps, like that's that my that lifestyle of 885 - > wearing all the hats. 886 - > SPEAKER_01: And yeah, and I just want to say, um in in case not 887 - > everybody caught that that joke of ours, um, you know, people 888 - > put those two words together, DevOps team, which is like the 889 - > antithesis of of DevOps.
890 - > DevOps is where you put the dev uh the development and the 891 - > operations together, that idea, as I think of it, of like you 892 - > build it, you run it. 893 - > Um so that's why the term DevOps team is so funny. 894 - > SPEAKER_00: Yeah, I the nothing against specialization and 895 - > everything. 896 - > Like for people who do uh who are like are DevOps engineers 897 - > and everything like that.
898 - > I think that that specialization has been has been really 899 - > powerful. 900 - > Even even at Amazon, where you know we're we we just we're all 901 - > software developers, there are some who are just a lot better 902 - > at at automating ops. 903 - > And that's just their kind of flavor and focus. 904 - > And and like, and I think I'm one of those, like compared to 905 - > whereas I'm not as good at other like maybe abstractions and 906 - > frameworks and stuff like that.
907 - > SPEAKER_01: To me, the significant thing, obviously, 908 - > individuals vary in their strengths and inclinations and 909 - > and backgrounds and stuff like that. 910 - > Um, to me, one of the significant things has to do 911 - > with incentives, because if you have all the developers over 912 - > here and then all the SRE people over there, um the the 913 - > developers might have needs that the SRE people don't have any 914 - > natural reason to care about, uh other than like hopefully out of 915 - > the goodness of their heart, but that's not like why people do 916 - > things at work.
917 - > Um and and so that creates really uncomfortable situations. 918 - > Whereas if you slice that up differently and put some devs 919 - > and some SRE people in the same team, then the the interests are 920 - > are more aligned. 921 - > SPEAKER_00: That's right. 922 - > And there's this a natural like back pressure and reacting to 923 - > pain, uh, you know, which is which is I think so important.
924 - > So instead of like negotiating a contract and everything like 925 - > that, which is you know, and there are a lot of philosophy, 926 - > there are a lot of these that work. 927 - > I've talked to so many customers of AWS over the years, like back 928 - > and forth about, hey, what do you do? 929 - > Oh, what do you do? 930 - > And like, and we just kind of there are a lot of different 931 - > places and models and everything that work for different 932 - > different companies.
933 - > I of course have my I'm you know, I'm very biased and and 934 - > opinionated about DevOps being this what we're describing as 935 - > DevOps, like to be the true Scotsman fallacy of like you 936 - > know, defining well, you're not doing real DevOps, yeah. 937 - > Okay, so but uh I'm a fan of this model. 938 - > I think I I think it works with how I like to to work and and I 939 - > think it results in really good outcomes for customers because 940 - > the developer teams are just you know, ultimately ops is a and 941 - > observability are a lens with which to understand the customer 942 - > experience.
943 - > And I uh and I really like that. 944 - > I like being so connected to it to how how are our customers 945 - > doing? 946 - > Well, let's let's look at how the service is operating. 947 - > Oh, yeah, okay.
948 - > Like, well, I don't know how if customers are happy based on my 949 - > observability data. 950 - > It's like, well, then we're missing some observability data. 951 - > Let's let's measure customer experience better. 952 - > SPEAKER_01: Yeah, and that's a whole podcast episode in itself, 953 - > at least.
954 - > It's probably many, many podcast episodes, is like um, how much 955 - > of the story uh do metrics tell? 956 - > Um it's of course not the whole thing. 957 - > Um the I I had an experience where uh numbers were being 958 - > cited for months and months, and then finally uh we were like, 959 - > hey, let's like go and actually talk to some people. 960 - > Um and it turns out we had totally the wrong idea because 961 - > we were just looking at the numbers.
962 - > SPEAKER_00: Yeah, measure it from the right place, got to 963 - > talk to people, constantly check your assumptions about whether 964 - > that's one of the kind of cultural things that's just so 965 - > important. 966 - > And I think what helps with DevOps helps a lot is that you 967 - > know it at AWS teams uh get together every week. 968 - > Like, and actually all of AWS gets together every week. 969 - > Every week, like for two hours.
970 - > And we talk about ops. 971 - > We talked about what's going on with of things things that are 972 - > size of formula, like every what what were the wins? 973 - > What did somebody do that they want to just like tell the world 974 - > about? 975 - > Hey, we just used this new tool or we made this new tool, we had 976 - > this big efficiency gain in the service, or or whatever.
977 - > SPEAKER_01: And what do you mean? 978 - > What do you mean everybody in AWS gets together like on a call 979 - > in person? 980 - > SPEAKER_00: Uh yeah, every uh we used to be in person, but every 981 - > every team is represented by at least one person. 982 - > Um, a lot of people are on uh like a live stream uh watching.
983 - > Anybody can participate on like we joined these days, kind of we 984 - > were kind of forced this way, uh, but I think for the better, 985 - > um, with COVID and everything, we went we went virtual for 986 - > this, and we've stayed that way. 987 - > We find that it's easier for everybody to participate uh 988 - > without the kind of intimidation of being in a large room where 989 - > you're trying to yell from the back and you don't necessarily 990 - > want to. 991 - > So um we have a channel, a Slack channel where we're all talking 992 - > and uh about it during the meeting.
993 - > Uh the kind of the side conversations are very 994 - > interesting and ongoing. 995 - > But yeah, there are people from from every team, um, every 996 - > single like not just every service, but every team within 997 - > every service um who attends uh and and watches and uh and and 998 - > discusses like what what what went well, like what maybe we'll 999 - > talk about retrospectives of things that that didn't go well 1000 - > that we wanted to share with everybody and and share how we 1001 - > think.
1002 - > Here's how I think about things that didn't go well and how to 1003 - > improve. 1004 - > Uh like uh and then uh we'll the agenda varies, but sometimes 1005 - > we'll actually what we used to do, especially more was we would 1006 - > pull up a random services dashboard, like in with a 1007 - > metric, a metric dashboard, and say, well, and then we would we 1008 - > would ask each other, like there would be uh a line that would 1009 - > show, say, uh no, here's uh oh, here's an error graph.
1010 - > Uh but these are just the 400s, these are the four XX HTTP 1011 - > status code responses. 1012 - > So those aren't those aren't server faults, those are clients 1013 - > calling us wrong. 1014 - > And people would ask, well, okay, like if that spikes 1015 - > though, like somebody isn't happy, they're not successful. 1016 - > Like, is there some some sharp edge in the service that's 1017 - > causing that?
1018 - > And we'll discuss that. 1019 - > Oh, no, those are just people exceeding their rate limit. 1020 - > Well, okay, like are they happy that they exceeded their rate 1021 - > limit? 1022 - > Maybe you could just increase their rate limit for them.
1023 - > And then we're we have good discussions about what like what 1024 - > limit, you know, about limits around input validation and 1025 - > everything about and that really do get to the question of, okay, 1026 - > is is this the right customer experience? 1027 - > And we kind of challenge each other on that. 1028 - > SPEAKER_01: Yeah, yeah. 1029 - > I so I I like can't help but think in analogies and something 1030 - > okay, so I've I've been like re-studying calculus lately, and 1031 - > I found this really nice video.
1032 - > Um there's this YouTube channel that's just called Math and 1033 - > Science. 1034 - > It seems to be this one guy. 1035 - > Um, and he did a really nice job of explaining what calculus was 1036 - > all about. 1037 - > And he used what I guess is a fairly common example of like 1038 - > here is a graph of um a particle's uh position over 1039 - > time.
1040 - > And then if you take the derivative of that, you get the 1041 - > um let's see, there's there's a position, velocity, and I don't 1042 - > remember what. 1043 - > Sorry? 1044 - > Exactly, yeah. 1045 - > And and so I like to take these things and take like the the 1046 - > derivative of each one.
1047 - > It's like okay, we have a culture of culture of 1048 - > firefighting and and we need to turn that around. 1049 - > We need to like be more preap proactive and have better 1050 - > monitoring and alerting so that we don't have that reactive 1051 - > culture of firefighting. 1052 - > It's like, hang on a second, why do we have a culture of 1053 - > firefighting? 1054 - > One of my favorite quotes is things are the way they are 1055 - > because they got that way.
1056 - > And so, like, if this is the way, like, why are we just 1057 - > looking at this now? 1058 - > Like, why didn't somebody look at this five years ago and 1059 - > straighten this out five years ago? 1060 - > Like, why now and not back then? 1061 - > Unfortunately, it's not very common to dig into that.
1062 - > unknown: Yeah. 1063 - > SPEAKER_00: You're describing kind of one of these ops 1064 - > meetings pretty well. 1065 - > Like they this gets second and second derivative pretty fast of 1066 - > like, okay, it's like, and people say, oh, okay, why 1067 - > there's this, you know, something happened in a service. 1068 - > Oh, like well, why does that happen in services?
1069 - > Why do like why isn't this just out of the box? 1070 - > So oh yeah, we didn't have this uh alarm on on like file 1071 - > descriptors or something. 1072 - > Oh, okay, well, why doesn't why aren't those already always the 1073 - > case? 1074 - > And so we do talk about that whole that whole like shift left 1075 - > aspect of it of how do we just get everybody to be better, 1076 - > which which like back to that universality point you're 1077 - > talking about.
1078 - > Like this is like agents are really that you go, there's a 1079 - > lot of the culture you cannot bypass the culture around of 1080 - > Roundups about making things better. 1081 - > There is no magic wand. 1082 - > There's no compression algorithm for experience. 1083 - > But gosh, agents are pretty, pretty universally adaptable in 1084 - > both root-causing issues and plugging into and understanding 1085 - > any system.
1086 - > Doesn't matter what framework it has, doesn't matter what 1087 - > observability provider it has, there's an MCP server for it. 1088 - > It'll figure out how to get the right metric out that it's 1089 - > looking for. 1090 - > Um, for for the shift left aspect of it, scan through an 1091 - > application, look at all the observability, look at all the 1092 - > incidents, look at the code for it, match it against things that 1093 - > we know are are like not good, things that opportunities to 1094 - > improve, and just do those.
1095 - > So this like the both the React and also the improve, the 1096 - > proactive part are what we've built into DevOps Agent uh last 1097 - > sales pitch on DevOps Agent. 1098 - > But it it's like that universality and like being able 1099 - > to adapt to to to really any learn an architecture, learn an 1100 - > application just by looking at everything it can tell from it, 1101 - > uh the whatever tools you throw at it, whatever MCP servers. 1102 - > Um it's just it's working better than I expected almost.
1103 - > Like it's it's pretty incredible the universality of uh of 1104 - > agents. 1105 - > SPEAKER_01: By the way, this is by far the fastest uptake I've 1106 - > gotten on on the term universality and having it make 1107 - > its way into the conversation. 1108 - > Um yeah, I got it from this book, um, The Beginning of 1109 - > Infinity by David Deutsch, or maybe it was The Fabric of 1110 - > Reality by David Deutsch. 1111 - > Both those books are are excellent and they've they've 1112 - > changed the way that I think about programming and just like 1113 - > the life in general.
1114 - > Um life-changing books, especially the beginning of 1115 - > infinity. 1116 - > Um where we're about at time, um there's there's one more thing 1117 - > that I have to ask, but let me know if you get to where where 1118 - > you need to have a hard stop. 1119 - > Um I want to ask about postmortems. 1120 - > Um something that really frustrates me in po postmortems 1121 - > is when a large number of people are brought together for a short 1122 - > amount of time, and the the incident is analyzed maybe 1123 - > superficially, and then out of the post mortem come uh several 1124 - > action items.
1125 - > It's like here here are the seven things we have to 1126 - > implement immediately to make sure this never happens again. 1127 - > And and I'm like, uh, is is this really the the way to do it? 1128 - > I think maybe we should have fewer people for a much longer 1129 - > time instead of like an a one-hour postmortem. 1130 - > You know, some postmortems maybe maybe they call for 30 minutes, 1131 - > some maybe they call for like three days of of picking it 1132 - > apart because different things call for different levels of 1133 - > thought and stuff like that.
1134 - > I don't know to you what what makes a good postmortem? 1135 - > SPEAKER_00: I one of my favorite things to do is to is to write a 1136 - > postmortem doc. 1137 - > We call them COE correction of error. 1138 - > You can take any name.
1139 - > Um, but we have a pretty pretty standard formula for these, but 1140 - > I love writing them. 1141 - > They're a document um that a team will write. 1142 - > Like a generally like a maybe one or or two people will spend 1143 - > a lot of time with the pen, but then that they kind of just can 1144 - > keep riffing on it until with the like as a team and as then a 1145 - > larger team if it again, if it needs more scrutiny, uh talk 1146 - > about it. 1147 - > But first, what makes up a one of these COEs, we call it.
1148 - > Um, it's a description that we we write them for anybody in the 1149 - > company to be able to read. 1150 - > Like to to and and so we want to make sure we're not using too 1151 - > much jargon about oh, the the uh the BSF service called the you 1152 - > know, MVP service called that we do, you know, we don't want to 1153 - > use too many acronyms, or we certainly will put a like a 1154 - > description of it in it. 1155 - > But we say, here's what it is, here's what happened, here's the 1156 - > customer impact, like pre- uh here's the timeline of like what 1157 - > break it down really granular.
1158 - > The timeline includes uh when we learned about something, when an 1159 - > alarm went off, when the impact happens, when we took any 1160 - > action, every when everything. 1161 - > Um, and then the meat of it is the the whys, the five whys. 1162 - > My understanding is this is a Toyota method for for kind of 1163 - > looking back at when something happens. 1164 - > Say, well, uh service fails.
1165 - > Why did it fail? 1166 - > Uh it ran out of something. 1167 - > Well, why did it run out of something? 1168 - > Uh well, it ran out of something because we didn't have an alarm.
1169 - > Why didn't we have an alarm? 1170 - > So you just keep asking why. 1171 - > You don't ask five. 1172 - > And and these answers, it's not just a one-sentence response to 1173 - > the why.
1174 - > Like, that's where I put all the thinking of like, well, okay, 1175 - > gosh, well, how do I really keep this from happening again? 1176 - > And then invariably you think, like, well, how would anybody 1177 - > else like I have people have that kind of, I think as 1178 - > engineers, you really maybe something just built in as being 1179 - > a programmer, but culturally, of you really want to help other 1180 - > people, like other teams, other people not have not face that 1181 - > same thing.
1182 - > That's why people really like writing frameworks, I think, is 1183 - > because you're like, oh, let's help other people do that thing. 1184 - > Uh, similarly with this, like, how do how would I make it so 1185 - > that uh that this like that everybody has an alarm on high 1186 - > CPU? 1187 - > Exactly this way. 1188 - > It's like I had an alarm, but I've alarmed on the median host 1189 - > CPU instead of really, it should always be the well, you should 1190 - > have that, but you should also have the the max, the P100, 1191 - > hundredth percentile host CPU to see like because maybe you're 1192 - > because it's really a distribution across your fleet.
1193 - > So all these little nuanced things. 1194 - > How do I make it so that and who do I talk to to who might have 1195 - > some some leverage, some tool that could be uh could apply 1196 - > this to everyone kind of out of the box? 1197 - > What is the abstraction that we could have? 1198 - > And just I love just unpacking and writing one of these 1199 - > documents that um where you're just digging into the data for a 1200 - > long time.
1201 - > Like this isn't a is a short thing. 1202 - > Um and uh it's a team, like I said, so in terms of review, 1203 - > then um some of these are like it they're always be reviewed as 1204 - > as a team, like uh whether that team's like 10 people or 1205 - > something, but then and also as a as a larger like a group of 1206 - > people, uh you know, maybe a larger team, maybe you bring in 1207 - > outside, help to be, oh, I wonder what this person thinks 1208 - > about how I got here.
1209 - > And so I'll just invite other people to review it. 1210 - > And and some of them, like this is part of that that Wednesday 1211 - > ops meeting I was talking about. 1212 - > Some of them uh we'll actually review every in a given week for 1213 - > for a while. 1214 - > Um, and and as a and talk about it and talk about our own the 1215 - > patterns that we've noticed over time.
1216 - > People will weigh in with, oh, I've seen this happen a bunch. 1217 - > And and uh and so like what do we do to what do we do to keep 1218 - > this from happening for any service? 1219 - > Um we also are analyzing. 1220 - > SPEAKER_01: I don't know, David.
1221 - > This all sounds like it takes a lot of time. 1222 - > That doesn't sound very efficient. 1223 - > SPEAKER_00: Oh, I don't know. 1224 - > SPEAKER_01: I I'm I'm speaking as a character right now.
1225 - > SPEAKER_00: Of course, yeah. 1226 - > Oh, there's no compression algorithm for experience. 1227 - > Like the thing is with like, and this isn't you know, maybe it 1228 - > there is just a a there are different businesses with 1229 - > different incentives. 1230 - > Like I like that ultimately that is totally totally the case, 1231 - > different different requirements.
1232 - > Um, I mean at AWS, like everybody is counting on us to 1233 - > be like we people to be perfect, right? 1234 - > So we we have to that that is the product, like is that we are 1235 - > doing ops so you don't have to. 1236 - > Like over the last 20 years, like that's like that's what we 1237 - > are doing. 1238 - > We are we are your ops team in a way.
1239 - > Uh not not not your application ops team, obviously, but like 1240 - > for the infrastructure, you're relying on so we spend an 1241 - > enormous amount of effort on this, obsessing over this. 1242 - > Um and uh yeah, I mean, we even have I didn't even describe 1243 - > everything. 1244 - > Like there are also teams. 1245 - > If you are a team, when I was on that framework team, for 1246 - > example, or when I was on the observability team, I would look 1247 - > at every one of these uh COEs, like we have it's a database of 1248 - > them, I can mine them and you know learn, like actually with 1249 - > LLMs, now it's even easier for me to mine them for for like 1250 - > patterns.
1251 - > And I'm like, well, if I own CloudWatch, that's the service 1252 - > I'm building, how could I make that just so much more out of 1253 - > the box and easier so that everybody has the alarm that 1254 - > they need all the time? 1255 - > Like, what do I need to build into my service? 1256 - > Like, how can I improve my service by seeing how how people 1257 - > maybe struggle with it or when some things don't always go 1258 - > exactly right? 1259 - > Um, every team I was on, it was like that.
1260 - > Okay, like, you know, when I was on on Dynamo DB, people would 1261 - > say, well, yeah, I had an outage because I like hit a throttling 1262 - > on my table. 1263 - > It's like, okay, how do I how do we add auto-scaling to DynamoDB? 1264 - > Or like make it so we we started DynamoDB with this notion of 1265 - > provisioned throughput because people wanted super predictable 1266 - > performance. 1267 - > So we would pre-provision capacity and like you could 1268 - > because it was elastic, you could call an API to add more or 1269 - > take away database capacity, and we would just do that behind the 1270 - > scenes, but um, we didn't have auto-scaling at first.
1271 - > And so when we saw people run into uh, oh yeah, like it's it's 1272 - > hard to actually monitor this all the time and then dial it 1273 - > up. 1274 - > So let's just build that into the product, this like real-time 1275 - > elasticity into it. 1276 - > So, like by understanding where things go wrong for your 1277 - > customers, you you make your product better. 1278 - > Uh whether it's whether it's going wrong because the service 1279 - > has a rough edge, or whether you know we need to just be better 1280 - > ourselves.
1281 - > Uh, and both both cases apply. 1282 - > So these these these root uh retrospectives, these uh COEs, 1283 - > they we get so much out of this process uh because it is, it 1284 - > takes a lot of time, but we we we really get a lot out of it 1285 - > from different angles, from the team itself, from other teams 1286 - > not running into the same issue, to just making the underlying 1287 - > services that we build for all of you better by looking at how 1288 - > we use it ourselves.
1289 - > SPEAKER_01: And I imagine that the patience and rigor pay 1290 - > themselves back many times over. 1291 - > SPEAKER_00: I hope so. 1292 - > I mean, uh yeah, I hope I hope that yeah, I mean, I hope that 1293 - > people, you know, I hope that we are having the operational 1294 - > outcomes that our customers need. 1295 - > That's all that matters to me.
1296 - > So it I uh we're we're always looking And and wishing we could 1297 - > do more, spend more time on it, uh it'd be better. 1298 - > So uh always looking for that next next tool to give our 1299 - > leverage uh that will make things better. 1300 - > So uh always looking to to to figure out how to how to keep 1301 - > keep that bar going up. 1302 - > SPEAKER_01: Mm-hmm.
1303 - > Well well David, I have about a million more questions for you, 1304 - > um, but unfortunately we don't have unlimited time. 1305 - > Um I have to say that this has been this this episode has had 1306 - > among the highest uh nuggets of wisdom per minute uh uh of of 1307 - > any episode I've done. 1308 - > So I really I personally have gotten a lot out of this. 1309 - > SPEAKER_00: And that's partly because different directions.
1310 - > It's been it's been a lot of fun just visiting different things. 1311 - > SPEAKER_01: Yeah, yeah. 1312 - > Um so yeah, I really enjoyed it. 1313 - > Uh before we sign off, anything that you want to share, links 1314 - > you want to send people to, that kind of stuff.
1315 - > SPEAKER_00: Uh oh sure. 1316 - > I mean, certainly. 1317 - > Um I guess okay, two two things. 1318 - > Of course, AWS DevOps agent, uh, you know, that that's a uh I 1319 - > think it's I I think it's gonna change a lot in terms of where 1320 - > as you code more and ship more so much faster with agents, the 1321 - > the the the bottle the code just piles up unless you can get it 1322 - > shipped and run it.
1323 - > So I think that's very important. 1324 - > The second thing, um, when we talk about what we've learned as 1325 - > AWS over the last 20 years, um one thing that we did a few 1326 - > years ago kicked off. 1327 - > Um we we were trying to help like just capture this like to 1328 - > for ourselves of like, well, what have we learned? 1329 - > Like how what is the right way to build a service and operate 1330 - > it?
1331 - > So it was like a like we we were just doing that for ourselves. 1332 - > We had been a lot of um talks over the years that we give 1333 - > every week that that so how would how do we kind of distill 1334 - > that? 1335 - > So we wrote um we launched this thing called the Amazon Builders 1336 - > Library. 1337 - > We wrote these really long articles, honestly, like they 1338 - > got a little longer.
1339 - > It's like, oh, just write down write down what does it mean to 1340 - > do instrumentation of a service. 1341 - > It was like, oh, wait a minute, that's gonna take me like 12 1342 - > pages to to touch on, right? 1343 - > And so we wrote we published this Amazon Builders Library, 1344 - > um, uh AWS.amazon.
com slash it yeah, AWS.amazon, AWS.com or 1345 - > slash builders library, just just Google it, I guess. 1346 - > Builders Amazon Builders Library.
1347 - > And it uh it has a lot of articles in depth about 1348 - > everything from fairness in multi-tenant systems to how we 1349 - > do CI C D. 1350 - > Um, a lot of ways to go about that. 1351 - > I think we have three articles about CI C D to uh to shuffle 1352 - > sharding of like how to make multi-tenant systems at large 1353 - > scale be um uh seem like single tenant systems. 1354 - > Uh so yeah, I would I would recommend that folks check that 1355 - > out.
1356 - > It's uh uh something we've put a lot uh of in into, but uh you 1357 - > know there's all we're always learning something new, so we 1358 - > will add articles from time to time. 1359 - > SPEAKER_01: Wow. 1360 - > Okay, that is gonna be a gold mine for me and my team. 1361 - > So I already Googled that on my phone just now.
1362 - > I'm gonna be digging into that later. 1363 - > Um Yeah, I I certainly will. 1364 - > Um again, this has been extremely educational. 1365 - > And David, thanks so much for coming on the show.
1366 - > SPEAKER_00: Thanks a lot. 1367 - > Uh oh yeah, and I guess last link is uh yeah, I'm on Twitter, 1368 - > LinkedIn, etc., uh the Barrett Blue Sky, whatever social 1369 - > network is in the fragmented current world. 1370 - > Find me on there, David Janichek.
1371 - > Uh, and uh happy to chat with anybody anytime. 1372 - > So uh but yeah, thank you so much for having me on. 1373 - > It was this was a really fun conversation. 1374 - > SPEAKER_01: Thank you.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.