
Arrested DevOps · 2024-01-18 · 48 min
Chelsea Troy's journey to staff data engineering reveals a crucial insight about technical leadership: as you advance, the work shifts from writing code to managing systems that are fundamentally about people, context, and coordination. At Mozilla, her MLOps team is tackling the challenge of operationalizing machine learning across the organization - moving beyond individual data scientists heroically shipping models to a consolidated, turnkey solution. The key innovation isn't just selecting products like Weights & Biases, but recognizing that satisficing versus optimizing metrics matter deeply. While engineers typically over-optimize across all dimensions (speed, features, scale), Troy argues that responsive, available support is the real optimization metric for MLOps tools, especially early-stage products from bootstrapped companies. Her team had to fight for this perspective, but discovered that products with weak support never made it to production despite potentially better code or ideology. The lesson extends beyond Mozilla: knowledge work creates value through context-building and socialization of changes, not through lines of code pushed per day.
Chelsea Troy is a staff data engineer at Mozilla working on the machine learning operations team. The team helps other teams across Mozilla get their machine learning models into production by providing standardized tools, processes, and recommendations - consolidating effort that was previously scattered across individual data scientists making independent deployment decisions.
Because success in senior roles depends on people understanding and wanting to use what you build. You must socialize changes, coordinate across teams, manage context loss, and ensure people know tools exist - all of which are management and marketing functions, not just code-writing.
Optimizing metrics are things you always want more of (like speed or features), while satisficing metrics have a threshold of 'good enough' beyond which more improvement doesn't matter. Most organizations treat everything as optimizing metrics and get stuck in decision deadlock; better decisions come from identifying just a few true optimizing metrics.
Because early-stage MLOps products constantly have things breaking and not matching documentation. The support team ends up being the most important interaction point for client engineers - Chelsea's team interacts with support multiple times weekly. Products with weak support never made it to production despite other advantages.
Assembly-line work creates linear visible artifacts (nails) throughout the day. Knowledge work creates value by coalescing context from many places into understanding, which looks like nothing externally until a big thing gets released - months of talking to users, understanding problems, and socializing changes happen before code is even written.
Computed from the transcript - who did the talking, and the words that came up most.
Jessitron is joined by Chelsea Troy, Staff Data Engineer at Mozilla, and one of the all-around most interesting people in software today, to discuss staff engineering, machine learning operations, and maybe also surfing.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Foreign. It's time for Arrested DevOps, the podcast where we help you achieve understanding, develop good practices and operate your team and organization for maximum DevOps awesomeness. I'm Jessica Kerr and here is a sponsor.
Speaker B: So Eufizi is a platform for platform teams. You can stand up your developer platform in minutes, not months. What I like about Eufizi is that it gives platform teams control and dev teams autonomy. It's kubernetes, native and extensible, so you can customize it with tooling that meets your team's evolving requirements. And these clusters, they spin up fast. Like super fast out of the box. Eufizi combines a great dev experience, secure, multi tenancy and cost efficiency. But try it out for yourself@uh, eufizi.com download their CLI and you can spin up your first sandbox cluster in under a minute on their free starter tier. That's ufizi.com, u f f I z z I.com let's face it, no one likes writing or maintaining documentation. But when you start a technical project or pick up a new task, missing information can cost you valuable time. Gitbook is a technical knowledge platform that fills that information gap, making it easy for your team to capture, maintain and find information from a single source of truth. For example, with Git sync you can set up a two way sync between your repository and Gitbook so you can turn markdown files into awesome user friendly docs. And if you make a change in your code base, the edits sync between the two automatically. Or what about when you need to find something in that knowledge base? Forget about searching, just ask Gitbook AI. You'll get a neat summarized answer that is sourced directly from your docs. These are a few examples of what Gitbook can do, so why not give it a try? Head to orestadevops.com gitbook to find out more.
Speaker A: All right, today I get to talk to Chelsea Troy, staff data engineer at Mozilla and one of the most all around interesting people in software. Today I get to ask her about staff engineering, machine learning, operations and maybe also surfing. So Chelsea, what's your job now?
Speaker C: I work as a staff, uh, data engineer at Mozilla and I am working on the machine learning operations team. The idea being that we have various teams doing machine learning work around Mozilla and in order to best consolidate effort and attention, we have this team that's dedicated to helping uh, the other teams get their models into production. So that's what I'm doing most of the time these days, and I like it. It has been a shift from work that I did previously because I had this idea at the beginning of my career that what happens is you're an engineer and you go along and you get promoted and you get promoted and then eventually you have to choose whether you're going to go along like a senior engineering track or a management track. And my plan was I'm going to not do management. I'm just going to keep going along the senior engineering track. And that way I will be able to write code for my entire job forever. And the thing that I have since learned, that was wrong. It turns out that's not how it works. My new conclusion that I have reached is that all jobs, if you do them long enough, become either management or marketing or some combination of management and marketing. Um, because ultimately you're in charge of products that you need to make sure that everybody knows how to use. You need to make sure that everybody who needs them knows that they exist, whether that's internally or externally. And, and then you also have the coordination within teams, the coordination across teams. They're ultimately, if you're going to do something for a long time and get to the point where you are wielding the leverage to ensure that products happen, then at, uh, that point the problems are. I don't want to say that they are people problems because the people are not the problem. But at that point the skills that we, that early career developers tend to think of as soft skills and inessential and not related to their promotion become the entire job, just about. And it's interesting to now be here and think about ways to word this because clearly the ways that we word this don't resonate across tech broadly because you get sort of the people who've already experienced it and they agree with it, and then you get the people who haven't experienced it yet who think it's complete BS and it's, it seems to be a mistake that people have to make themselves. Like, you know, when you advise a friend who's in a really terrible relationship that they should break up with this person, it's like more or less useless. They're not going to listen. There are certain conclusions people just have to reach themselves sometimes. And I think this might be one of those.
Speaker A: Right, right. So it's not that you need to break up with your technical skills, it's that you need to open the relationship and also consider other skills.
Speaker C: I do. Because it turns out you don't get to use the technical skills to build things unless people want them. So unless you've established people wanting it, you don't get to build it anyway. At least not if you want a thing that doesn't become shelfware.
Speaker A: Right. So the problems that you're working with now as a staff data engineer, are there system problems? Uh, there are. And that system includes a lot of people.
Speaker C: It does, absolutely. I think that's ultimately what happens here is you do something long enough that you end up working on the system and what you learn. To many people's dismay, the realization is that the system includes a lot more than just the repository or repositories themselves. It includes all of the people, it includes all of the regulations, it includes all of the corporate bureaucracy or what have you. But there are a lot of things besides the code that play into how a product gets delivered.
Speaker A: There's a big difference between writing the code and getting something done.
Speaker C: Yes, there is a huge difference between writing the code and getting something done. And I think we really do ourselves a huge disservice, um, particularly in software engineering when we try or. Yeah, in software engineering, when we measure ourselves and others on productivity metrics that we have inherited from an assembly line type of production cycle where there are five people in a line and they all are responsible for one of the jobs of running a machine that creates nails. And when you show up in the morning, there are zero nails. And then the number of nails that you have produced increases linearly until you get to the end of the day when you've created as many nails as you are going to create. I think we have this idea that you continuously do a visible thing which incrementally increases the amount of work done until you get to the end of the day. Unfortunately, um, we run into a problem when we try to conceptualize software engineering that way, because it's not production work in that way, it is knowledge work. And as a result, the way that things get produced looks very different. We don't create value by doing the same thing over and over. We create value by coalescing context from a bunch of different places into an understanding of how a problem needs to get solved right now and how to maximize. Um. The term that Kent Beck uses for this that I really like is optionality. How do we ensure that our system can move in as many of the likeliest directions as possible so that if what's going to create value changes later, we're also able to change the system? That kind of work doesn't. You don't get artifacts in a linear fashion until you get to the end of the day. And instead what you get when it's done in a way that the actual product is going to mean anything or last when you've done it that way, what it looks like from the outside is nothing, nothing, nothing. And then a big thing gets released at the end because you are spending months understanding the problem, talking to the people who are going to use the thing, figuring out what's going to work, what's not going to work. Then you are spending time socializing the changes that you are going to make. Because the bigger the change is, the more your team's understanding of the system is going to have to change. Because again, if we're in knowledge work, we don't create value by doing things as much as we create value by, uh, having this context that allows us to do the right thing. So if we change things in a way that torches a whole bunch of people's context, we've changed their power in the organization. We've taken power away from them. We've made it a lot harder for them to do their jobs. So we have to socialize these changes to them and ensure m that we minimize that context loss. Once we have done these things, which are hard things, then it is time to write the code, which I find tends to be by far the easiest step and the one that comes last.
Speaker A: And it's the only one that looks like anything, right?
Speaker C: We are treating lines of code like nails in a lot of our conception of what it means to be productive and like. We m think we can determine whether engineering work happens based on how many nails are at the end of the assembly line at the end of the day, how many lines of code have been pushed to the repo at the end of the day? And I found that it's just not true. Like, I wish it were. I know people want it to be true. I know that it feels very satisfying to see the line of code on the screen. And, uh, the problem is it's just not true. The problem is it's wrong. It's not that everybody who says that it's not like this is like airy fairy and never got good and can't write code or whatever. It's just that the entire conception of lines of code as the unit of production for a knowledge worker is incorrect.
Speaker A: Our job is not what we do, it's what we know.
Speaker C: Right? Yes, precisely.
Speaker A: People ask me, well, what's your typical day like? And I cannot explain it, but I can tell you what I know about right and that knowledge lets me say a lot of things and type a lot of things that are useful.
Speaker C: Right. And a lot of what I end up doing day to day is either possessing or finding for someone the knowledge that they need and then routing it to that person.
Speaker A: Um, because if that person is a junior engineer, people like you do a lot of work to set it up so that they can type and feel productive.
Speaker C: Right? Exactly. The thing that I never understood when I was an early career engineer was that my ability to write code had a lot less to do with how good I was as an engineer than it had to do with how good the people who were making the decisions about the repository I wrote code in were at, uh, doing their jobs.
Speaker A: So you mentioned a couple things in there. Uh, you mentioned optionality, the ability to do a lot of things that we might do. And also you mentioned that we don't do the same thing over and over. So to bring that back to operations for machine learning, what does your team do to create optionality for the rest of Mozilla?
Speaker C: Yep. For the entirety of its history, Mozilla has managed to sort of have a product with an impact that is outsized relative to the size of our team. People are surprised to hear the size of Mozilla around 750 people, um, which is fewer than people are expecting because they think of Firefox and they think of Firefox as a browser and they think about companies that make browsers and then they think of Google. And Google has enough engineers to fill the city of Chattanooga, which is not an exaggeration. That's how many engineers Google has. And so a lot of that comes from figuring out how to do the most impactful things with the relatively small, motley crew of people that we possess. And in sort of the early era of deploying machine learning models, the way that this largely worked was that individual teams around the company, um, would individually make decisions about how to get to production. Usually it was one data scientist or one machine learning engineer making those decisions. And that worked for a while, but now there's more opportunity to scale the way that we bring machine learning models to production. There's an additional complication in it for Mozilla in that we have a series of values that we want to stick to, a series of values that include not keeping data that we don't need to keep. It's a very data privacy oriented company, and using machine learning models in ethical ways. We deliver whole podcasts and stuff about how other companies are doing it wrong. So there's a great deal of pressure to ensure that we do this correctly with an eye towards data privacy, with an eye towards ethics. And those things, for better or worse, are difficult to prioritize if, uh, folks are having to expend a lot of energy on getting machine learning models to production and at all. And what we had going on a lot of the time was individuals, intrepid, laudable individuals, coordinating heroic efforts to get things into production relatively poorly. And the hope was to build sort of a production pipeline with recommendations for specific products, with licenses with specific products that would make that entire process easier. And so in the fall we launched this company that started doing this or not, not this company. In the fall we started launching this team that was responsible for doing a sprint to evaluate a bunch of different products with the ultimate hope of having a turnkey solution for machine learning engineers and data scientists at Mozilla for getting machine learning models into prod such that that wasn't the big obstacle that everybody was anticipating and, and instead became something that felt fairly routine and fairly straightforward, such that we have the additional bandwidth to devote to the design questions and sort of the bigger questions that we find ourselves wanting to approach differently than kind of the industry status quo.
Speaker A: Okay, so how did that work out?
Speaker C: Well, we've got some products in mind, and as you can imagine, there's a fair amount of coordination. The contract negotiation itself is above my pay grade. And so that was a nice thing to be able to sort of hand off. But there's still an amount of coordination involved within the organization to help people understand, first of all that we have these tools available for them and they don't need to like put things in random places on the Internet and then ping them from random APIs anymore. The second part is making sure that those things are easy to use within our system, which can be relatively difficult because we've got a variety of categories of data. And some of the most sensitive categories of data are the ones in which people need to be able to deploy these models. And so they have to be in very contained environments. And sometimes the products don't really jive with that. One of the interesting things about the machine learning operations sort of product landscape right now is that due to the fact that pushing machine learning models to production is relatively new, we're talking about less than, less than 10 years now for broad appeal. Right. A lot of these products are relatively new. They're these like bootstrapped upstart companies with 15 employees and just a lot of the things that larger clients are asking for, they can't do yet, or they don't have that ability yet, or sometimes things just don't work the way that the documentation says that they should work. And I know that that is a universal, I know it is. But it's uh, especially a problem with these sort of products in my experience. And so I have found that the, as a result, I have found that the characteristics that it's really important to evaluate some of these products on are a little bit different than one might initially imagine. For example, a thing that I think engineers have a tendency to do is over optimize on too many metrics at once. They have these metrics that they want to judge products on and they uh, they don't necessarily consider the difference between an optimizing metric, which is, there is no limit to how good this could be and it would be a benefit to us. You always want fast available, maybe potentially, but sometimes that's. Sometimes. Sometimes an optimizing metric would be speed. For example, in programming language design, I think speed is an important optimizing metric because the faster your programming language is, the more mistakes somebody writing it can make and still have the code be fast enough. And so if a programming language isn't fast enough, then people won't end up using it. And for that reason that does need to be an optimizing metric. But not everywhere. A lot of times in end user application development, we focus on keeping the code legible and understandable over optimizing for speed. Particularly in cases where what we're doing is we're iterating over a relatively small collection of things. If we're iterating over 14 keys and values, I would rather that be done in a slightly slower and legible way than in some kind of obscure way that's a little bit faster. Because in that case my optimizing metric is different because I'm working with a team of people who is building this product that people are going to look at a screen and use. And I have found that with end user application development, at least where people are looking at screens, your thing doesn't have to be faster than people are going to be able to visually observe the change.
Speaker A: Okay, so there's such a thing as fast enough, beyond which you don't care that much, Right?
Speaker C: In certain contexts, I think different metrics can be the optimizing metric, uh, can be an optimizing metric depending on your context. But what we tend to do as engineers is treat everything as an optimizing metric. Everything needs to be better all the time. And, and we put ourselves in these decision deadlocks about building products and using products when really the way to make the decision that would be the best fit for us is figure out what the minimum number of optimizing metrics we really need is, what is most important to us, what has to be really good.
Speaker A: So what is most important to you in this decision? I need some concrete examples.
Speaker C: Um, I will get there. Um, but the other thing that we need to do is then treat some of the metrics instead as satisficing metrics. So you end up in a situation where you have a few optimizing metrics, things that um, you would always want more of them if it were available. And then satisficing metrics for which you have a threshold that's good enough and better than that doesn't really make that much of a difference to you. And I find that decision making becomes a lot easier when I stay very mindful of decision deadlock around optimizing metrics. Um, but treating too many things as optimizing metrics and then just understand what's going to be good enough for most things and then what areas additional improvement would always be helpful. So for example, when evaluating some of these products, one of the things, uh, I think as an example, um, when people are evaluating operationalization products, things that feel really important are speed. Things that feel really important are full featuredness as proxy by asking a sales representative whether the features are available and or looking at the demo. I think something that people really focus on is, I think a thing that they consider certainly is cost at various levels. I think something they think about is GPU availability. Um, and I have found that uh, although those metrics can be useful depending on your organization, depending on the organization's scale, there are satisficing metrics for things like scale. Typically if you're talking about a relatively small startup, there are just maybe not situations in which you need to worry about what happens when 1 billion users log onto your system at once because it's not going to happen for a
Speaker A: long enough time problems you want to have.
Speaker C: Right? Exactly. And m, when we are refusing to choose deployment solutions because we are stuck in decision deadlock around insufficient scale based on these, let's call them, aspirational expectations, then, uh, we do ourselves a disservice. One of the things that I have found, particularly with a lot of these machine learning operational solutions, is that one optimizing metric that tends to not even make it into the, into the decision grid for a lot of organizations is the availability and flexibility of support. Support, yes. If it is a product where they redirect you on the website in the contact form in like three different locations and then they try to give you an automated chat thing and then you got to email them and then there's like a week turnaround or whatever that really big companies that are making universal products, I won't name names but you can probably guess they can get away with that. These companies that are fighting for market share in the ability to help people get their machine learning models to prod, they have stuff breaking all the time, they have stuff not uh, not jiving with the docks all the time. They have a lot of custom situations that people are running into that they didn't necessarily think of. And so it's really, really important for them to have very responsive, very responsive and very involved support teams in my experience because ultimately once the contract is signed, that's chiefly who engineers of the client company are going to be interacting with. And we're going to be interacting with them a lot. At least in my experience that has been the case where it is a multiple times weekly thing where a team needs something and I'm going to m, for example the weights and biases team. And I'm saying, all right, somebody wants to make a custom chart that looks like this. They've tried this and that, but they're running into this problem. What do you see? And have been able to get a hold of, for example the support engineer on Slack who will take the code, put it in one of their own ways and biases projects and try to figure out why the chart's not doing the thing.
Speaker A: Hooray.
Speaker C: Yeah, it's lovely. And it was an optimizing metric for me going into choosing these products and I had to fight pretty hard for it. But it's ended up being really, really valuable because the products we chose that have really um, responsive support teams are the products that we have managed to get into production at this point. And the ones that don't have that, the ones that meet. To be honest, a lot of our uh, most deeply held ideological ideas about what it is that we want. These like free open source things where everybody just sort of contributes. The idea is that it's going to be this community maintained free and open source software. Those, I mean they don't have dedicated support teams because they're these open source projects. And that's wonderful for lots of reasons. But what it's not is a quick way to get to prod and so Some of those decisions, um, and I won't even call them decisions, some of those potential paths forward ended up not being viable despite the fact that we love the code because, wow.
Speaker A: It's almost like the code isn't the whole thing. Right.
Speaker C: The coordination part isn't there. The assistance part isn't there. The documentation part isn't there. The support is not available. There's not somebody whose job it is to ensure that this is working for the people who need to use it. And it turns out that that job, that support engineer job, is like the most important job for a lot of these products.
Speaker A: Yeah, yeah. At Honeycomb, we're a small enough company that we put a ton of effort into support, and our, uh, support is really good. And you're right, it doesn't show up on most people's evaluation sheets, but it sure does afterward.
Speaker C: Right. Once they've put down all of the money and all of their bargaining leverage is completely gone, at that point they realize that support is a really important thing that they probably should have checked on before they signed this.
Speaker A: Yep, yep. It gets us renewals.
Speaker C: Yeah, it's great for retention, which is nice.
Speaker A: Okay, Okay. I have to ask you, you've been talking about buying a lot of ML ops things. How is this different from regular DevOps?
Speaker C: That's a good question. My tenure in the industry didn't go through DevOps, so I'd be interested to hear how it sounds like it differs to you. My path was that I started as a machine or I started as a software engineer. As a software engineer for years. Then I switched over into data science. That kind of became machine learning engineering. Then I was back to software engineering for a while. Then I was freelancing and kind of doing all three. And then I went to Mozilla. I was working there as a software, um, engineer for machine learning teams. And that sort of became this Data engineer is an interesting title because I think it means to different companies, different things. But most broadly, what I have found it to mean is this is a person we expect to be able to like, fling into the model or fling into the data science code, or fling into the app code and expect them to be able to do what needs to be done. And my understanding of DevOps is that this is more focused specifically on taking, uh, as an example, end user applications and getting them into production, and getting them into production in a way where they're going to stay up and functional through the various circumstances that are going to befall it.
Speaker A: Once it arrives, you can see what's going on.
Speaker C: Right. And I think. I don't know that it is necessarily different in kind.
Speaker A: Well, it's got to be different or there wouldn't be all these brand new startups helping you do it.
Speaker C: Well, there are special considerations I think for machine learning models that are different than the considerations for what we might call deterministic code systems or at least if they're working. Deterministic?
Speaker A: Oh yeah, yeah. It's the new ICE car. Right. The internal combustion engine. Yeah. It's now a special category in vehicle development. So now we have. What did you call it? Deterministic program execution. Yes, that's regular DevOps for us now.
Speaker C: Yeah, I think. Right, exactly. If the code is receiving the same input and giving you different outputs every time, that would be concerning to a DevOps person. That would be not necessarily concerning yet to an mlops person. However, that puts a different decision tree in place because there are a lot of reasons that it could be changing and some of them are fine and uh, some of them are awful. One of the things that I think I'm particularly aware of in a role as a machine learning operations person is that once we are putting these models in producing, it can be a lot harder to diagnose why a machine learning model is malfunctioning. Because many of the, because the particular nature of the internals has been automated in such a way that we don't necessarily have an exact answer for why it categorized this as this thing or why it uh, predicted that thing. It's been interesting to watch people toy with ChatGPT because for, for a lot of people this is their first experience that they know of interacting with a non deterministic system like this and they don't understand how ChatGPT works, uh, which is fine, but then they will be upset because ChatGPT will give them a response that they know to be not factually correct. The most memorable example that I have is a conversation that I had with someone over Discord who was just distraught Jess, because they had asked ChatGPT for some code and um, and then they had asked ChatGPT to run the code on some input and GPT said I ran the code on your input and the output is blah and, and the output was not what the answer would have been had the code been correct. It was made right and the person. And I was. And they asked in the discord, why would chatgpt, um, why would chatgpt say that it ran code that was correct when the output demonstrates uh, an incorrect response based on Correct code. And I misunderstood the question to be like, can you explain to me the mechanism by which ChatGPT operates that would result in this outcome happening? And so I tried. I think you actually witnessed this. I tried for far too long to attempt to explain the mechanism. But what this person was asking wasn't, how does Jack work? What they were asking was, how could ChatGPT do this? To me, these are very different questions.
Speaker A: This is a question of justice.
Speaker C: Right. This is not about. And finally, the person was like, why is it so pie in the sky of me to imagine a world in which this thing has actually has an interpreter in it, that it's actually running on this code, that it actually ran in this specific circumstance? And I was like, you can imagine that world all you want, it's just not the one we're in.
Speaker A: Uh, Yeah, I mean, ChatGPT probably does have access to an interpreter at this point. I don't think it did then. But.
Speaker C: And even if it did, it certainly wasn't the case that it had run it on this code because the input and the output demonstrated that that hadn't happened.
Speaker A: Yeah, yeah. ChatGPT is not a deterministic program execution, and you do not get to decide what reasoning it uses.
Speaker C: Right. And so right there I think that the. And so I think that the decision tree for understanding how these things work and what you're going to get out of them is different enough for machine learning Ops relative to DevOps, that a background in software engineering by itself is insufficient to be able to debug and diagnose some of the issues with these things. And the products that are designed for operationalizing deterministic code don't possess some of the features that are really important in order to be able to do that as well.
Speaker A: Like what?
Speaker C: Well, in machine learning, one of the things that we want to be able to check is if we have built a set of training data and a set of test data. And the way that we have built this machine learning model is that we have used this training data to create this model, and we've used this test data to ensure that the model is going to operate with, um, whatever metric we are using, some metric of accuracy that's important to us. Right. And then we put it in production. One of the things that's really important to check on periodically is that the data in production continues to match the test data that you tested the model on in all of the ways that are important for ensuring that the model is going to give accurate predictions.
Speaker A: So, uh, so wait Wait, are you saying you need to test the input in production?
Speaker C: Sometimes, yeah. Yes. And uh, that way you know that if your model is starting to give, if you're, if you're starting to see suboptimal results of using a machine learning model in production, sometimes the reason is that the model was trained for a situation that does not match, often inadvertently, the situation that we're running into in production. The canonical example that Andrew Ng talks about in machine learning, Yearning is this model that is designed, I think it was to like identify cats in photos or something like that. And they gave them um, all of these images of cats, high quality images of cats. Right. And then they put it in production. And the actual cat pictures that people. It was a fake example. I think it was a contrived example. But if you put that in production and then people are offering you their grainy, blurry cell phone pictures of cats, then a model that did very well on these very high quality images of cats is not going to do as well. Because the data that you are getting from production doesn't match the data that you were testing on when you put the model in production in the first place. And there are ways that this can happen subtly over time. For example, I work on a system at Molazilla that is responsible for sanitizing search data, by which I mean we are very, very careful about saving search data. And uh, it's a balance that Mozilla is often trying to strike between sort of sustainability for making our products better and not storing people's data. So one of the ways that we do this is that we have a system that is designed to detect if somebody has put personal information in the search box, and if they have put personal information in the search box that goes like to the incinerator, more or less. We do not keep that data. We attempt to keep the searches that are like looking for a specific product, looking for information without any personal data in them. And then we attempt to get rid of the ones that contain that personal data. And it's pretty conservative by which I mean we throw away a lot of stuff that is not personal because it contains like a character that is often in personal data. A, ah, good example of this is numerals. Anything with numerals in it. We don't, we don't, we don't keep, with the exception of like an allow list that has been specifically created by a person.
Speaker A: Like one.
Speaker C: Yeah, yeah, exactly.
Speaker A: So if I searched for like number one camera, you could keep that. But if I searched for anything that might be a phone number you would not keep.
Speaker C: That depends. If the literal exact phrase number one camera is in the allow list, then we will keep.
Speaker A: Oh, wow.
Speaker C: Otherwise we will not.
Speaker A: Wow.
Speaker B: Yeah.
Speaker C: Uh, but it contains some automated pieces. For example, we use a named entity Recognizer to determine if somebody has put a name in the search box. And as you can imagine, if there's a name, uh, if there's a name according to this classification, then that gets classified as personal data and we don't keep it. But the named Entity Recognizer, we use spaces named Entity Recognizer. And one of the things that we want to be very careful about is ensuring that folks who are using this feature in Firefox are, um, in the population that was used as the training and test population for the named entity recognizer. Now spacy tends to do so is
Speaker A: that like, what if people are searching for names from, I don't know, Russia or Madagascar or someplace where those nationalities of names weren't in the training data?
Speaker C: You got it. That's exactly what we would be concerned about. So what we start with is, um, using the product in places where we know training data came from for these models. Then as we begin, if we were hypothetically then to want to use that product in a wider variety of places, a thing that we would need to make sure we knew was the case was that our named Entity Recognizer was still going to work. So we know what languages the named Entity Recognizer sort of performs the most sensitively on. We're optimizing for sensitivity there, finding the names. And, uh, we are looking for our search data to contain, uh, the largest proportions of those languages. We are monitoring the distribution of language proportions of our search data over time. We're not storing the search data itself, but we're storing this aggregate metric about the search data in order to ensure that the data that this model is being exposed to is data that the model will be able to interpret correctly. According to our tests of the model, if that input data were to change, we want to know about it before the model is able to start storing stuff for a long period of time. I think one of the things that is similar about machine learning operations and DevOps really is that a lot of our role is in increasing the catchability of problems. There are sort of three big risk amplifiers that I think about when I'm trying to determine how bad a risk is. One is catastrophicness. How bad are the consequences if this happens? One is likelihood. How likely is it that this thing is going to happen? But the Third, and I think the most overlooked one in a lot of discussions of risk is insidiousness. How likely is it that this thing goes uncaught if it happens? And I think security engineers, DevOps and MLOps are all focused perhaps more heavily than, um, some of the other roles within our profession on that insidiousness, risk, and ensuring that if something is going to go wrong, we will be able to identify it and track it so that we can figure out when it happened, figure out where it happened, and figure out how to fix it.
Speaker A: Nice. Reduce insidiousness. That is an objective. There's some differences with machine learning and generative AI and all of this exciting, uh, new stuff versus deterministic program execution. When should you use each?
Speaker C: That's a good question. I think that I have a particular perspective on this.
Speaker A: I want your perspective, Chelsea.
Speaker C: Okay. So my perspective is that when we think about automating things, we tend to imagine some, like, generalized system doing all of the steps for us. I think the most operable and most effective systems are going to look pretty different from that and are instead going to be a series of steps with humans in the loop that allow us to get to where we're trying to get to. The example that I think about a lot is, um, who to follow models on social media. It's one example that I end up using a lot because it's something that feels familiar to enough folks. Uh, we have this sort of feed on social media on something like Twitter, right. And there are a number of different ways that you can recommend who else people follow. Social media systems survive by their ability to get folks connected into each other because that will allow people to keep checking them. That will create the legitimacy for the platform for people to decide that they need more followers on those platforms. These sorts of things drive up engagement. Usually on these systems, the people using the people, the constituents of the system are largely the product, and the way that they're staying sustainable is through advertising. So these usage metrics are really important. And a big part of that is making sure that people can find folks they want to hear from in order to follow them and how to do a, uh, who to follow model. There are a lot of different ways that you can do this. Twitter has published a variety of papers about how they did it. Twitter has also published a variety of papers about how the ways that they suggested in the original papers were, it turns out, not that great. And, uh, it's an interesting academic paper trail to follow. But when organizations think about this, I think they Imagine like, oh, I'm going to make popularity metrics and people who are. We're going to recommend people who a lot of people follow because those seem like very followable people. And what they end up with is the Beyonce problem. Like, what they're doing in these situations is they're taking somebody who's already popular and sort of spiraling that popularity up, which is a way to do it. It's sort of a critical mass function. But what it's not necessarily doing is finding is connecting niche appeal folks to other niche appeal folks in a way that's actually going to keep them online. Because it turns out that even if a machine can't tell the difference between somebody who has figured out how to game the engagement system, a lot of times a person can. And that's a lot of the reason that people hated Twitter by the end is that it was a lot of like, or X or whatever we're calling it now. Right. Is that there are these very clear optimization mechanisms for getting recommended and they turn out to not be things that people. It's not actually people want to look at this thing. It's very, very hard to take a single step and use that single step as a proxy for human judgment. It turns out human judgment is very hard to automate. Part of the reason for that is that human judgment is a relatively varied thing. The other reason is that it's very hard to do better. This is maybe a cancelable take. It's very hard to do better than human baseline with a machine learning model. And I would say human baseline on judgment, spotty at best. Right. So, um, but I think a way
Speaker A: better would be more consistent.
Speaker C: Right, right. And so I, uh, I think about a problem like that and I think a more an approach that would better approximate what people are actually looking for on these platforms would be to break it down and determine, um, for example, to start with, like, what topics are people talking about about. And a topic recognizer can be useful. You can even start with hashtags. Like, there are a million occasions when folks on Twitter have created features that made Twitter valuable. Follow Fridays were a feature of Twitter. Those were created by users. Hashtags are a feature of Twitter. Those were created by users. Assuming that we don't, that we don't think of a feature as exclusively a good thing. Twitter main character was a feature that was created by users. Twitter engineers didn't do any of that. It was all people who were using, uh, the platform. And so I think that there's something to be said for understanding how people are trying to engage with a thing, breaking it down into steps, and then building, a lot of times classical machine learning models based on tabular data in order to answer those sorts of questions. So an example here would be figure out what topics people are talking about. Figure out which people are knowledgeable on those topics. There's probably a human in the loop there. Then figure out how to recommend the folks who are knowledgeable on these topics to the people who are interested in learning about these topics. That's three steps. Each of them is much more simple and you end up with something that better approximates what at least some people are trying to do online than using a popularity model that recommends somebody, a celebrity talking about COVID because, uh, that person has X x million followers and this person's interested in learning about COVID that's not a match. But if you did a three part more simple system or a three part system in which each component were more simple, then you would be more likely to get a match there.
Speaker A: I think I hear you recommending for any problem that you think you might use generative AI or machine learning for. Uh, break it down and you might find some parts where yes, a model is useful for what topic is this message about? Um, and some places where a deterministic. Hey, I hear you talking about this topic and them talking about this topic and I can say, uh, deterministically, I might want to connect you, um, and other places where there's a human involved, um, that are low volume and high value enough to ask, to ask a real person. M okay, okay, Chelsea, this has been amazing. Chelsea, where can people find more?
Speaker C: Oh, I do a lot of writing@chelseatroy.com. that's my blog where I talk a lot about this stuff. I'm also. Hey, Chelsea. Troy. H E Y Chelsea, like the neighborhood in New York. Troy, like Helen of Troy on most social media platforms.
Speaker A: Great. Find those links in the show notes over@re arresteddevops.com ML Ops. Find Arrested DevOps on various podcast systems that I don't even know about and leave us a review if you want to help other people find us or just be obnoxious, but give us like a high star review and then be obnoxious in it. Okay, we'll find that amusing. Thank you for listening. I'm Jessica Kerr at ah, Jessatron. This has been arrested DevOps in the banana stand.
Speaker C: Oh, sorry, immediately.
Speaker A: No, we're there, we're there.
Speaker C: There's always DevOps in the banana stand.
Speaker A: Uh, yes, good. I like that one. Okay, we're keeping that, um.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.