Engineering Founders · 2026-07-24 · 35 min
Fireworks AI, founded by former Meta and PyTorch engineers including Benny Chen, has built a focused infrastructure platform for training and serving open-source generative AI models. The company initially pivoted from recommendation systems to generative AI when ChatGPT launched, and further narrowed focus to text models after recognizing that image generation models were predominantly closed-source - a deliberate choice to stay aligned with their open-source thesis. Customers like Cursor and Genspark use Fireworks to customize models through reinforcement learning, leveraging their proprietary user data and PM judgment to create high-quality, application-specific models. Chen emphasizes that competitive advantage increasingly comes from data flywheels - the continuous loop of collecting user request data, having teams judge outputs, and retraining models. This mirrors recommendation systems more than traditional ML. On hiring, Chen notes the criteria have fundamentally shifted: rather than seeking candidates with specific prior experience (the old hardware model of finding someone who worked on Bluetooth at a specific company), Fireworks now prioritizes high agency, quality consciousness, and change-management experience. AI coding models have made acquiring domain skills on-the-job feasible; what matters is hiring people who ask good questions and maintain uncompromising standards for infrastructure and testing so AI agents can operate effectively.
A data flywheel occurs when application companies collect user requests and traces, have PMs and engineers judge whether model outputs are good or bad in their specific context, then feed that judgment back into reinforcement learning to continuously improve customized models. Over time, this creates a virtuous cycle where the model learns user preferences and application-specific behavior.
Most image generation models are closed-source, which conflicts with Fireworks' core thesis of supporting open-source models. While image generation is profitable for some companies, Fireworks chose to stay focused on open-source text models where they can build stronger customer relationships around data flywheels and model customization.
High agency (ability to learn on the job by asking good questions of AI models), strong quality consciousness (refusing to ship bad code to production), and change-management experience are now more important than domain-specific skills, because AI coding models can help engineers acquire technical knowledge if they know how to use them effectively.
Companies use user request data, agent trajectories from those requests, and most importantly, judgment signals from PMs and engineers about whether model outputs are good or bad in their application context. These judgments train the reinforcement learning loop to bias the model toward desired behavior.
Fireworks provides infrastructure to scale training workloads across multiple regions and data centers, enabling companies to run reinforcement learning on their proprietary data at scale. This allows them to customize base open-source models (like Llama) into application-specific models optimized for coding (Cursor) or search/slide generation (Genspark).
Computed from the transcript - who did the talking, and the words that came up most.
Benny Chen, Co-Founder@ Fireworks AI, joins the EF pod to discuss his founder journey and share valuable insights on navigating common founder / product dev challenges in today’s agent-first landscape. He and Jerry cover strategies for creating effective messaging, staying competitive in a crowded market space, hiring top-tier talent / what qualities to look for in high-performing engineers, navigating the cultural shift to managing agents, creating data flywheels & how this can help your customers, and more. ABOUT BENNY CHEN As co-founder and early product architect, Benny Chen shaped Fireworks AI’s infrastructure strategy, spearheading the design of scalable systems to support high-throughput AI model serving. Benny’s contributions established the technical foundation for Fireworks AI’s robust and cloud-native architecture, which underpins its ability to meet enterprise demands. Formerly Meta’s Ads Infrastructure Lead, Benny optimized large-scale ad-serving pipelines and developed significant expertise in distributed systems and cloud infrastructure. He holds a B.S.
Transcribed and scored by The B2B Podcast Index.
Speaker A: This episode is brought to you by Sidero. If your team runs Kubernetes, chances are upgrades can feel risky. That's what the Telos platform is for. It's built on immutable minimal OS designed for Kubernetes. Immutable means clusters cannot drift so they stay identical and upgrades are a non event. Minimal means there is almost Nothing to attack. 50 binaries, no shell, no ssh. So it's secure by default. Upgrades get boring, security gets boring. And that's exactly the point. Check it out@siderealabs.com that's S I D
Speaker B: E R O labs.com at the end
Speaker C: of the day you're hiring people to do certain things and then acquiring the skill early on is very, very difficult in hardware. I think the joke was that like you just have to find a person who worked on Bluetooth for this manufacturer because that's what that person worked on for his whole life. No one else have the training data for it. As the coding model got better, I think it's more and more important to hire people who have high agency and can figure out all the things they need to learn on the job. You don't need to learn all those things before you come into the job. It's more important to figure out what the right question to ask so you can extract all those information out of the model. I emphasize much more on people who worked on change management and people who have a high quality bar for their work to make sure that we can set up all the right parameters for the coding agent to work in.
Speaker D: Welcome to Engineering Founders, the show for engineering leaders making the daring leap to start their own company.
Speaker B: Hey Benny, thanks for spending time with us today to share your journey building Fireworks AI. Uh, really excited to learn more about your experience and transitioning to entrepreneur the journey and there are a lot of things we talked about last time, so I'm really uh, eager to get into it. To start, can you share who you are in your journey building the company?
Speaker C: Absolutely. I'm Benny Chen, I'm one of the co founders at Ah Fireworks. So we've been working on large language model training inference for about four years. Before this I was at Meta working on ADS model, serving with GPU Asics and how we got started is one. I was working on capacity planning for ADS infrastructure and all the demand for AI infrastructure inside Meta was going through the roof so. So it was very clear to me that yeah, infrastructure is a important area for me to play in. And on top of that there were other people thinking about starting the company around the same time. So I hopped on the train and joined them.
Speaker B: How did it take from an early idea into a. Like a rocket ship now?
Speaker C: Yeah. Most of the early members on the founding team were Pytorch veterans. So Dima was uh, a core contributor. Lin was the manager on the team or originally we knew that we want to work on the AI infrastructure space. It wasn't clear to us whether the area to work in was recommendation system or generative AI, so we started working in recommendation system first. As ChatGPT came out and as generative AI workloads uh, started to take off, we quickly pivoted to work on generative AI. I would say the demand is definitely much higher than anything I could have anticipated. At the same time I think we're in a good position to make a lot of impact.
Speaker B: Tell us more about as a founder in the early days, how does messaging matters? Because this is something that when you and I chatted earlier that you shared this particular point about you can't outsource messaging to a very experienced go to market leader.
Speaker C: Yeah, I think that's a good question. It's not so much that uh, uh, you can't outsource. The stories are ideally told by the people who are working on the product. Articulating an idea often is a multi fold, multi pronged process. And then the people who work on the product have the full story sometimes in their head. And communicating that clearly takes a lot of work. It is easier to tell that story by yourself. It is not scalable and ideally it is something that tells the messaging can be scaled to many, many other people. I would say like for uh, fireworks. Early on we knew AI, ah, infrastructure was going to take off. That was very clear to us and that was unequivocal how it would take off. I think that's still up in the air. So directionally we knew that we want to work in AI infrastructure. We could work on AI infrastructure on recommendation system or generative AI. And uh, inside generative AI we also have text and image and audio models. So it wasn't so much that we had to pivot in the middle. We are uh, very focused on open source and we are very focused on AI infrastructure and that were the two main focus of the company. At the end of the day it turns uh, out text models, open source models were the area for us to play in. So we were able to focus the company on these efforts. But I think the transition is pretty natural because we are big believers in open source and we want to focus
Speaker B: on infrastructure that ties to the prior experience of the founding team. You mentioned Pytorch.
Speaker C: Yeah, Pytorch. Uh, most of the founding members were Pytorch and Pytorch's JSON teams. And for us it was clear that open source will play a huge role in generative AI, whether at the modeling layer or the infrastructure layer or the layers below. We do see that open source models iterate much faster than closed source models because there are many, many people exchanging ideas in the open. So supporting open source model was a uh, big, big hypothesis for us. And I think three or four years ago people were still wondering whether training or inference was going to be the main focus. We were very clear about inference from day one. We always look at our prior experience in supporting Meta's AI infrastructure and most of the spend goes to inference. Training is definitely very very important but training models is what you would do before you do inference. So it's an iteration cycle. Making sure that we focused on inference early on was something that's contrarian early on. I don't know why for the industry but uh, we were very focused on inference and then as we ramped up our inference efforts, the whole reinforcement learning wave came. Reinforcement learning is all about having really good inference. So we were able to contribute to the training area as well.
Speaker B: It feels like your team are focusing on the right thing and you're getting ready to take the big wave of uh, the market momentum just in time when it's ready.
Speaker C: We honestly had a pretty simple thesis. This area though, it's very execution heavy. The execution quality is very, very important. We make a lot of mistakes still every day and we try our best to avoid those mistakes. We want to make sure we deliver the best infrastructure for our customers.
Speaker B: Can you share how that played out? Like you have a unique insight that are not agreed by the mainstream which give you the advantage to start early. How does that early insight carry the company's messaging over time since we're talking about that topic?
Speaker C: So initially it was mostly about serving the open source model as is at the very beginning we were serving llama uh, models and those models are pretty hard to use and the usage surface area was also relatively small. So a lot of initial workloads uh, are sort of AI assisted workload. As the industry matured m we started to see one, the open source models are being more sophisticated to the areas of which the AI models is deployed is also more mission critical. So, so then model customization became more and more important. We want to help people customize models so they can optimize the model for their application. And so initially it was more geared towards serving the model as is. And then as time went on, the model customization and the training piece became more and more important. It's less about whether the capacity or the revenue of the company is coming from training itself. It's more that the customization of open source models are becoming more and more important in the industry. And I believe it will be very important in the future.
Speaker B: Can you share a bit more about what kind of customization that becomes more popular?
Speaker C: Yeah, uh, recently we went public with Cursor on their, like composer two and um, 2.5 training. So cursor, for example, is a company that has great proprietary data on coding. They want to be able to scale up their training aggressively and we help them scale up their training workload onto multiple regions to be able to do reinforcement learning across many small data centers instead of one giant data center. Uh, there are also many other companies we went public with like genspark, like Vercel, where we help them customize models so then they can use our infrastructure for training and then later on use the same infrastructure for inference. All these application companies have really good data and really good PMs who can describe what is good and what is bad. Those information can be all baked into the model to help the model understand how to behave in certain conditions so that their customers can derive more value out of the model.
Speaker B: So every company is different, their use case is different, the way they want to train the model is different. So there's uh, almost like infinite variations of use cases. So that's the demand.
Speaker C: Absolutely. At the end of the day I think data is where you have the moat. And all these application companies are in a great position to capture all the customer intent and turn those into really cost efficient and high quality models.
Speaker B: And going back to our earlier question about messaging, so what uh, is the messaging that you have landed for now?
Speaker C: Yeah, we want to train and serve open source models. That's sort of the core thesis of the company and all our messaging revolves around that. At the end of the day you want to come to fireworks to be able to customize models and serve those customized models. And that uh, sort of evolved over time. From early on we were just serving the open source model as is now.
Speaker B: You have now learned the messaging, you have a unique insights, you gain the momentum and benefit from the, the big shift. There's also a lot of competition in the space. There are a lot of money putting into the AI industry. So how do you stay competitive and be a winner? So in this very competitive market, yeah,
Speaker C: I think staying focused is very, very important. The whole field is very execution quality heavy and we are very focused on making sure that we deliver uh, the best inference and training stack to our customer. A lot of companies in the area are often distracted about different things and chasing different shiny objects. We just want to train ourselves open source models.
Speaker B: Do you have an example of walking away from a, uh, shiny object? I'm sure there's a lot of them along the way.
Speaker C: I think one good example was image generation. So we used to be a big proponent for image generation as well. It turns out most image generation models are not open source. So it is very different from our thesis where we want to stay focused on serving and training open source models. So yeah, we slowly diverted away from it as all the models in the uh, field become close source. It was very tempting to think about, hey, like how do we figure out different forms of partnerships or setup where we can serve these proprietary models as well. But at the end of the day, if we can convince ourselves that we can help our customers set up the data flywheel. I don't think it was as interesting as the uh, open source text models. Maybe it will come back as the text model gets the capability to generate images as well, but it may not come back anytime soon. And that's something we had to walk away from because sort of like the ecosystem around open source models died out.
Speaker B: Yeah, it felt uh, like you have your team have early and have the core conviction about open source, how open source are going to be really critical for AI infrastructure. So that conviction guided the team away from the distractions.
Speaker C: So that's, yeah, uh, distraction or not, maybe like they're definitely companies in the area that makes a lot of money on image generation, that's for sure. I think for a small company you have to stay focused. So there are definitely areas where we have to make certain sacrifices to stay focused.
Speaker A: This episode is brought to you by Sidero. If your team runs Kubernetes, chances are upgrades can feel risky. One CVE patch can put a whole fleet in doubt. What should be routine takes too much time and it often ends with your best engineers fighting fires. Instead of shaping a roadmap, the root cause is underneath a uh, general purpose operating system that was never built for this. The moment someone runs a package manager or hotfixes a box at 2am you have a drifting node. And worse, the OS comes with a large Attack surface that you never needed. That's what the Telos platform is for. It's built on immutable minimal OS designed for Kubernetes. Immutable means clusters cannot drift, so they stay identical. And upgrades are a non event minimum. Means there is almost Nothing to attack. 50 binaries, no shell, no ssh, so it's secure by default. It handles the fleet's lifecycle for you. Provisioning, upgrading and retiring machines automatically. That design matters most when the stakes are highest, including the edge with no one on site and AI clusters where drift is expensive. Upgrades gets boring, security gets boring. And that's exactly the point. Check it out@ah sidearlab.com that's S I D E R O Labs com.
Speaker B: You mentioned about customer data flywheel. What do you mean by customer data flywheel?
Speaker C: Uh, every application have unique insights into how to use their data. They will collect customer traces about how to, you know, vibe code a webpage or come up with a design or I don't know, generate some 3D assets. The nuance is that how to evaluate whether the output is good or bad and whether the output is good or bad in the context of the application itself. All those work that the PMs and the engineering people on those application team do to improve their product can eventually be articulated and translated into language models judge so that the model can learn the same thing as well. So all these information can flow into the model and all the customer uh, usage on the application can flow into the model so that they can have really really good model and also derive really good ROI from those models. I do believe in the long run at scale, all these generative AI models will look more and more like recommendation systems where we learn the preference uh, for different people or how they use different applications and the model can respond accordingly. So yeah, that's very important for us because we had very strong belief on data flywheel and model customization and we help all the application developers in the space who want to train models to use our platform to customize the model and serve them in production.
Speaker B: Can you give us one example of a customer that are building the data flywheel through Fireworks AI and what is the value they deliver back to their customers or users?
Speaker C: Yeah, I think Composer, uh, the recent Composer launch is definitely a uh, very good example. Like I love the model, the model is great, they work a lot. Customizing the model to the cursor harness itself so that you're able to train a really good coding model. Genspark is another example where we went Public with that customized the model for different search and slide generation functionalities in the application so that they're able to get really good slides and uh, really good search results for their users. All these customers we work with are ah, big proponents in setting up the data pipeline, setting up all the customer preferences they know about as a training loop and then um, putting those information into the model itself. And we believe there are more and more people like this who have the appetite to customer customize models and have the conviction to deliver the best model for their customers.
Speaker B: Got it. As you said, AI models in the long run will more become a uh, recommendation system so that what you are doing is helping your customers to take their consumer or customer data ah, so that they learn more about what their customer needs so that they deliver a much more nuanced context driven or personalized experience back to their end users.
Speaker C: Absolutely, absolutely.
Speaker B: And that's the trend for all the application to be right in order to win this market?
Speaker C: I think so, yeah. As the spend on gen AI goes up I do believe more and more application companies will lean towards customized models because they simply don't have the balance sheet versus these frontier labs. They cannot afford to burn as much money so they need to figure out how to customize models and they need to figure out uh, how to beat the frontier labs in areas where uh, they have have a moat.
Speaker B: The question about what kind of data company leverages to train their own customized model. So you mentioned earlier the product usage data. What are the common type of data company collect and use to train their model?
Speaker C: So oftentimes it comes in the form of user request and the agent trajectory from those requests. And then what's important for reinforcement learning is that the application developers on Those teams, the PMs on those team can clearly articulate those requirements as language models are judged. So then we can use the language uh, models to judge the outcome of the model in the reinforcement learning loop and then sort of bias the model towards what the training process wants the model to do. It's like initial user request and then the final language model judge and everything else in the middle is generated on the fly through uh, reinforcement learning.
Speaker B: Got it. In the case of Jenspark, you mentioned slide generation, they provide input in terms of judging whether a slide generated by their application is good or not.
Speaker C: Absolutely, yeah.
Speaker B: So now assuming customer has built the data flywheel, what are the next thing they do is that ongoing uh, basis with fireworks AI or they will unlock more things that you can help them with?
Speaker C: Yeah, I think ongoing is Definitely happening a lot. Especially when you have new open source model release. You need to sort of retrain the whole setup with the new uh, base model. You also derive more insights from the user using these applications and then for reinforcement. Often it's like a whack a mole game because the model is often lazy and will try to figure out what's the easiest thing to do to fit your requirements. You need to often add more data and more judgment to the process so that the model quality can be better at the end of the day. So it is a repeated process.
Speaker B: And typically what team do you work
Speaker C: with at those customers, companies, the application developers or the MLEs, uh, in those companies.
Speaker B: Got it. So going back to your own building, your own company competition not only means product level competition, messaging level competition, there also means talent competition. So in the new age of AI, tell us how you see the bar of hiring, uh, becomes for engineering talent.
Speaker C: I think it's different. We used to at least uh, when I was a meta, focus on whether this person has worked on certain things before. Because at the end of the day you're hiring people to do certain things and then acquiring the skill early on is very, very difficult. I think a good example is in hardware. I think the joke was that like you just have to find a person who worked on Bluetooth for this manufacturer because that's what that person worked on for his whole life. No one else have the training data for it. As the coding model got better, I think it's more and more important to hire people who have high agency and can figure out all the things they need to learn on the job. You don't need learn all those things before you come into the job. It's more important to figure out what the right question to ask so you can extract all those information out of the model. What is important has also been changing. I emphasize much more on people who worked on change management and people who have a high quality bar for their work to make sure that we can set up all the right parameters for the coding agent to work in, set up all the tests and all the release process or the design thinking so that the agent can follow and deliver really good infrastructure for our uh, customers. The actual skill is I think less
Speaker B: important because those can be acquired on the fly through AI. And how do you test those qualities?
Speaker C: Honestly, I think it's very, very hard. So we try our best to go through uh, cultural round discussions with candidates. We try to understand if they go above and beyond to fix something previously in their old job. But I do think this is where, you know, at certain point, the joke I always like to make is that the monkey brain is more useful than the prefrontal cortex. I don't know how much you can analyze to understand whether this person has high quality bar or this person has high agency. I do think this is where it's more nuanced. And oftentimes you just have a conversation with this person for 45 minutes and try to pick up those, uh, signals, which sometimes is quite sparse.
Speaker B: If you look at the best performing engineers you hired or talent in general, are there some commonality or patterns in terms of their past experience? In terms of other things that people in your audience can use as a hint?
Speaker C: Yeah. I ask myself every day, like what I was saying about the monkey brain. I haven't been able to come up with very clear guidelines, even for myself, on what to hire. But I do think when you have a long conversation with a person, you go through all the different things they talk about for their whole job and the attitude they project for the work they've done. At some point you do pick up on how they value craft, whether they're not shipping anything bad into production, if, like, if the customer is having a bad experience or not. Those things are, uh, often hard to fake. But in terms of, like, clear guidelines, I honestly think it's very hard. It's more like a, uh, detective role play kind of thing at this point where you try to understand what this person's personality is like. I do think sometimes it is hard to fake personalities, and it's also hard to shape personalities through culture in the sense that, like, our culture is very focused on open source models. We're also very focused on customers. Those things are to a certain extent, something you can educate and something you can influence, but it's just very, very hard to change. And whether this person is being genuine about making sure the quality of their work is high. I think humans have a unique advantage where they have a monkey brain that's like evolved over hundreds of thousands of years to tell whether someone's lying. It's very small, but I'm sure everyone has the skill to tell. And those things I think are more
Speaker B: useful if we take a step back and try to explain why the criteria for good talent have shifted this way. So I guess it's more to the nature of the work now in the AI age that human beings are making the judgment call. So you all have AI to use, but what is our taste of quality, uh, where we feel comfortable to stop so that on an ongoing basis will kind of shape the product, also shift the culture, shape the customer experience. That's why you want to find people that are naturally, even without looking over their shoulders, they have a high bar for a lot of things so that you can trust.
Speaker C: That uh, trust piece is definitely important. It's also much easier to ship like slot. In this day and age, having someone who has the mentality and the attitude to ship high quality software is very important for us internally and very important for our customers. It's very, very easy to have code that rots very quickly these days. You can spit out like 2000 line PRs overnight very very easily. Does that mean it uh, contribute positively to the team and to our customers? May or may not be true.
Speaker B: Tell us more about shifting away from the high bar, high uh, agency. So what about IQ? So the models has definitely higher IQs than typical human beings. So what's your bar on um, interview questions to test IQs?
Speaker C: I don't explicitly test for IQs, but these things are often correlated. People who are very proud of their work and have high quality barrier tends to have very high iq. There are certain cases where this person seems to be really smart but also seem to be really sloppy. That also happens, but I think happens quite rarely. I do think people realize this day and age it's very important about like quality is more important than anything. So like if they don't realize that, I think imply that the IQ may not be high enough.
Speaker B: Another thing we observed is the nature of engineering works are kind of shifting from doing the work to managing or management in general because you're, they're managing agents and the ability to delegate, the ability to know how to measure success and provide clear instructions become more important. What's your experience shifting the culture of uh, your engineering team from doing to more of a management of uh, agents mentality.
Speaker C: Yeah, I think the full mentality is also skill. Reading speed is very important these days. Being able to read just a lot of text very quickly so that you can quickly siphon out what the agents are doing. I don't know how to test it in interviews, but I do think people who are able to read very quickly have a unique advantage these days. So maybe, yeah maybe if you're an uh, English major who can read very, very quickly and remember very long texts and try to connect the dots very quickly. So that's often very useful. But then going back to your original question, talent is more focused on management. Yeah, I do think uh, the mentality is different. Whoever can adapt to that kind of mentality would definitely be more successful in the future. Being a manager of agents often is also about having good hygiene on remembering to put down the skills of the things that uh, these agents made mistakes on. Remember to set up automation so that you don't have to repeatedly do the same thing over and over again. All those hygiene, all those focus on dev efficiency is very very important. People who are good at improving dev efficiency for themselves as well as the team often make a uh, very large outsize impact because the models are getting better. And then like making sure the models themselves have very good dev efficiency will help speed up the whole process for the team. And often people who are managing these agents also need to take responsibility on the continuous integration process. Making sure we have super high quality CIs is important. The CI one is a good form of automation so then you can just leave the agent do its own job. Second is a good way to prevent regression in the code base. Preventing those regression is very important for agent managers I would say because it will save you the trouble of reverting all the changes made. So I often recommend like people setting up the CIs before they set up all the PRs to make the changes. It's slightly different from just like test driven development because it's very very easy to get the agent to give you one uh, hundred tests that's completely useless. But then like CI driven development may be more important as in like thinking about what's the right end to end test to set up, thinking about where to uh, push for end to end tests, thinking about how to maintain the health of these tests. All those things are very very important.
Speaker B: You made a very interesting distinction preventing regression Instead MCI is very different from test driven development. It's easy for agent to get all the tests passed and then you want
Speaker C: to call out specifically for this point. I think agents are too smart, they're too high iq. So then I think the old school test driven development is that you set up some tests and then you as a human write out all the functionalities and in your head ah, you have like an intended way to implement these things. But then the test is just making sure that you didn't do something stupid. And the test is still useful because it's human who have all these tribal knowledge and context are writing the code. So then there's no weird edge cases that you're skipping through. When the agents are writing the test and writing the code, they may not have the same context. They also may not have the same mentality of like making sure I can maintain this code for the long run. So they will do whatever it takes to get those tests passed. Right. That may not be the right way to do things. So then if you set up very rigorous end to end tests and make sure that all the tests are set up in the way that like from a user's perspective, the functionality works as intended, then there's no way for the agent to cheat. Then I think it's unfair to say that it's not test driven development. It still is. It needs to be designed more carefully, it needs to be thought through more carefully and uh, it needs to be tamper proof.
Speaker B: Can you give us an example of a good test to make sure agents are doing the right thing for training infrastructure?
Speaker C: We often test the training inference, consistency. So what we call like KL divergence, it's very easy for the agent to mess around locally and come up with a design that have low KR uh divergence locally. It's very hard to make sure that KR divergence is low end to end in production. So we often have like uh, end to end tests that uh, test for low KR divergence for our reinforcement learning stack. We're very proud of that. But it takes a lot of effort and it does take even like the best model a lot of iterations to get through it. And I can see in my cursor like chat history the models like oh, I test this locally, why does it not work uh, end to end. And then he's like, keep trying to make sure like tease out what's missing to be able to test these end to end. So it is a very rigorous test that both as a human and as an agent, very hard to lie about.
Speaker B: So in a way human test is more watching out for potential edge cases, making sure both the habit case and ad case works. But for agent they can think reversely. But the test is more to make sure that they are optimized for the right thing, not just getting tests passed.
Speaker C: Yeah, oftentimes it's also thinking about what are uh, the right metrics to look at for your organization. It used to be that if you are a leader of an organization, as long as the metrics is like correlated with what the outcome you are trying to follow, it's most of the time okay, unless you get to like a giant organization and then the people in the company exploit those metrics as much as they can. But the agents are different given any metrics, they will try to exploit it. So your metric Better be very hard to exploit. KR divergence I think is uh, one of the metric I shared earlier for like the quality of our training stack. Yeah, I think that may be the best example I can think of right now.
Speaker B: What are some of the best engineers, best performing engineers in your team, how they spend their day. So this is kind of getting into if you are a really AI native team and you're using AI to push AI to the edges, what uh, the day to day looks like for tougher front engineers.
Speaker C: Most of the coding time for me is probably very similar to everyone else. I often push for improvement for our continuous uh, integration from time to time and spend a lot of time also improving the CIS itself or push other people on my team to improve the CIs and also a better release process. That is probably like 20% of my time on a constant basis these days because it's very easy to have these things regress. It's important to keep them healthy. Uh, does take a lot of time and effort to make sure they're healthy.
Speaker B: That's really helpful to know. So another question I have, there are a lot of engineers, they have been working in industry for a long time and their organization may not be as advanced in terms of adoption than some of the companies are making a lot of progress. So how do you help or advice? Do you have to share with those engineers to be involved in this new AI driven age?
Speaker C: Start small. Start with uh, cursor or cloud code. Get your hands dirty. Yeah, start with codecs. Codex is very, very cost effective these days. It's incredibly cost effective to be honest. The joke I always make is uh, I think the researchers may go and the uh, infrastructure engineer may go. The SREs will stay. The production engineers, site reliability engineers, they will stay. I would definitely shift the mentality to more like end to end ownership and quality of the software and uh, prepare for the world where everyone just becomes SREs.
Speaker B: Uh, so your engineers on a regular basis they are responsible for production operation?
Speaker C: Oh absolutely, yes.
Speaker B: How's your um, release cycle?
Speaker C: Looks like release cycle. So we have a different way to do change management for different things. I would say probably like weekly.
Speaker B: Weekly across the org.
Speaker C: Across the org, let's say. Yeah, most of the time. For smaller changes it might be daily, but uh, for big changes, uh, probably like roughly once every week.
Speaker B: How do you see the role of uh, QA play in this new way of software development?
Speaker C: I do think the role collapses. There's probably not going to be separate QAs and uh, everyone just need to do their own QA and do their own SRE work. I think it's very hard to decompose these roles. It's more important these days than ever to make sure whatever you're shipping is high quality and doesn't need a separate person to do quality assurance.
Speaker B: Got it. So putting 18 years doing testing product operations so that they really own it end to end.
Speaker C: Absolutely. And ah, making sure we have good CIS in the team so that their job can be easy.
Speaker B: Great. Cool. Uh, I think there's all the questions I want to ask about the kind of the journey. I have a few rapid fire questions if you're uh, answering. So uh, my first question is what are some of the readings you have? What do you read and listen to to stay up to date?
Speaker C: To stay up to date, honestly. Twitter for example. Like uh, the Colossus deal with Anthropic. That was because Elon started following a bunch of people from uh, anthropic computing. A lot of these things just go on Twitter. Uh, they also happen first on Twitter.
Speaker B: What are some of the emerging trend that you have observed but has not made to the mainstream yet?
Speaker C: Everyone is becoming a model builder because clock code can do everything. It's not as hard to customize model these days. We are seeing a lot of people who have never tuned a model ever before. Able to just uh, talk with clockcode and tune models. That's always been amazing for me and I'm seeing this more and more every day. I'm sure this will catch on. We have a lot of tracking to show that a lot of people who are able to tune models never ever read our docs. Never. Their email never came across our docs website either because they're clearing the cookie and preventing tracking or something. But the percentage is way too high. Most people just don't even read our docs and still able to tune models. That has always been amazing to me. If you try to convince me a year ago you can tune a model without reading our docs, I would never believe you. But that's the fact on the ground today.
Speaker B: That's a good one. Last one. If you ask to give one advice to founders in this AI age to build their startup, what advice you would have given them?
Speaker C: Invest in young people. I do think it's a lot of companies are very cynical these days. It's like Anthropics is not hiring any of the E5s and below. They're only higher E6 and above. I think a lot of other companies are going through very aggressive layoffs. I honestly don't know if that's the right approach. These days. The young people are just as capable as when we were young, 10, 15 years ago. So yeah, it is very hard to convince yourself to convince young people these days. At the same time, I do think for some of the best and brightest, you should still do it and just accept the risk. I mean, like, if you're doing startup, you're accepting so much risk already, why not accept a little bit more, you know?
Speaker B: Yep. I think that's good for our whole initiative because you have to be young people first. Right. Otherwise the talent will not have a healthy inflow of talent over time.
Speaker C: Absolutely. Absolutely.
Speaker B: Great. Well, I learned a lot of really good insights, uh, unique insight that I never heard from other founders. So I think those are really good. Thanks for spending time and I truly enjoyed the conversation and wish had more time. Thanks again for spending time with us.
Speaker C: Thank you for having me, Jerry. Thank you.
Speaker D: If you're listening to this and you're wondering how can I connect with other engineering leaders in my city? Pull up your phone right now and go to elc.community click our chapters page. You can see that on the menu on the left. Find your local chapter and click Join. We're hosting virtual and in person events all the time Time. And this is the best way to help you get involved, expand your network in your city and support your leadership and career growth. So pull up your phone, head to ELC community, join your local chapter and get involved. A huge thank you to all of our local leaders who make community happen. And thank you for listening to the Engineering Leadership Podcast.
Speaker B: Sat. Mhm.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.