The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/The Engineering Leadership Podcast
The Engineering Leadership Podcast artwork

Hiring top-tier talent, leveraging open source models, and staying competitive in the age of AI w/ Benny Chen #267

The Engineering Leadership Podcast · 2026-09-02 · 35 min

0:00--:--

Key moments - from our scoring

Substance score

66 / 100

Five dimensions, 20 points each

Insight Density14 / 20
Originality13 / 20
Guest Caliber16 / 20
Specificity & Evidence12 / 20
Conversational Craft11 / 20

Benny Chen shares Fireworks AI's journey from its inception at Meta to becoming a key player in open source AI infrastructure. Most of the founding team came from PyTorch and adjacent teams at Meta, giving them deep conviction about open source's critical role in generative AI. Rather than chasing every opportunity, Fireworks deliberately focused on inference from day one - a contrarian bet that proved prescient - and evolved toward model customization as companies like Cursor, Genspark, and Vercel sought to build competitive advantages through proprietary data flywheels. Chen emphasizes that founder-led messaging isn't about outsourcing but ensuring the product story is told accurately; companies can't delegate this early-stage work to experienced go-to-market leaders without losing nuance. On competition, Fireworks survives by staying laser-focused on open source models and deliberately walking away from tempting distractions like closed-source image generation. The hiring paradigm has fundamentally shifted: rather than seeking people with narrow domain expertise (like 'Bluetooth experience'), Chen now prioritizes high agency, quality standards, and change management skills - traits that enable engineers to ask the right questions and extract knowledge from AI coding models on the fly. This conversation reveals the underlying conviction driving Fireworks' strategy: in the long run, generative AI models will function like recommendation systems, learning from customer data and user trajectories to deliver increasingly personalized experiences.

Key takeaways

  • →Founders must lead company messaging themselves rather than delegating to experienced go-to-market leaders, because the people building the product hold the full story and nuances that can't easily be transferred.
  • →As AI coding models improve, hiring criteria should shift from domain-specific expertise to high agency, quality standards, and change management skills - traits that enable learning on the job rather than requiring pre-existing knowledge.
  • →Staying competitive in AI infrastructure requires ruthless focus; Fireworks deliberately walked away from closed-source image generation despite its market potential to remain true to its open source thesis.
  • →The long-term value for application companies lies in building data flywheels - using customer traces, user requests, and reinforcement learning feedback to continuously customize and improve models for competitive advantage.
  • →Model customization through reinforcement learning is becoming essential for application companies competing against frontier labs; they can't match the capital burn of OpenAI or Anthropic, so they must build moats through proprietary data and custom models.

Guests

Benny Chen

Topics in this episode

Reinforcement learningEngineering managementVercelgenerative AICompetitive advantagePyTorchopen source modelsGenSparkFireworks AIdata flywheelsengineering leadershipelcmodel customizationCursor ComposerMeta AI infrastructure

Questions this episode answers

What should engineering leaders prioritize when hiring in the age of AI - domain expertise or other qualities?

High agency and quality standards matter more than domain expertise, since AI coding models allow engineers to acquire specific skills on the job; the focus should shift to hiring people who know how to ask the right questions and have a high quality bar for their work.

How should application companies use customer data to stay competitive against frontier AI labs?

By building data flywheels: collecting user request data and agent trajectories, training language model judges based on PM and engineering feedback, then using reinforcement learning to customize open source models to their specific use cases - effectively creating competitive moats through proprietary data.

Why did Fireworks focus on inference from the start rather than training?

Looking at Meta's AI infrastructure spending, inference consumed far more resources than training; training is an iteration cycle that happens before inference, so focusing on inference early proved contrarian but correct as the market matured.

What messaging should an early-stage AI infrastructure company use to position itself?

Fireworks settled on 'train and serve open source models' as its core thesis; clear focus here allows founders to guide the team away from distractions like closed-source image generation and concentrate on areas where they can deliver the best infrastructure.

How do you assess a candidate's quality standards and agency during hiring when these traits are hard to measure?

Through long, exploratory conversations exploring their full job history and attitudes toward craft; look for signs they won't ship bad code to production and whether they genuinely go above and beyond - these things are difficult to fake but the assessment remains more art than science.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

14 / 20

The episode delivers solid, actionable insights on hiring in the AI era, testing methodology for AI agents, and the evolution of infrastructure choices. However, substantial padding from ads, softball follow-ups, and repetitive reinforcement of single themes (open source focus, data flywheels) limits density. Key ideas are present but not packed tightly - much airtime is spent on summarization and agreement rather than exploration.

As the coding model got better, I think it's more and more important to hire people who have high agency and can figure out all the things they need to learn on the job.
I do think uh, the mentality is different. Whoever can adapt to that kind of mentality would definitely be more successful in the future. Being a manager of agents often is also about having good hygiene on remembering to put down the skills of the things that uh, these agents made mistakes on.

Originality

13 / 20

Benny articulates genuinely fresh thinking on CI-driven development versus TDD in the age of agents, and the shift from technical skills to quality bar + agency as hiring criteria. However, the core thesis (open source models matter, focus on inference, data flywheels as moats) reflects industry consensus already circulating by 2024. The specific application to agent management is novel, but broader positioning is incremental.

CI driven development may be more important as in like thinking about what's the right end to end test to set up, thinking about where to push for end to end tests, thinking about how to maintain the health of these tests
the monkey brain is more useful than the prefrontal cortex. I don't know how much you can analyze to understand whether this person has high quality bar or this person has high agency.

Guest Caliber

16 / 20

Benny Chen is a credible practitioner: co-founder of Fireworks AI (a real, funded company), former Meta infrastructure engineer, PyTorch-adjacent background, and actively building in the inference/training space. He speaks from direct execution experience and customer interaction (Cursor, Vercel, Genspark). This is not a career podcaster or pure theorist. His seniority and operational depth are genuine, though the company is relatively young and not yet a household name.

I'm Benny Chin, I'm one of the co founders at Fireworks. So we've been working on large language model training, inference for about four years. Before this I was at Meta working on ADS model, serving with GPU Asics
we went public with cursor on their composer two and um, 2.5 training. So cursor, for example is a company that has great proprietary data on coding.

Specificity & Evidence

12 / 20

The episode includes named customer examples (Cursor, Vercel, Genspark) and specific technical metrics (KL divergence testing). However, most claims lack concrete numbers: no revenue figures, customer counts, infrastructure scale metrics, or timelines on adoption rates. The hiring philosophy is explained but not backed by hiring data or success rates. The data-flywheel concept is illustrated with broad examples rather than quantified results.

we went public with cursor on their composer two and um, 2.5 training
For training infrastructure, we often test the training inference consistency. So what we call like uh, KL divergence, it's very easy for the agent to mess around locally and come up with a design that have low kr uh divergence locally.

Conversational Craft

11 / 20

The host asks relevant follow-ups and shows genuine curiosity (e.g., 'Can you give us one example of a customer...' and questions about testing agents). However, many opportunities for productive pushback are missed. When Benny claims hiring is 'very hard,' the host merely asks for patterns rather than probing the contradiction. Softball agreement ("sounds like your team are focusing on the right thing") replaces sharp interrogation. No tension or genuine disagreement emerges.

Can you give us one example of a customer that are building the data flywheel through Fireworks AI and what is the value they deliver back to their customers or users?
But in terms of clear guidelines, I honestly think it's very hard. It's more like a, uh, detective role play kind of thing at this point where you try to understand what this person's personality is like.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B72%
  • Speaker D21%
  • Speaker A5%
  • Speaker C2%

Most-used words

model48models41important30sure23infrastructure22source22open21team20test20training20data20high19different18quality17application14hard13

Episode notes

Benny Chen, Co-Founder @ Fireworks AI, joins the show to discuss his founder journey and share valuable insights on navigating common founder / product dev challenges in today’s agent-first landscape. He and Jerry cover strategies for creating effective messaging, staying competitive in a crowded market space, hiring top-tier talent / what qualities to look for in high-performing engineers, navigating the cultural shift to managing agents, creating data flywheels & how this can help your customers, and more. ABOUT BENNY CHEN As co-founder and early product architect, Benny Chen shaped Fireworks AI’ s infrastructure strategy, spearheading the design of scalable systems to support high-throughput AI model serving. Benny’s contributions established the technical foundation for Fireworks AI’s robust and cloud-native architecture, which underpins its ability to meet enterprise demands. Formerly Meta’s Ads Infrastructure Lead, Benny optimized large-scale ad-serving pipelines and developed significant expertise in distributed systems and cloud infrastructure. He holds a B.S.

Full transcript

35 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: This episode is brought to you by Cinch. Here's the thing about Cinch. You know them better than you think. That text confirming your order, the call that checked was really you logging in the email that landed right when it should. Behind each one is real technical complexity. Getting a message delivered reliably at scale across a hundred countries and carriers that all play by their own rules. It's usually not a problem that any engineering team wants to own, so Cinch owns it for you. They build a powerful infrastructure for messaging, voice, email list delivery, routing, compliance and a fraud prevention. Building APIs for builders who want to go deep ready made applications for the team who just need it to work. It's the communication infrastructure that the AI era runs on. Um, already trusted by over 200,000 businesses globally and empowering a trillion interactions every year. So the next time a message test works, there's a good chance it's powered by Cinch. Learn more@cinch.com that's s I n-dot com

Speaker B: at the end of the day you're hiring people to do certain things and then acquiring the skill early on is very, very difficult in hardware. I think the joke was that uh, like you just have to find a person who worked on Bluetooth for this manufacturer because that's what that person worked on for his whole life. No one else have the training data for it. As the coding model got better, I, uh, think it's more and more important to hire people who have high agency and can figure out all the things they need to learn on the job. You don't need to learn all those things before you come into the job. It's more important to figure out what the right question to ask so you can extract all those information out of the model. I emphasize much more on people who worked on change management and people who have a high quality bar for their work to make sure that we can set up all the right parameters for the coding agent to work in.

Speaker C: Hello and welcome to the Engineering Leadership Podcast brought to you by elc, the Engineering Leadership Community.

Speaker D: I'm Jerry Lee, founder of eoc.

Speaker C: And I'm Patrick Gallagher and we're your hosts. Our show shares the most critical perspectives, habits and examples of great software engineering leaders to help evolve leadership in the tech industry.

Speaker D: Hey Benny, thanks for spending time with us today to share your journey Building fireworks AI really, uh, excited to learn more about your experience and transitioning to entrepreneur the journey and there are a lot of things we talked about last time. So I'm really uh, eager to get into it to start. Can you share who you are in your journey building, uh, the company?

Speaker B: Absolutely. I'm Benny Chin, I'm one of the co founders at Fireworks. So we've been working on large language model training, inference for about four years. Before this I was at Meta working on ADS model, serving with GPU Asics and how we got started is one, I was working on capacity planning for ADS infrastructure and all the demand for AI infrastructure inside Meta was going through the roof. So it was very clear to me that y infrastructure is a important area for me to play in. And on top of that there were other people thinking about starting a company around the same time. So I hopped on the train and joined them.

Speaker D: How do you take from an early idea into a. Like a rocket ship now?

Speaker B: Yeah. Most of the early members on the founding team were Pytorch veterans. So Dima was a, uh, core contributor. Lin was the manager on the team. Originally we knew that we want to work on the AI ah, infrastructure space. It wasn't clear to us whether uh, the area to work in was recommendation system or generative AI. So we started working in recommendation system first. As ChatGPT came out and as generative AI workloads started to take off, we quickly pivoted to work on generative AI. I would say the demand is definitely much higher than anything I could have anticipated. At the same time I think we're in a good position to make a lot of impact.

Speaker D: Tell us more about as a founder in the early days, how does messaging matters? Because this is something that when you and I chatted earlier, that you shared this particular point about you can't outsource messaging to a very experienced go to market leader.

Speaker B: Yeah, I think that's a good question. It's not so much that uh, you can't outsource. The stories are uh, ideally told by the people who are working on the product. Articulating an idea often is a multi fold, multi prong process. And then the people who work on the product have the full story sometimes in their head. And communicating that clearly takes a lot of work. It is easier to tell that story by yourself. It is not scalable and ideally it is something that the messaging can be scaled to many, many other people. I would say like, uh, for Fireworks, early on we knew AI infrastructure was going to take off. That was very clear to us and that was unequivocal how it would take off. I think that's still up in the air. So directionally we knew that we want to work in AI infrastructure We could work on AI infrastructure on recommendation system or generative AI. And uh, inside generative AI we also have like text, image and audio models. So it wasn't so much that we had to pivot in the middle. We are very focused on open source and we are very focused on AI infrastructure and that were the two main focus of the company. At the end of the day it turns out text models, open source models were the area for us to play in. So we were able to focus the company on these efforts. But I think the transition is pretty natural because we are big believers in open source and we want to focus

Speaker D: on infrastructure that ties to the prior experience of the founding team. You mentioned Pytorch.

Speaker B: Yeah, Pytorch. Uh, most of the founding members were Pytorch and uh, Pytorch adjacent teams. And for us it was clear that open source will play a huge role in generative AI, whether at the modeling layer or the infrastructure layer or the layers below. We do see that open source models iterate much faster than closed source models because there are many, many people exchanging ideas in the open. So supporting open source model was a uh, big hypothesis for us. And I think three or four years ago people were still wondering whether training or inference was going to be the main focus. We were very clear about inference from day one. We always look at our prior experience in supporting metas AI infrastructure and uh, most of the spend goes to inference. Training is definitely very, very important, but training models is what you would do before you do inference. So it's an iteration cycle. Making sure that we focused on inference early on was something that's contrarian early on. I don't know why for the industry, uh, but we were very focused on inference. And then as we ramped up our inference efforts, the whole reinforcement learning wave came. Reinforcement learning is all about having really good inference. So we were able to contribute to the training area as well.

Speaker D: It feels like your team are focusing on the right thing and you're getting ready to take the big wave of the market momentum just in time, when it's ready.

Speaker B: We honestly had a pretty simple thesis. This area though is very execution heavy. The execution quality is very, very important. We make a lot of mistakes still every day and we try our best to avoid those mistakes. We want to make sure we deliver the best infrastructure for our customers.

Speaker D: Can you share how that played out? Like you have a unique insight that are not agreed by the mainstream which give you the advantage to start early. How does that early insight carry the company's messaging over time, since we're talking about that topic.

Speaker B: So initially it was mostly about serving the open source model as is at the very beginning we were serving llama, uh, models and those models are pretty hard to use and the usage surface area was also relatively small. So a lot of initial workloads, uh, are sort of AI assisted workload. As the industry matured we started to see one, the open source models are being more sophisticated to the areas of which the AI models is deployed is also more mission critical. So their model customization became more and more important. We want to help people customize models so they can optimize the model for their application. And so initially it was more geared towards serving the model as is. And as time went on, the model customization and the training piece became more and more important. It's less about whether the capacity or the revenue of the company is coming from training itself. It's more that the customization of open source models are becoming more and more important in the industry and I believe it will be very important in the future.

Speaker D: Can you share a bit more about what kind of customization that becomes more popular?

Speaker B: Yeah, uh, recently we went public with cursor on their composer two and um, 2.5 training. So cursor, for example is a company that has great proprietary data on coding. They want to be able to scale up their training aggressively and we help them scale up their training workload onto multiple regions to be able to do reinforcement learning across many small data centers instead of one giant data center. There are also many other companies we went public with like genspark, like Vercel, where we help them customize models so then they can use our infrastructure for training and then later on use the same infrastructure for inference. All these application companies have really good data and really good PMs who can describe what is good and what is bad. Those information can be all baked into the model to help the model understand how to behave in certain conditions so that their customers can derive more value out of the model.

Speaker D: So every company is different, their use case is different, the way they want to train the model is different. So there's almost like infinite variations of use cases. So that's the demand.

Speaker B: Absolutely. At the end of the day, I think data is where you have the moat. And uh, all these application companies are in a great position to capture all the customer intent and turn those into really cost efficient and high quality models.

Speaker D: And going back to our earlier question about messaging, so uh, what is the messaging that you landed for now here

Speaker B: we want to Train and serve open source models. That's sort of the core thesis of the company and all our messaging revolves around that. At the end of the day you want to come to fireworks to be able to customize models and serve those customized models. And that, uh, sort of evolved over time. From early on we were just serving the open source model as is now.

Speaker D: You have now learned the messaging, you have unique insights, you gained the momentum and benefit from the big shift. There's also a lot of competition in this space. There are a lot of money putting into the AI industry. So how do you stay competitive and be a winner? So in this very competitive market, yeah,

Speaker B: I think staying focused is very, very important. The whole field is very execution quality heavy and we are very focused on making sure that we deliver the best inference and training stack to our customer. A lot of companies in the area are often distracted about different things and chasing different shiny objects. We just want to train ourselves. Open source models.

Speaker D: Do you have an example of walking away from a, uh, shiny object? I'm sure there's a lot of them along the way.

Speaker B: I think one good example was image generation. So we used to be a big proponent for image generation as well. It turns out most image generation models are not open source. So it is very different from our thesis where we want to stay focused on serving and training open source models. So yeah, we slowly diverted away from it as all the models in the uh, field become closed source. It was very tempting to think about, hey, like, how do we figure out different forms of partnerships or setup where we can serve these proprietary models as well. But at the end of the day, if we can convince ourselves that we can help our customers set up the data flywheel. I don't think it was as interesting as the uh, open source text models. Maybe it will come back as the text model gets the capability to generate images as well, but it may not come back anytime soon. And that's something we had to walk away from because sort of like the ecosystem around open source models died out.

Speaker D: Yeah, it felt like you have your team have early and have the core conviction about open source, how open source are going to be really critical for AI infrastructure. So that conviction guided the team away from the distractions. So that's.

Speaker B: Yeah, uh, distraction or not, maybe like there are definitely companies in the area that makes a lot of money on image generation, that's for sure. I think for a small company you have to stay focused. So there are definitely areas where we have to make certain sacrifices to stay focused.

Speaker A: This episode is Brought to you by Cinch. Here's the thing about Cinch. You know them better than you think. That text confirming your order, the call that checked was really you logging in the email that landed right when it should. Behind each one is real technical complexity. Getting a message delivered reliably at scale across a hundred countries and carriers that all play by their own rules. It's usually not a problem that any engineering team wants to own, so Cinch owns it for you. They build a powerful infrastructure for messaging, voice, email list delivery, routing, compliance and a fraud prevention. Built in APIs for builders who want to go deeper, ready many applications for the team who just need it to work. It's the communication infrastructure that the AI era runs, um, on already trusted by over 200,000 businesses globally and empowering a trillion interactions every year. So the next time a message just works, there's a good chance it's powered by Cinch. Learn more@cinch.com that's s I n c

Speaker D: h.com you mentioned about customer data Flywheel. What do you mean by customer data flywheel?

Speaker B: Uh, every application have unique insights into how to use their data. They will collect customer traces about how to, you know, vibe code a web page or come up with design or, I don't know, generate some 3D assets. The nuance is that uh, how to evaluate whether the output is good or bad and whether the output is good or bad in the context of the application itself. All those work that the PMs and the engineering people on those application team do to improve their product can eventually be articulated and translated into language models judge so that the model can learn the same thing as well. So all these information can flow into the model and all the customer uh, usage on the application can flow into the model so that they can have really really good model and also derive really good ROI from those models. I do believe in the long run at scale, all these generative AI models will look more and more like recommendation systems where we learn the preference for different people or how they use different applications and the model can respond accordingly. So yeah, that's very important for us because we had very strong belief on data flywheel and model customization and we help all the application developers in the space who want to train models to use our platform to customize the model and serve them in production.

Speaker D: Can you give us one example of a customer that are building the data flywheel through Fireworks AI and what is the value they deliver back to their customers or users?

Speaker B: Yeah, I think uh, the recent Composer launch is Definitely a really good example. Like I love the model, the model is great, they work a lot. Customizing the model to the cursor harness itself so that uh, you are able to train a really good coding model. Genspark is another example where we went public with that uh, customize the uh, model for different search and slide generation functionalities in the application so that they're able to get really good slides and uh, really good search results for their users. All these customers we work with are uh, big proponents in setting up the data pipeline, setting up all the customer preferences they know about as a training loop and then um, putting those information into the model itself. And we believe there are more and more people like this who have the appetite to customize models and have the conviction to deliver the best model for their customers.

Speaker D: Got it. As you said, AI models in the long run will more become a recommendation system so that what you are doing is helping your customers to take their consumer or customer data, uh, so that they learn more about what their customer needs so that they deliver a much more nuanced, context driven or personalized experience back to their end users.

Speaker B: Absolutely, absolutely.

Speaker D: And that's the trend for all the application to be right in order to win this market?

Speaker B: I think so, yeah. As the spend on Genai goes up, I do believe more and more application companies will lean towards customized models because they simply don't have the balance sheet versus these frontier labs. They cannot afford to burn as much money so they need to figure out uh, how to customize models and they need to figure out uh, how to beat the frontier labs in areas where they have a moat.

Speaker D: The question about what kind of data company leverages to train their own customized model. So you mentioned earlier the product usage data. What are the common type of data companies collect and use to train their model?

Speaker B: So oftentimes it comes in the form of user request and the agent trajectory from those requests. And then what's important for reinforcement learning is that the application developers on Those teams, the PMs on those teams, can clearly articulate those requirements as language models judge. So then we can use the language uh, models to judge the outcome of the model in the reinforcement learning loop and then sort of bias the model towards what the training process wants the model to do. It's like initial user request and then the final language model is a judge and everything else in the middle is generated on the fly through uh, reinforcement learning.

Speaker D: Got it. In the case of genspark, you mentioned slide generation, they provide input in terms of judging whether a slide generated by their application is good or not?

Speaker B: Absolutely, yeah.

Speaker D: So now assuming customer has built the data flywheel, what are the next thing they do is that ongoing, um, basis with fireworks AI or they will unlock more things that you can help them with?

Speaker B: Yeah, I think ongoing is definitely happening a lot. Especially when you have new open source model release. You need to sort of retrain the whole setup with a new uh, base model. You also derive more insights from the user using these applications and then for reinforcement learning. Often it's like a whack a mole game. Because the model is often lazy and will try to figure out what's the easiest thing to do to fit your requirements. You need to often add more data and more judgment to the process so that the model quality can be better at the end of the day. So it is a repeated process.

Speaker D: And typically what team do you work

Speaker B: with at those customers, companies, the application developers or the MLEs, uh, in those companies.

Speaker D: Got it. So going back to your own building, your own company competition not only means product level competition, messaging level competition, there also means talent competition. So in the new age of AI, tell us how you see the bar of uh, hiring becomes for engineering talent.

Speaker B: I think it's different. We used to at least uh, when I was a meta, focus on whether this person has worked on certain things before because at the end of the day you're hiring people to do certain things and then acquiring the skill early on is very, very difficult. I think a good example is in hardware. I think the joke was that like you just have to find a person who worked on Bluetooth for this manufacturer because that's what that person worked on for his whole life. No one else have the training data for it. As the coding model got better, I think it's more and more important to hire people who have high agency and can figure out all the things they need to learn on the job. You don't need learn all those things before you come into the job. It's more important to figure out what the right question to ask so you can extract all those information out of the model. What is important has also been changing. I emphasize much more on people who worked on change management and people who have a high quality bar for their work to make sure that we can set up all the right parameters for the coding agent to work in, set up all the tests and all the release process or the design thinking so that the agent can follow and deliver really good infrastructure for our customers. The actual skill is I think less

Speaker D: important because those can be acquired on the fly through AI. And how do you test those qualities?

Speaker B: Honestly, I think it's very, very hard. So we try our best to go through cultural realm discussions with candidates. We try to understand if they go above and beyond to fix something previously in their old job. But I do think this is where, you know, at certain point, the, uh, joke I always like to make is that the monkey brain is more useful than the prefrontal cortex. I don't know how much you can analyze to understand whether this person has high quality bar or this person has high agency. I do think this is where it's more nuanced. And oftentimes you just have a conversation with this person for 45 minutes and try to pick up those, uh, signals, which sometimes is quite sparse.

Speaker D: If you look at the best performing engineers you hired or talent in general, are there some commonality or patterns in terms of their past experience, in terms of other things that people in your audience can use? Um, as a hint?

Speaker B: Yeah. I ask myself every day, like what I was saying about the monkey brain. I haven't been able to come up with very clear guidelines, even for myself, on what to hire. But I do think when you have a long conversation with a person, you go through all the different things they talk about for their whole job and the attitude they project for the work they've done. At some point, you do pick up on how they value craft, whether they're not shipping anything bad into production, if the customer is having a bad experience or not. Those things are, uh, often hard to fake. But in terms of clear guidelines, I honestly think it's very hard. It's more like a, uh, detective role play kind of thing at this point where you try to understand what this person's personality is like. I do think sometimes it is hard to fake personalities. And it's also hard to shape personalities through culture in the sense that, like, our culture is very focused on open source models. We're also very focused on customers. Those things are, to a certain extent, something you can educate and something you can influence, but it's just very, very hard to change. And whether this person is being genuine about making sure the quality of their work is high. I think humans have a unique advantage where they have a monkey brain that's like, evolved over hundreds of thousands of years to tell whether someone's lying. It's very small, but I'm sure everyone has the skill to tell. And those things, I think are more

Speaker D: useful if we take a step back and, um, try to explain why the criteria for good Talent have shifted this way. So I guess it's more to the nature of the work now in the AI age that ah, human beings are making the judgment call. So you all have AI to use. But what is our taste of uh, quality where we feel comfortable to stop so that on an ongoing basis will kind of shape the product, also shift the culture, shape the customer cares. That's why you want to find people that are naturally, even without looking over their shoulders. They have a high bar for a lot of things so that you can trust.

Speaker B: That uh, trust piece is definitely important. It's also much easier to ship like slot. In this day and age, having someone who has the mentality and the attitude to ship high quality software is very important for us internally and very important for our customers. It's very, very easy to have code that rots very quickly these days. You can spit out like 2000 line PRs overnight very very easily. Does that mean it uh, contribute positively to the team and to our customers? May or may not be true.

Speaker D: Tell us more about shifting away from the high bar, high uh, agency. So what about IQ? So the models has definitely higher IQs than typical Ah, human beings. So what's your bar on um, interview questions to test IQs?

Speaker B: I don't explicitly test for IQs, but these things are often correlated. People who are very proud of their work and have high quality barriers tends to have very high iq. There are certain cases where this person seems to be really smart but also seem to be really sloppy. That also happens, but I think happens quite rarely. I do think people realize this day and age is very important about like quality is more important than anything. So like if they don't realize that I think imply that the IQ may not be high enough.

Speaker D: Another thing we observed is the nature of engineering works are kind of shifting from doing the work to managing or management in general because you're they're managing agents and the ability to delegate, the ability to know how to measure success and provide clear instructions becomes more important. What's your experience shifting the culture of uh, your engineering team from doing to more of a management of agents mentality?

Speaker B: Yeah, I think the full mentality is also skill. Reading speed is very important these days. Being able to read just a lot of text very quickly so that you can quickly siphon out what the agents are doing. I don't know how to test it interviews but I do think people who are able to read very quickly have a unique advantage these days. So maybe, yeah maybe if you're an English major who can Read like very, very quickly and remember like very long texts and try to like connect the dots very quickly. So that's often very useful. But then going back to your original question, talent is more focused on management. Yeah, I do think uh, the mentality is different. Whoever can adapt to that kind of mentality would definitely be more successful in the future. Being a manager of agents often is also about having good hygiene on remembering to put down the skills of the things that uh, these agents made mistakes on. Remember to set up automation so that you don't have to repeatedly do the same thing over and over again. All those hygiene, all those focus on dev efficiency is very, very important. People who are good at improving dev efficiency for themselves as well as the team often make a very large outsize impact because the models are getting better. And then like making sure the models themselves have very good dev efficiency will help speed up the whole process for the team. And often people who are managing these agents also need to take responsibility on the continuous integration process. Making sure we have super high quality CIs is important. The CI is a good form of automation. So then you can just leave the agent to do its own job. Second is a good way to prevent regression in the code base. Preventing those regression is very important for agent managers I would say because it will save you the trouble of reverting all the changes made. So I often recommend like people setting up the CIs before they set up all the PRs to make the changes. It's slightly different from just like test driven development because it's very, very easy to get the agent to give you 100 tests. That's completely useless. But then like CI driven development may be more important as in like thinking about what's the right end to end test to set up, thinking about where to push for end to end tests, thinking about how to maintain the health of these tests, all those things are very, very important.

Speaker D: You made a very interesting distinction. Preventing regressions in MCI is very different from test driven development. It's easy for agents to get all ah, the tests passed and then you

Speaker B: want to call out specifically for this point. I think agents are too smart, they're too high iq. So then I think the old school test driven development is that you set up some tests and then you as a human write out all the functionalities and in your head uh, you have like an intended way to implement these things. But then the test is just making sure that you didn't do something stupid and the tests are still useful because it's Human who have all these tribal knowledge and context are writing the code. So then there's no weird edge cases that you're skipping through. When the agents are writing the test and writing the code, they may not have the same context. They also may not have the same mentality of like making sure I can maintain this code for the long run. So they will do whatever it takes to get those tests passed. Right. That may not be the right way to do things. So then if you set up very rigorous end to end tests and make sure that all the tests are set uh up in the way that like from a user's perspective the functionality works as intended, then there's no way for the agent to cheat. Then I think it's unfair to say that it's not test driven development. It still is. It needs to be designed more carefully, it needs to be thought through more carefully and uh, it needs to be tamper proof.

Speaker D: Can you give us an example of a good test to make sure agents are uh, doing the right thing?

Speaker B: For training infrastructure, we often test the training inference consistency. So what we call like uh, KL divergence, it's very easy for the agent to mess around locally and come up with a design that have low kr uh divergence locally. It's very hard to make sure that uh, KR divergence is low end to end in production. So we often have like uh, end to end tests that uh, test for low KR divergence for our reinforcement learning stack. We're very proud of that. But it takes a lot of effort and it does take even like the best model a lot of iterations to get through it. And I can see in my cursor uh like chat history, the model's like oh, I test this locally, why does it not work uh, end to end. And then he's like keep trying to make sure like tease out what's missing to be able to test uh, these end to end. So it is a very rigorous test that both as a human and as an agent, very hard to lie about.

Speaker D: So in a way human test is more watching out for potential edge cases, making sure both the habit case and ad case works. But for agent they can think reversely. But the test is more to make sure that they are optimized for the right thing, not just getting tests passed.

Speaker B: Yeah, oftentimes it's also thinking about what are the right metrics to look at for your organization. It used to be that if you are a leader of an organization, as long as the metrics is like correlated with what the outcome you are trying to follow. It's most of the time okay, unless you get to like a giant organization and then the people in the company exploit those metrics as much as they can. But the agents are different given any metrics, they will try to exploit it. So your metric better be very hard to exploit. KR divergence I think is uh, one of the metric I shared earlier for the quality of our training stack. Yeah, I think that may be the best example I can think of right now.

Speaker D: What are some of the best engineers, best performing engineers in your team, how they spend their day. So this is kind of getting into if you are a really AI native team and you're using AI to push AI to the edges, what uh, the day to day looks like for tougher from engineers.

Speaker B: Most of the coding time for me is probably very similar to everyone else. I often push for improvement for our uh, continuous integration from time to time and spend a lot of time also improving the CI itself or push other people on my team to improve the CIs and also a better release process that is probably like 20% of my time on a constant basis these days because it's very easy to have these things regress. It's important to keep them healthy. Uh, does take a lot of time and effort to make sure they're healthy.

Speaker D: That's really helpful to know. So another question I have, there are a lot of engineers, they have been working in industry for a long time and their organization may not be as advanced in terms of adoption than some of the companies are making a lot of progress. So how do you help or advice do you have to share with those engineers to be involved in this new AI driven age?

Speaker B: Start small. Start with uh, cursor or cloud code. Get your hands dirty. Yeah. Start with Codex. Codex is very, very cost effective these days. It's incredibly cost effective. To be honest. The joke I always make is uh, I think the researchers may go and the uh, infrastructure engineer may go. The SREs will stay. The production engineers, site reliability engineers, they will stay. I would definitely shift the mentality to more like end to end ownership and quality of the software and uh, prepare for the world where everyone just becomes SREs.

Speaker D: So your engineers on a regular basis they are responsible for production operation?

Speaker B: Oh absolutely, yes.

Speaker D: How's your um, release cycle look like?

Speaker B: Release cycle. So we have a different way to do change management for different things? I would say probably like weekly, weekly across the org. Across the org, let's say. Yeah, ah, most of the time for smaller changes. It might be daily but uh, for big changes, uh, probably like roughly once every week.

Speaker D: How do you see the role of QA play in this new way of software development?

Speaker B: I do think the role collapses. There's probably not going to be separate QAs. Everyone just need to do their own QA and do their own SRE work. I think it's very hard to decompose these roles. It's more important these days than ever to make sure whatever you're shipping is high quality and doesn't need a separate person to do quality assurance.

Speaker D: Got it. So putting engineers doing testing, product operations so that they really own it end to end.

Speaker B: Absolutely. And making sure that we have good CIS in the team so that their job can be easy.

Speaker D: Great. Cool. Uh, I think there's all questions I want to ask about the kind of the journey. I have a few rapid fire questions if you're uh, answering. So my first question is what are some of the readings you have? What do you read and listen to to stay up to date?

Speaker B: To stay up to date. Honestly, Twitter for example, like uh, the Colossus deal with Anthropic. That was because Elon started following a bunch of people from uh, Anthropic's computing. A lot of these things just go on Twitter. Uh, they also happen first on Twitter.

Speaker D: What are some of the emerging trend that you have observed but has not made to the mainstream yet?

Speaker B: Everyone is becoming a model builder because cloud code can do everything. It's not as hard to customize model these days. We are seeing a lot of people who have never tuned a model ever before able to just uh, talk with clock code and tune models. That's always been amazing for me and ah, I'm seeing this more and more every day. I'm sure this will catch on. We have a lot of tracking to show that a lot of people who are able to tune models never ever read our docs. Never. Their email never came across our docs website either because they're clearing the cookie and like preventing tracking or something. But the percentage is way too high. Most people just don't even read our docs and still able to tune models. That has always been amazing to me that if you try to convince me a year ago you can tune a model without reading our docs, I would never believe you. But that's the fact uh, on the ground today.

Speaker D: That's a good one. Last one, if you ask to give one advice to founders in this AI age to build their startup, what advice you would have given them?

Speaker B: Invest in Young people. I do think it's a lot of companies are very cynical these days. It's like Anthropics is not hiring any of the E5s and below, they're only higher. E6 and above. I think a lot of other companies are going through very aggressive layoffs. I honestly don't know if that's the right approach these days. The young people are just as capable as when we were young, 10, 15 years ago. So, yeah, it is very hard to convince yourself to convince young people these days. At the same time, I do think for some of the best and brightest, you should still do it and just accept the risk. I mean, like, if you're doing startup, you're accepting so much risk already, why not accept a little bit more?

Speaker D: You know, I think that's good for our whole initiative because you have to be young people first. Right. Otherwise the talent will not have a healthy inflow of talent over time.

Speaker B: Absolutely. Absolutely.

Speaker D: Great. Well, I learned a lot of really good insights, uh, unique insight that I never heard from other founders. So I think those are really good. Thanks for spending time and I truly enjoyed the conversation and wish had more time. Thanks again for spending time with us.

Speaker B: Thank you for having me, Jerry. Thank you.

Speaker C: If you're listening to this and you're wondering how can I connect with other engineering leaders in my city? Pull up your phone right now and go to elc.community click our chapters page. You can see that on the menu on the left. Find your local chapter and click Join. We're hosting virtual and in person events all the time. Time. And this is the best way to help you get involved, expand your network in your city and support your leadership and career growth. So pull up your phone, head to ELC.community, join your local chapter and get involved. A huge thank you to all of our local leaders who make community happen. And thank you for listening to the Engineering Leadership podcast.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Building a $4 Billion AI Infra Company | Benny Chen, cofounder of Fireworks AIInfinite Curiosity Pod with Prateek Joshi · features Benny Chen73 / 100
  • Hiring top-tier talent, the importance of adopting an evolutionary mindset, and taking critical feedback to solve technical problems w/ Sergiy Nesterenko @ QuilterEngineering Founders · on Engineering management88 / 100
  • What Marketers Can Control When AI Changes Everything with Nick Wedewer, VP of Growth Marketing at Hims & HersThe Partnership Economy · on generative AI87 / 100
  • Less about Models; More about ArchitecturePractical AI · on open source models85 / 100
  • Code Review Is a Taste Problem | David Poll ⁨@GitHub⁩Hangar DX Podcast · on engineering leadership81 / 100
  • 657. Waziri Garuba, CEO of Harlem Labs, Introducing G.R.I.O.TUnleashed · on Vercel80 / 100

More from The Engineering Leadership Podcast

All episodes →
  • Determining bets, using customer insights to pivot, and gaining developer buy-in & trust w/ Tomas Reimers #26483 / 100
  • Utilizing AI internally to iterate faster and empower smaller teams to upskill w/ Vivek Raghunathan #26396 / 100
  • The Product Paradigm Shift: How Livekit Navigated High Stakes Scaling Challenges to Build the Future of Voice-First AI Interfaces w/ Russ d’Sa #26276 / 100
  • Building an empowered career w/ Jean Hsu & Cate Huston #26163 / 100
  • Redefining profit, centering human flourishing, and building an incorruptible mission-driven roadmap w/ Eric Ries #26075 / 100
Explore the best B2B Engineering & DevTools podcasts →
All The Engineering Leadership Podcast episodes →