The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/Platform Engineering Podcast
Platform Engineering Podcast artwork

Green CI and Merge Queue Mastery with Trunk’s Eli Schleifer

Platform Engineering Podcast · 2026-04-15 · 50 min

0:00--:--

Key moments - from our scoring

Substance score

72 / 100

Five dimensions, 20 points each

Insight Density16 / 20
Originality13 / 20
Guest Caliber17 / 20
Specificity & Evidence14 / 20
Conversational Craft12 / 20

Trunk addresses a critical bottleneck in modern software delivery: unreliable CI pipelines and merge queues that collapse under the velocity of AI-generated code. Eli Schleifer shares his journey from Microsoft to Google to Uber, where $600-700M in annual engineering salaries were blocked by broken CI - the same problem plaguing teams today as Cursor, Claude, and Copilot flood repositories with pull requests. Trunk's merge queue uses predictive testing and dynamic parallelism to eliminate backed-up queues, while its flaky test quarantining system uses AI-powered failure fingerprinting to distinguish truly broken tests from infrastructure-related timeouts. The platform integrates with any CI system (GitHub Actions, GitLab, etc.) via JUnit XML output and a Rust-based analytics CLI. For teams adopting agentic engineering, Trunk's MCP integration feeds historical test data into Claude and Cursor so agents avoid chasing false failures. Cory O'Daniel discusses his three-person team's experience: generating 15-20 PRs daily with AI tools has exposed gaps in their CI - particularly a flaky deployment timer test that Claude once tried to "fix" by converting their entire test suite from async to sync, turning 17-second runs into 15 minutes. The conversation highlights how teams with strong CI foundations see the most success with generative AI, while those without it face exponential waste as AI velocity outpaces testing reliability.

Key takeaways

  • →Merge queues with dynamic parallelism become necessary at lower team sizes (10+ engineers with AI) than historically required (20-30), because AI agents generate code faster than humans can land it.
  • →Flaky test quarantining using AI-powered failure fingerprinting distinguishes infrastructure failures from actual bugs, preventing agents from wasting cycles "fixing" timeouts or rate limits that have existed for years.
  • →AI agents given access to historical test failure data via MCP integrations (Claude, Cursor) make better decisions about which failures to address versus which to ignore.
  • →Teams with strong CI practices and test reliability see the most success with generative AI adoption; weak CI becomes a compounding problem when AI multiplies PR velocity.
  • →Trunk's merge queue can dynamically create parallel testing lanes based on domain-driven design boundaries in monorepos, eliminating unnecessary waiting when changes don't interact.

Guests

Eli Schleifer

Topics in this episode

ClaudeCursorCopilotMerge queueFlaky test detectionFailure fingerprintingDynamic parallelismTest quarantiningJUnit XMLMCP integration

Questions this episode answers

What is a merge queue and when do teams need one?

A merge queue uses predictive testing to catch logical merge conflicts (where two green PRs together break the build). Historically needed at 20-30 engineers, AI-accelerated teams need it at 10+ people because agents generate PRs faster than sequential testing can handle.

How does Trunk distinguish between a flaky test and a genuinely broken test?

Trunk uses AI embeddings to create failure fingerprints - analyzing the failure reason in JUnit XML and comparing it to historical failures. If a test fails in a new way (different error type or signature), Trunk alerts you it's likely a real break; if it matches a 6-year-old timeout pattern, it's known flakiness and can be quarantined.

How do AI agents access Trunk's historical test data to avoid fixing false failures?

Trunk provides an MCP integration that lets Cursor and Claude pull historical test failure information directly into their context, so agents can see if a CI failure is a known flaky test before attempting a fix.

What CI systems does Trunk integrate with?

Trunk is agnostic to CI systems (GitHub Actions, GitLab, etc.). It accepts JUnit XML output from any test runner and posts data via a Rust-based analytics CLI that tracks failures and manages quarantining.

How does dynamic parallelism in merge queues prevent teams from waiting on unrelated code?

Trunk analyzes build system metadata (Bazel, Buck, Nx, or custom patterns) to identify domain boundaries and create separate testing lanes in parallel; if you're changing frontend code, you don't wait behind backend tests.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

16 / 20

The episode delivers substantial technical content with specific concepts like logical merge conflicts, failure fingerprinting, dynamic parallelism, and anti-flake protection mechanisms. However, there are extended tangential discussions (e.g., lengthy riffs on AI-generated frontend code, landing page redesigns) that dilute the density of novel CI/merge queue insights.

The silliest way to get around that is to make sure that every time Main changes, you just have to retest everything... all your active pull requests. That's insanely expensive.
we have another bite at that apple. Right behind me is your PR that also has my changes and your changes. So together, if that change passes, we know that I'm good to go also. That it's probably just a flake.

Originality

13 / 20

The technical framing around merge queues and failure fingerprinting is solid but largely industry-standard. The novel angle is applying flaky test quarantining and fingerprinting specifically to AI-generated code workflows, which is timely. However, most core concepts (merge queues, CI reliability, testing strategies) are well-established patterns rather than fresh thinking.

we're really looking to become that outer loop company. We want to keep engineering moving and basically make CI work.
The agents want to have full context. A monorepo is the best way for the agent to understand the totality of the system.

Guest Caliber

17 / 20

Eli Schleifer is a strong operator: 11 years at Microsoft (Windows Mobile), startup acquisition to Google, Uber ATG (self-driving cars), and currently CEO of Trunk. He has direct experience building at scale and confronting real CI bottlenecks. His perspective on inner loop vs. outer loop problems in the context of AI-generated code is grounded in operational reality, not theory.

Inside Google, I was just blown away by how easy it was to build software. The Google three stack, the DevX experience in there was just bar none.
We had, you know, some six hundred, seven hundred of, probably at that time, the most expensive engineers you could hire - these self driving engineers. And we couldn't get code into the repo. It was just impossible.

Specificity & Evidence

14 / 20

Solid use of concrete examples: Uber ATG with 600-700 engineers and multi-day merge queue backlogs, Fair as a customer with 20 parallel lanes, Brex as a recent flaky test system customer, Trunk's GitHub Ruby Monolith stress test limits. However, many claims lack numbers: how much faster are merge queues? What percentage flakiness reduction? How many engineers does Trunk actually have (stated as 16, 11 engineers, but exact count unclear)?

We can handle more pull requests than GitHub itself can handle. So we basically maxed out the testing on that at the limits of the Ruby Monolith at GitHub.
At Fair, they literally had twenty parallel lanes running inside their merge queue and none of the code interacting.

Conversational Craft

12 / 20

Cory asks solid foundational questions (what is a merge queue, when do teams need one, how does it integrate with CI systems) but rarely presses back on claims or explores trade-offs. When Eli claims teams can skip code review for generated 10,000-line features, Cory engages conversationally but doesn't challenge the security or maintainability assumptions. The discussion often meanders into AI culture war territory rather than sharply interrogating CI-specific claims.

So for folks who haven't worked with a merge queue before, could you tell us a little bit about what a merge queue is?
Yeah, so how do you guys handle... let's say it's a flaky test, but let's say we actually just... somebody broke it, like three changes up from me.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

code56merge29queue27test27tests22flaky19build18teams16sure16slike16system15product14pull13engineers12engineer11testing11

Episode notes

When a flaky test can stall a merge queue, “just rerun CI” stops scaling fast. Cory talks with Trunk co-founder and CEO Eli Schleifer about the outer loop problems that show up as teams ship more code - especially with AI-assisted development increasing PR volume. They break down what a merge queue is, why logical merge conflicts happen even when individual PRs are green, and how predictive testing helps protect main without forcing constant retesting. Eli also explains how Trunk approaches flaky tests: collecting JUnit results, using quarantines so known flakes don’t block delivery, and fingerprinting failures to tell the difference between “this always times out” and “this was just broken by a recent change.” The conversation closes on how review and quality practices may shift as code generation accelerates - and what still needs strong guardrails like tests, security checks, and reliable CI signals. Guest: Eli Schleifer , co-founder and CEO of Trunk Eli Schleifer leads Trunk’s technical vision and product strategy, focused on closing the gap between AI-speed code generation and human-speed delivery by removing the bottlenecks that slow modern engineering teams.

Full transcript

50 min

Transcribed and scored by The B2B Podcast Index.

Welcome back to the Platform Engineering Podcast. I'm your host, Cory O'Daniel. My guest today has spent his career inside of some of the most demanding engineering organizations on the planet - Microsoft, YouTube and Uber. He's currently the co-founder and CEO of Trunk, Eli Schleifer welcome to the show.

Thanks for having me, Cory. So I'm very excited for you to be here. It is a collision of my worlds. So I am very much a TDD engineer.

I love tests, I love green pipelines and I feel like honestly you guys are kind of at a very interesting place in how our engineering world is changing, so very excited to talk to you today about CI, keeping greens build... but would love to just learn a little bit about your background, like how you ended up in the CI space and... yeah, we'll go from there. I spent obviously a bunch of my career in big tech in the early days, did eleven years at Microsoft and then had a startup that then ended up selling to Google.

Inside Google, I was just blown away by how easy it was to build software. The Google three stack, the DevX experience in there was just bar none. And then I went to Uber, worked on self driving cars at Uber ATG. We had, you know, some six hundred, seven hundred of, probably at that time, the most expensive engineers you could hire - these self driving engineers.

And we couldn't get code into the repo. It was just impossible. CI was completely unreliable. We had a merge queue that would back up for days.

If you queued something on Thursday, it would be like, "By Sunday that thing's going to clear and then I can get back to work." And it's kind of funny because in my early days at Microsoft I worked on Windows Mobile, before even the iPhone came out, and we would have days when you'd wake up and the build would be broken and be like, "Sorry guys, no simulator to run today, nothing to do." And I was like, this is crazy that we're still... many years later...

still hitting these problems. We're just like, "We can't get work done. This is so inefficient." And engineers want to build, they just want to land their code and go on to the next thing.

And it's really hard to do that. So that's why we started Trunk. It was like, "Can we build the Google developer experience that is just...that just works?"

Really, you just want the thing to work, you want it to give you signal when you need the signal and otherwise you want it to get out of your way. No one is like, "I wish CI was taking longer." It should just be instantaneous. That would be the perfec world.

So that's kind of like why we started Trunk. It's like, "Can we do that?" We started with our code quality product, but we quickly launched our merge queue product, which is the most performant merge queue on planet Earth right now. We can handle more pull requests than GitHub itself can handle.

So we basically maxed out the testing on that at the limits of the Ruby Monolith at GitHub. And then we've launched, you know, in the last year and a half, our flaky test product, which we built from scratch. Really to identify flakiness in CI and really looking, you know, at the broader spectrum of like, "What does it look like to manage tests for organizations that manage CI for organizations so that they can get to work?" And as you kind of hinted at, like, it's a crazy time to be doing this because the inner loop, where all that code generation is happening with Claude and Cursor and Copilot...

for people who still use Copilot... like, that's just generating a lot of PRs and a lot of code and the seams in the system, the outer loop where we're doing CI, where we're trying to actually test and verify things are working, that's falling apart. We are getting customers every single day just being like, "Our merge queue is backed up, it's totally posed, our tests are totally flaky, we can't get things through. The agents keep generating code, which is awesome, but we can't actually get into the system reliably."

And they come to us, and we're really looking to become that outer loop company. We want to keep engineering moving and basically make CI work. Yeah. So for folks who haven't worked with a merge queue before, could you tell us a little bit about what a merge queue is and what size teams are typically...

or what change velocity are teams typically starting to look towards using a merge queue? Great question. And really it probably historically would have been something like twenty, thirty engineers would need a merge queue. Now, if you're like a really productive team, that's like ten people, plus each of those is running a bunch of agents, you might need it already.

The way the merge queue works, it's basically doing predictive testing against your pull request. And why this matters is that you, Cory, are writing a pull request and you make some new code that's using a function called "foo" and I go in another pull request and I rename "foo." I rename "foo" to "bar". So the whole code base is now going to have "bar" and you're still calling "foo."

You push your stuff up - green CI. I push up my stuff green CI. If our stuff basically both lands, it's broken. Because you've just introduced code that's using this old function.

We call that a logical merge conflict. The silliest way to get around that is to make sure that every time Main changes, you just have to retest everything... all your active pull requests. That's insanely expensive.

Like, you never get any work done. So the industry basically started building merge queues. And what a merge queue does is it does predictive testing. So in a pipeline, as you might imagine, now my same pull request is in that pipeline and your pull request gets queued behind it.

So when your pull request goes into the merge queue, instead of testing against just the head of Main, you're testing against my changes, which renames and gets rid of the "foo" function. And now your code which is using that function will fail. Like, it'll fail to compile, it'll fail to test and you'll get an error, you'll get rejected from the queue, you'll be like, "Oh, I didn't know Eli was doing that. I'll fix it up," and no problem, you resubmit to the queue and everything works.

So instead of having a broken build, which basically means no one can get anything done because Main is host, instead you move to a model where you're protecting Main and Main is basically always ready to go. And in a world where you have really good CI, you could basically push directly from Main. Many organizations do it through the CD flow and that's really giving you, like, "I can change code really quickly and get it into the hands of customers." No one wants to change things faster than ever...

but AI so being having reliable merge queue in the place is important. Now, as you might imagine, flaky tests will basically bring a merge queue to a grinding halt. Oh yeah. If you're sitting in a queue...

and just like all queues and you learned in early computability around pipelining... if that head of that pipeline gets hosed, well, you're going to do a whole pipeline flush and you're going to have to start over again. We have anti-flake protection built into our system that actually can take advantage of that stacking behavior of a merge queue. Again, I submit that change.

I'm renaming a function, you're behind me and now you're also testing my renamed code. Now if my PR fails because of flakiness, well, we have another bite at that apple. Right behind me is your PR that also has my changes and your changes. So together, if that change passes, we know that I'm good to go also.

That it's probably just a flake. So that's one of the features we have with this anti-flake protection we have inside our merge queue that basically will keep things moving really quickly. So it's a very cool feature that you can kind of tune. And then we also have batching and we have dynamic parallelism.

So companies that basically understand...like one of the things that engineers hate about merge queues is like they'll make a change to a doc or the front end code, they're like, "Why am I sitting behind all this back end code? I'm not touching that at all. I'm just trying to change a color.

Like, this is nonsense." So we have a dynamically parallelizable merge queue that basically can use your build system like Bazel or Buck or Nx or your own globbing pattern of like, "Here's my front end code, here's my backend code." You give that to us and we can actually dynamically build parallel lanes for your merge queue in real time so that we get a fan in, fan out action as you're changing different parts of the code base. The net effect of all that super nerdy stuff is that the merge queue just works and you don't have these backed up merge queues that last for days.

Instead everything's just flowing through the system and we push massive numbers of pull requests through the system. And you also get this really cool graph that shows all the PRs and how they're interconnected. That's like where we kind of geek out about we know we see this. We had one customer, Fair basically is one of our longtime merge queue customers.

At some point they literally had twenty parallel lanes running inside their merge queue and none of the code interacting. And that just meant that everyone just... like it's as if they had their own merge queue, on their own, every one of those PRs, and it was just running behind the scenes for them. It was a super cool looking graph.

Yeah, I'm thinking about that. In our own code base we do a lot of domain driven design. So it seems like if you have very welldefinwell-definededdomainswithinyourmonolith,you couldprobably have evenalaneperdomain.So even teams aren't waiting behind each other if you're not working on the same product line effectively.

Absolutely. I think that as we go into deeper and deeper agentic engineering, the agents want to have full context. A monorepo is the best way for the agent to understand the totality of the system. And if you have multiple repos now you have to tell the agent, "Hey, all these repos interact with each other and they're going to deploy at different times."

That's kind of craziness. It's hard for humans to do it. It's why Google moved to a monorepo a long time ago. And that's why largest organizations gravitate towards that structure.

I think agents are going to push more and more organizations into a monorepo world. And when you're in a monorepo world then you really want a merge queue with dynamic parallelism, otherwise you get into that stuck single lane problem. Yeah, and I feel like that's... you know, we're a fairly small team.

We're, we're three engineers, you know, at my day job. Like we've leaned in pretty hard on AI development over the past six months or so. I was very skeptical last year but, you know, just... we started getting into it originally I was like, "Ah, it's hit or miss the output."

And then you know, we kind of all took some time over Christmas, like sat down, played with the tools, figured out like what worked, what didn't, how to tie it into our code base. And we've gotten to a point where we produce pretty good code. But now like we're already at this point... you know, you're saying like twenty, thirty engineers, you might need some things...

but like we already seeing it where it's just like I go in and I'm like... the hardest part of our job now is we go in and it's like there's fifteen, sixteen PRs. And you know we're one of those teams, like we try to keep our PR small still. Like we still like the idea of trunk based development, small short-lived branches and go.

So like we're not going in and be like, "Refactor the entire code base" and getting like a 20,000 line change. But even a three, four file change, a hundred lines, two hundred lines... when you have twenty or thirty o those stacked up and now you're thinking across... I'm thinking across a lot of changes, I have to think, "How does that affect that?"

It does seem like we are speeding up in our production of code, but we're also speeding up in a lot more engineering and cultural problems in our space that are just cropping up really, really quickly. And I feel like CI and the CI pipeline is going to be just a huge bottleneck for teams as they start to adopt this. And I feel like it's one of the things that is probably going to trip a lot of teams up in adopting AI, because they don't have great CI, they've got flaky tests, they can't trust the output, and now they have AI that's flakin tests faster.

Like, how are you seeing teams that are starting to adopt AI fi into this world? Yeah, I mean, I think that those teams are now all reaching out to us and becoming our customers. If your CI is not reliable, you can't keep up with velocity. You want to, like, you're so excited, "Oh, look at all these PRs, I want to land that.

The product will move forward like what it used to take a month in week." But if I can't land those PRs, then I can't actually make progress. So you need to basically be looking at your entire testing infrastructure - How quickly are things running? How much flakiness is in the system?

Do these things have to be babysat? We always talked about... without a flaky test system in place, what you're basically going to do is have a flake, everyone's going to keep track in their head and their minds what are the flakes, and then they're going to go and rerun it inside CI. And that's really ridiculous.

It just slows things down. Engineering time is always at a precious minimum. So having to wait for a CI to run, even if that's twenty minutes, is a big pain. But having to do that agentically, you're going to talk about massive amount of burn.

One of the things we actually have built and are seeing really cool opportunities with is the ability to use the information we collect historically around all your tests to actually identify flakes to the agent. We have a quarantining system which is critical to keep CI green. So you push up your test, the tests that are flaky, you could basically quarantine. We'll basically ignore those failures.

So you only are being blocked by CI when the red is truly red and trying to give you signal, because that's what you want. At the end of the day, when you're running a test, you want to be like, "Did I break something?" Because you don't want to break something, you want to fix it. But when a test fails, because that's a timeout and it fails 5% of the time that times out, you're like, "What am I going to do with that?"

I'm just going to run it again and hope that this time it's fast enough. And that's very silly. So, with our quarantine system, you can basically ignore all of those kind of well known flaky situations. Now, on top of that, when your agents are doing this, they really need to know, "Did this ever fail in this way?

Did I break something or is this a known problem?" So we have... you can basically pull into Cursor or into Claude... with some of our skills, you can basically, on our MCP integration, you can pull in all of your test historical information.

So when the agent is looking at CI results and saying, "Oh, this thing failed, do I need to fix something?" We don't want it to go and try to fix something flaky that's been in the system for six years because it's just going to be chasing its tail, like, and also create a PR that, who knows, is completely off track of what it's trying to do. What you want it to be like "Oh, this thing has always been flaky. I'm going to ignore that and not try to fix it and move forward."

It's funny that you say that. We program in Elixir and the Elixir test suite, just the one that comes with it, is amazing and has this really great ability to run tests asynchronously. So I think our entire code base is like nineteen hundred tests or something like that. And it runs in seventeen seconds.

It is fast. But we do have this one flaky test that crops up all the time. And it is around like our... so we're on like the CD side of the world...

so like it's when the deployments are happening and, you know, it's a timer related thing. And so it just fails every once in a while. And I had Claude working on something overnight, left, came back in the morning and there was just hundreds of changes. And I was like, "Whoa, what in the holy hell is this?"

And I'm like going through it, I'm like, "Wait a second, it's trying to fix this flaky test." And what it did was it went to every single test file that we had and it switched the test framework to running synchronously instead of asynchronously. And it's like... it 100% fixed the test, but our test that also ran in seventeen seconds now took like fifteen minutes.

And I'm like, "Well, I'd rather just see that flaky test every once in a while." Cory, I have a solution for you. You've got data? We'll get you quarantined, no problem.

You guys only have three engineers. You're within our free tier. There you go. So for folks that are doing CI, maybe they're using GitHub Actions today...

How does Trunk fit into your world if you're using GitHub Actions or you're using GitLab? How do you tie in Trunk into CI? We're totally agnostic to what your CI system is. So basically you're still running your tests...

I don't know what the Elixir command is, but you're running your Vitest, you're running your Bazel test, you're running your Jest, whatever it might be, your Playwright test for sure. After that, those tools can all be written... can all be configured to basically output JUnit XML, which is the lingua franca of test output, test results. Which is funny because it's like a pseudo spec developed by IBM that is the loosest XML spec you can imagine.

It's a travesty that the industry uses this, but this is what it is. The next step, you're going to call our analytics CLI. That CLI is basically... it's a Rust based CLI that will take that JUnit xml, it'll post that information up to our service.

It will then also check all the failing tests that were reported in that JUnit output. If all those tests are flaky and quarantined, we'll actually go and change that exit code from a failure to a success code... if you're using quarantining. Otherwise we'll just start tracking the information.

We'll understand, "Here's what happened in this test" and we'll print out inline NCI and also up in our SaaS dashboard, "Here's what happened in this particular PR for your testing." Yeah, so how do you guys handle... let's say it's a flaky test, but let's say we actually just... somebody broke it, like three changes up from me.

Like, it's just broken now. How does the quarantining detect the difference between something that's flaking and something that's actually busted now? Dockeractuallydoiswe'lllookatthefailurereasonsthatareinsidetheJUnitXMLandthenwe'llactuallydoAIembeddingstounderstand thedistancebetweenthatfailuretypeandanypreviouslyseenfailures.So wecallthisafailurefingerprint.

Andyou canbasically identifywhenthingsarebeingbrokeninanewway,whichisalsointerestingtotheagenticsolutioncaseaswell.Sosomeonegoesin,thisisagiantbigproblem.It'slikeabrokenwindowsituation.Like,"Oh,thistestisflaky.

Ignoreit."No,youactuallyjustbrokethistest.Sobecausewehavethisfingerprintingdesign,whatyoucanactuallysayis,"Thisthingwasfailingwiththistimeoutforthelasteightyears,butnowIhadthisnewfailurethatwasintroducedforthisexactcommit."Icantellyouexactlywhenithappenedandyoucangoandfixit.

Sothatislikeacriticalpiecetothissystemthatwe'vebuiltwherewe'redoinganalysisofwhatyou'veactuallyuploadedtous,doingthisfingerprintinganalysisandshowingyouthediscrete waysinwhichatestfails.Becauseifyou'rean engineerthat'staskedoranagentthat's taskedwithgoingtofixaflakytest,you wanttosay,"Whatarethedifferentwaysit'sflaky?"Becauseit'snotallflakythesameway.There'sasimilarparallelproblemoflikeanorganizationisusingsomedockerpulltodosometestsandlike300testsrelyonthisdockerpull andthedoctortimesoutorthere'saratelimitandallofasuddenallthosetestsaregoingtobemarkedasfailing.

Thenyou'llretryitandit'llpass.You'llbelike,oh,thesetestsareallflaky.That'sobviouslynotthecase.Itwasan infrastructureproblem.

Wealsohavean Docker pull to do some tests and like 300 tests rely on this Docker pull. And the Docker times out or there's a rate limit and all of a sudden all those tests are going to be marked as Wecanalertyou tothatsituationaswellandprotectthetestfrombeing "Oh, these tests are all flaky." ThatSo w Oh, that's very cool. It's funny, like, our flakiest test is also probably like the one that if it actually shipped broken, would piss people off the most because it's the actual deployment mechanism.

So it's like whenever I see it fail, I'm like, "Please for the love of God, tell me that's the flakiness." It's like, I know the signature. Like you're saying the fingerprint. The second I see it, I'm just like, "Rerun, whatever.

I know what it is. That's not a problem." Yeah, we're going to fix that for you, Cory. We're going to get you integrated.

We can do the whole thing in an hour, and never see that thing again unless it's actually broken in a new way. I would love that. Love a green build. So for folks that...

I know, some of the audience is a bit AI skeptical and I feel like... there's like two groups of people that I feel like are talking at each other and not talking with each other right now. There's like the far booster side of the world that's like, "Definitely figured it out." Which I feel like I'm starting to move towards that booster world.

But I'm also concerned about where AI is going. Like, as an industry kind of freaks me out. And then there's like the people that's like, "This stuff's just... this sucks and everybody's hallucinating over here."

Like, what are you starting to see team wise? Like, are you seeing teams succeed with their AI initiatives? I feel like the teams that do the best around CI and code quality tend to see the most success with generative output. But what are you guys seeing as somebody who works right around code quality?

Cursor theheart ofthematter,the future is AI driven development. AI is just much better and faster at writing code than humans are, at the end of the day. What we're seeing is massive success, especially of late. I think if you were using early versions of Copilot, you'd be like, "Okay, you can sort of do a for loop, good job, but sort of just as good as autocomplete and not really getting me there."

Now what I would say to anyone who's AI skeptic, I'd be like, "Go open up cursor or Claude on the latest model, give it a prompt and see where it goes for you." Right? And I would say the places where it's best, or places where it has the most context window, is like... Doing UI changes used to be such a pain.

You could have frontend engineers who'd be like, "I'm really good at CSS and making things pop up and bubbles and all stuff." That is just so easy now, it's my favorite thing too, because I never wanted to learn any of that. The last time I coded frontend was like HTML 2.0, whatever it was.

Now I'm like, "Okay, make a pop up, make the combo box, make it do this thing." And then it's like, "Wow, it worked great, let's go." I think that that is so exciting because you can deliver really cool, beautiful, polished UI that in the past would have taken you forever and you would have just agonized over it. Ops teams, you're probably used to doing all the heavy lifting when it comes to infrastructure as code wrangling root modules, CI/CD scripts and Terraform, just to keep things moving along.

What if your developers could just diagram what they want and you still got all the control and visibility you need? teams, you'reprobablyusedtodoingalltheheavyliftingwhenitcomestoinfrastructureascodewrangling,root modules,CICDscriptsandTerraform. Just tokeepthingsmovingalong,whatif yourdeveloperscould justdiagramwhattheywantand youstill gotallthecontrolandvisibilityyou need? That'sexactlywhatMassDriverdoes.

Ops teamsupload yourtrustedinfrastructure ascodemodules toourregistry. Your developers,theydon'thaveto touchTerraform,buildrootmodules,oreven copyasinglelineofCICDscripts. Theyjustdiagramtheir cloudinfrastructure.MassDriverpulls themodulesanddeploysexactlywhat's ontheircanvas.

Theresult,it'sstill managedas code, butwithcompleteaudittrails, rollbacks,previewenvironmentsand cost. more at Massdriver.cloud.StartmakingInfrastructureasCodesimplerwithMassdriver.

LearnmoreatMassdriver.cloud. Balsamiqrecently,or Ihadthisidearecentlyforsomethinginourproductthat'sbeenchallengingfor customers.AndIsatdownandIdrewuplike,IdowhatIalwaysdo.

Like,IdoveryBalsamiq-esque... Idon'tknowifyou'refamiliarwiththeBalsamiq...IjustdoveryBalsamiq-esquedesigns,usuallybecauseIdon'tlikepeopletothinkaboutcolors.AndIwaslike,"Thisismyidea."

AndIshowedittooneofmyteammatesandhe'slike,"Idon'tgetit,Idon'tgetit."Like,hedidn'tgetit,andI'mlike,"Damn,Ireally thinkthiswouldhelp."AndI'm likesittingherelookingatthisdrawingandhe'sjustlike,"Yeah,Idon'tknow."He'slike,"Itseemsalittleconvoluted."

AndthenIsatdownwithClaudeandIwaslike...Ihaven'tdoneHTMLinalongtimeeither.Idon'tknow...IknowjQuery.

rememberwhenjQuerycameout.Like,that'swhereI'mat,right?ButIamdefinitelynotatypescriptninja.I'mnotareact10xrockstar,anyofthatstuff.

AndsoI'mlike,Ireallywanttogetthisdesignslike,likewhatitwouldlooklikeinoursiteandthengethisopinion,right?SothenIjust,Iliterallyjusttook.IdumpedtheHTMLfromthepagewe'reon,like,"Ireallywanttogetthisdesign...likewhatjustdrew,putitinhere.

Andthen like,I'mgoingtoopinion."So then Iliterallyjusttook...Ihimascreenshotofitandhe'slike,ohyeah.He'slike,itinCursoranwaslike,"Hey,takelikebalsamiqthingsomethingoutthat'sanmvp.

Andit'slike,Iknowthecodemightbedogshitunderneathit,especiallygivenitput drew, put it in here." And goingtoiterate onittolikegetitlike howit isin mymind.AndthenI it is in my mind. And then I sent him a screenshot of it and he's like, "Oh yeah.

He's like, "That would absolutely work." And it was just like... the ability to get something out that's an mvp. It's like, I know the code might be dog shit underneath it, especially given it was just working from an HTML dump, like not our actual like full code base.

But like, that was an idea that will greatly benefit customers that I could not have communicated without it. And I was at a dead end. I drew it like four times. It's like every time I showed it to him it was like, "I don't get it."

And I was like, "Arghhhhah." And now it's something that's going into production and like we've shown it to people and they're like, "Oh, this makes way more sense than what you guys have today." And we're like, "That's fantastic." I feel like that's maybe like a place where people can start, because it is intimidating to go into your giant code base where you have fifteen flaky tests and so many red herrings that AI can chase.

But like to sit down and it's like, "Okay, let's think about this problem that we haven't been able to solve " a" you'resamethingmyself.LikewhenweneededaChromeextensiontobebuilt,Ijustlikestartedwithlike,"Youknow,allright,let'sgobuilda chromeextension.Brandnewcode."Likewedidthewholethingintwodaysand itwassopolished.

Thatwouldhavetakenusamonth.Wewouldneverhavescheduledthetimebecauseitwouldbenotworthitfortheamountofeffortitwouldtaketomaintain.AndwedidtheentirethinginCursorwith V0aswell.Icallthis thing"CodeFirstEngineering."

Iwrotepostaboutitandit'sreallylikethecostofcodeusedtobesuperhigh.Generatingcodehigh.Generatingcodewasexpensive-crazylet'snotstartwritinglike,"Let'sknowwhatwe'redoing.Anduntilweknowwhatwe'redoing."

Andwe'dbelike,"Let'ssitdownwithdownwithaPMandfigureoutifthisthingisgoingtobeviable.Andwe'lldoscreenshotsandifthisthingisgoingtobeviable,andwe'llgone.Nowit'slikeyoucouldbelike,hey,andallthatstuff."Andnowthat'sallgone.

Nowyoucouldbelike,"Hey,I haveadrawingfromBalsamiq,orFirstwetelllikeascreenshotthatlikeV0generated,ofyouisnowaCursor,engineer.intotheproduct."CodeFirst.Wetellourteamnow,tellallofnotengineers,Everyoneofaproductengineer.

Youbuildproduct.Thatwasalwaysourjobasengineers.Likewe'resupposedtobuildthingslikenoonecareswhatthecodeislookslikeunderthehood.Theywanttomakesureitworks,it'sreliable,and ourengineers,"Everyoneofyouis nowa product engineer.

You're notafrontendengineer,you'renot afullstackengineer,you're notabackend engineer, You'reaproductengineer.You "Every one of you is now a product engineer. You're not a frontend engineer, you're not a full stack engineer, you're not a backend engineer, you're a product engineer. You build product."

That was always our job as engineers, like we're supposed to build things. Like no one cares what the code is looks like under the hood. They want to make sure it works, it's reliable, and also that it's doing the thing that they want it t do. And I think that that is really like what product engineering is all about.

Now everyone's a product engineer because now you have a PM and you have a designer in your back pocket. You can go to all these tools, they'll do really great work for you. And when we say Code First we say like, "Code first, let's make sure it works and make sure it makes sense." You showed that to your peer, he's like, "Oh, I get this."

Great, now we can go and p..... it, lI 'dj"L". You know, like where you guys sit today... so, like, let's say teams...

because the other place I'm kind of seeing people get choked up, especially in our org, is like the actual... and I feel like some people in this space are saying that they're just not doing this anymore.... which, if you aren't, congratulations, I guess... but like, we still have to review the code, right?

The build is green. That's fantastic. The build is green. Like, I ran my Sobelow security scanner or whatever, right?

It ran all that stuff, that's all green, but is it green to me and you? It's like, we can catch some of that stuff in linters, we can catch some of it in style guides, but these are two quality pieces many teams still miss. They have their style guide in the Readme and they're like, "Hey, do it this way." But nothing to enforce it.

There's teams that have... let's say they've nailed all the stuff in CI. I've got my linters, I've got my style guides. It's all automated, my formatting, my quality, my security scanning, all that stuff.

But then at the end it's like, "Okay, but there's still this code that someone's going to maintain." Whether it has a heartbeat or, you know, a cpu, who knows? How are people dealing with the review process on the other side of the CI green building? And are you seeing people still get hung up there in just like the manual review of the code, or are you seeing people be a bit more, I guess, laissez faire about what they're merging?

I think that that's a really interesting point where, we'll have to as an industry figure out what do we do as humans to be reviewing this code. And I think that there definitely are cases where like the browser extension that this thing generated, in this case, every single GraphQL call got its own file. I was like, "That's really weird. Why are you doing that?"

So I was like, "Don't do that. Be normal. Do it like you would normally do." And I think that when we were making GraphQL changes to our frontend service, I sent it to our GraphQL expert and he was like, "Well, the best idiomatic way to do this is this."

And then he did the code review and he was like, "We should structure it this way." He left those comments, and what I just took to those comments was, "Hey, Cursor, look at all those comments, address them and then fix it up." And that was like... the part where like I had to do that job, that was silly.

Like he should just like write those comments and then an agent should be like, "Great comments, I'll go fix them and land it." Like, Eli doesn't care the way that it's done, we just make sure that it's done right. I'm not on the front line of engineering right now, so I don't have the context of like, "What should this look like?" I think that you need people to still have taste in how code is designed to make sure it works right.

But certain things, I don't want code reviewed. If it's a whole 10,000 lines of frontend code and you have integration tests to make sure everything is clickable and the right UI changes at the right time, ship it. Don't look at it because it's just going to be tossed out or changed the next day. Do you really want to look at what the CSS looks like here?

You need to care. We've always done this... it's like, let's care about the things that matter most. Like in the old days it would be like a nit.

Like you spelled this thing wrong in this comment, and if I update that, it's going to run CI again. Like, that's bonkers. Like you're going to spend 30 minutes rerunning test... well, in your case, 17 seconds rerunning test...

because I changed the comment to be spelled correctly. Like, I think you knew what I was talking about. So I think that the more we get... just like we never...

we don't look at assembly because we just know the compiler did the right thing... for the most part, there have been compiler bugs, but for the most part it just works. I think it'll be the same thing with the code we generate today. We're going to move further and further away from reviewing the code and just reviewing what the product does, because at the end of it, that matters.

Now, AI is really off in certain cases of building things that are scalable. You want to lay out this data in a reasonable way that's actually going to work when we actually have a billion users or a million users or... you know, we consume billions of test runs every single day. Is that a thing that I would just trust AI to correct the schema correctly?

No way. Like it's not there yet, it doesn't understand the complexities. You just toknowwhere. where you need to bring in your designthinking asalargesystems engineerto makethisthingwork.

Butthe code thing work. But thdoalotlessofthatreviewinginthefuturebecause we,therejustwon't in the future because we... there just won't be time and also it's not the best us React that'sinteresting.Icadefinitelystarttofeelthatincertaipartsofthecodebase.

It'sfunny,frontendengineeringisalwaysaplacewhereIfeellikethere'sjustsomuchgoingonthere,andIfeellikepeopledogonfrontendengineerssometimes.It'slike,"Isitengineeringornot?"It'slike,"Yeah,it'sfuckingengineering.We'rputtingtogetherthingstotellmachinestodothings.

It'sengineeringwork."Butatthesametime,like...Idon'twanttousethe wordthrowaway...Iknowthat'sthewordpeopleareusingaroundlike,"Yeah,that'sourthrowawaycodeversusourstablecode."

Idon'tnecessarilylovethattermbecauselikepeopleareworkingonthatandliketoknowthatyouroutputislikethethrowawaything...butyouknow,what'sreallyinterestingaboutlikeCSSandHTMLislike,it'snotyourbusinesslogic.YouknowwhatImean?Like,it'snotyourdomain.

It'snotthethingthatmakesyoumoney,right?It'sthelogicthatyou'vebuilt,theabstractionthatyou'vecreated,and thatsoftwareasaservicethatyou'vebuilt-that'sthethingthatmakesyoumoney.Likethisisapresentationlayerandteamswebuiltourlandingpageoriginally,Italkedaboutthisafewepisodesback,butdidn'tgolikeintonittygrittydetails,butlikewhenwebuiltourlandingpageoriginally,wespentfuckinglike...firstwespent$25onlikeareacttemplateoff ofsomethingwhenwefirstlaunched,anditwasugly.

Andthenwehiredmarketingandbrandingteamandweblewtwenty,thirtygrandonourhomepage,anditwasacatastrophicpieceofshit.Like,itwasjust...itwasgnarly,CSSjustgnarledeverywhere.Andweregenerated...

like,ourcurrentsitewasfourbucks.Fourbucks.Isatdownandhadthebestdayofmylife.Iwaslike,"Oursitelookslikeshit.

I'mgoingtomakeitlookgreattoday."Andwesatdown,Iburned$4,Igeneratedlike20versionsofit.AndwegottoapointwhereIwasjustfocusingonwhattheaestheticwaslike.WegotittoapointandthenwenttomyfrontendteamandIwaslike,"Hey,whatdowehavetohaveinplacethatwouldmakeyouhappyaboutacodebase?

Thiswasthenewfrontend.Whatwouldmakeyouhappy?"Hestartedputtingallofhisdetailsinthere,likewhathewouldwant.It'slikewefedthatUIourexistingsitemapandthenhisconstraintsandwegotjustthetightest,mostwelldesigned...

likeReactcomponentsforeverything,liketheCSSisextremelywellscopedinsteadofjustkindofallovertheplace.It'slike,wewouldhaveneverdonethat,right?Andatthesametakeonthis?Like,youcode.

Likeitcostme$4andmescrewingaroundthere,likewhathewouldwant.It'slikewefedthatUIourexistingsitemapandthenhisconstraintsandwegotjustthetightest,mostwelldesigned...likeReactcomponentsforeverything,liketheCSSisextremelywellscopedinsteadofjustkindofallovertheplace.It'slike,wewouldhaveneverdonethat,right?

Andatthesametimeit'sthrowawaycode.Likeitcostme$4andmescrewingaroundtogetit.Andlikethat'sthereality,Ithinkoflikewherealotofournon-criticalbusinesslinecodeisgoingtoendup.It'slikeyou'resittingthereandit.

Andlikeatyouremailsystemthatyoubuilt,builtinRailsthree,Ithink,oflikewherealotof ournon-critical businesslinecodeisgoing togoingtowork.It'slike,dude,pointitatthatthingthatyoudon'tcareaboutanymore andsay lik.... "T","W","." A"eEh"."

L...,i, L"I.L?" L."

, Yeah, yeah, exactly. You can also build things that are just so much prettier than you would have ever been able to do or ever be able to justify. I built like a tool inside our homepage where, like on a marketing site, you like right clic on our logo and download our assets. Because I was always like so annoyed...

like, I need this asset, I need it to be either the word mark or the trademark, and I need it to basically show up at different sizes. And I just like... I saw another company that had a nice one. I was like...

screenshot, "Make this thing work like this." And it was like three hours of just toying with it and then it was amazing. I was like, "Is this necessary? No."

But it was a really fun AI experiment of how well can you just build something that can do these kind of really cool things really quickly. And I was like, "This is so cool." And it's why I think it is the most fun time ever to be a builder. Youwant to build something?

You have this army of tools that can go and help you build faster than you ever could before. It's like... the equivalent of like carpentry, like if you wanted to do woodworking like before power tools - man, that's slow. That's just a slow way to do.

Now every one of us has like a full wood shop available to us - you got a table saw, you got a joiner, you got all these pieces. Everything ilikepoweredby240v andyou'rerockingandrolling, you'rejustpoundingthroughitandyoucanmake you know,furnitureat theendofthedaymuch faster. Samething. Like faster.

Same thing, like we can make software much faster. And it's so fun because the things that used to trip us up, especially like... we had one of our engineers wa like, "There's this like weird thing that fails all the time, and like it's not a big deal, but like wanted to fix it." Sent it to Claude on Monday, and Claude di like a four hour research project and was like, "Oh.

The reason why is because like there's this weird timing race condition the way they were doing xyz, I fixed it now, it'll never happen again." You would never go fix that. Like that's just not where the humans time... be like, "Can we make this never happen?

It only happens like 0.1% of the time." But now it's fixed and didn't cost that engineer any time. Just like, "O", , i ?

thinkwe're...tomeit'slikean extremelyfascinatingtimeandterrifying.Right?There'sthesemoments whereit'slike,there'stwosidestothecoinwhereyou seethesepeoplethataren'tengineersandthey'rebuildingstuffandthey'resuperexcitedandpeoplelike,"Eh,they'refulloninpsychosis."

Andit'slike,"Nothey'renot,they'reexcite thatthey'vehad thisideaforeverthatthey'veneverbeenabletocommunicatetoanotherpersoneffectivelytogetitcreatedandlikenowitexists."That'sverycool.Thescarypartofthatis...we'veseenlike thePostgresrootpasswordjustlikesittinginaJavaScriptfileandlikeallthatstuff.

Andsolike,that'sthepartofitthat'svery,veryterrifying.Especiallyseeing like thefullblown...I'mprobablygoingtogetshot bytheYCMafiaforthis...butlikethefullblown,likeYCpsychosisishappeningandlikethecurrentbatch.

I'mnotsureifyoucaughtupwiththatonline,butit'sexcitingbecauselikestuffishappening.It'slikethisisoneofthosepointswhereit'slike,Ifeellikewewerelike,"Well,what'sgoingtohappentoengineering?"Andit'slike,well,alotmorebusinessesaregoingtobecreatedbecausepeoplethatcouldn'tcreatethisstuffbeforecan,andsomeone'sgoingtohavetosupportthisandsecureiteventually.NowIthinkthatourliveswillbeabitmorefraughtforthosepeoplethatarelike,"Okay,I'mgoing togogetmyjobat,youknow,thecompanythat'sbeenAIdevelopedcompletelynativelybytheCEOfortwoyears,it'smaking $4millionayear."

Yougetintheredayoneandyou'relike,"Holyshit.It's7.6millionlinesofcodeand it'sallsomehowinCSS."Ithinktheothersideofit'sgoingtobealittlescary,butIdon'tthinkit'snecessarily goingtocollapsejobs.

Ithinkwe'regoingtoseeteamsprobablybecomealotsmaller,right?Andalotmoreteams.Butthethingthat'sreallycurious tomewithyourideaofthislikeproductengineer,thatyourteam'skindofbecoming,is,howdowetrainpeoplein thisnewworld?Whereit'slike,youknow,folkslikeyouandIhavebeenintheindustryfortwenty yearslike,howmuchlongerarewegoingtobeherethatwecan,youknow,apprenticeandteachourjuniorsonlike,howtothinkaboutthesesystemsholisticallyandhowtothinkaboutthem,youknow,scalewiseandformillionsorbillionsofusers?

like,howdowe trainfuturegenerationsofengineerstobeabletoworkwiththisoutputwhentheymaynothaveseencomplexsystemsbesideswalkingintothat7millionlineCSScodebase. anhadthisconversationmanytimeswithdifferentpeople.It'skindoflike... itremindsmeoflikewhat'shappenedlikein thenuclearpowerindustryintheUnitedStateslikefor,youknow,adecade.

Webuiltatonofnuclearpowerplantsand wetraineduppeopleandthentheyallslowlyagedoutofthesystemandbasicallynomorepeoplewereevertrainedin that.Andbelike,"Okay,howdo youbuildanotherwholegenerationofnuclearpowerreactorswithoutallthepeoplethatknowhowthisworks?"Andnowwehaveabunchofpeoplethat,youknow,maybetheygraduatedwithcomputersciencedegrees,butaretheyproductengineers?Dotheyknowhowtobuild?

They probablyhaven'tseenlargesystemsbecauseyou'renotgoingtobeexposedtothat.AndIthinkit'llmovereallylikeintoanapprenticesystem.thinkfor...youknow,inthelastdecadetherewasthislikeshift frompeoplegettingMBAstogettingCSdegrees andlike,"Oh,that'sagoodwaytomakea lotofmoney.

Thosearehighpaidjobs."Andit'slikeifyougotintoitjustbecauseitwasahighpaidjob,youdidn'tactuallyhavelikethatfireinsideofyoutobuild-it'sgoing tobeahardroad,Ithink,oryougottofigureouthowtofindthatfire.Becauseit'snotaboutlike,"Oh,IhaveaskillthatIcanwriteAssembly,IhaveaskillthatIcan writeC++."It'slikeyouneedtohaveaskillwhereyou'relike,"Iloveto buildthingsandIwanttosee themcometofruition.

Ithinkholisticallyaboutawholesystem,abouthodoIconnectallthesepiecestogethertomakesomethingreallycoolthatpeoplewillcareabout."Andthosearetheproduct engineersthatwillcontinuetohavegreatemployment.Iwon'tbesupercritical,butIdon'tknowhowyoufilltheskillsgapoflike,"Oh,you'veneverseenaweird situationlikethis."Youdon't.

Like,theotherweekIwasdebuggingthesamechromeextensionandtherewasasituationwherebasicallywhenourcookiewasbasicallybeingusedto callourGraphQLsystem,likeiftheorganizationhadn'tbeenspecifiedinthistokenthatitwouldn'tknowwhattodo andtheGraphQLquerieswouldfail. Andthecode...Cursorwaslike,"Ohno,we'regood,everythingshouldwork."AndIwaslike,"It'snotworking."

And itwas becauseinoursigninflow,afteryousignin,ifyou'reinmultipleorganizations,youhavetoclickonebecausethatcookiewas thenbeingusedwiththatorganization. IDwasstoredinthecookieaswell. AndIwaslike,"IknowthisbecauseIbasicallyintuitedwhatwasactuallyhappeningandthoughtthroughthewholeflow."Ittookmetwohours.

Iwaslike, "Oh,that'sthething."Engineershavealwayshadthatexperience.Like,"Oh,thatwasdumb,Ifigureditout."MaybeonedayAIfiguresthatout.

Butthat kindoflike,"Icandebugandfigurethroughasystem.Iunderstandhowallthepieceswork.Sowhenit'snotworking,Icankindofpointyouintherightdirection."That'slikealearnedskill.

Andthat'sgoingtobelikeapprenticeship.It'sgoingtobeinternshipatbiggercodingshopstounderstandhowthesethingsallworkandhavinganintuitionaroundbeingabuilder.AndIthinkthebuilderswillhavejobsandthepeoplewhoarelike,I,I'macoder.That'sgoingtobeasmallerandsmallerpartofthepieforsoftwareengineering.

,i, " "r"- t AutocompleteIthinkso,forsure.Ithinkhonestly,capitalis goingtodowhatcapitaldoes.AndIthinkorganizationsaregoingtostarttoseeAI nativeteams shippingquicklyandthey'regoingtowant toemulatethat.And thecatchwillbe,iswhether theycan orcannot.

Thatisan unknown.Just likemanyorgstriedtomove to thecloud andthey'resittinghalfwayinbetween hereandthere.But,yeah,Ithinkit's goingto be prettycritical.I almost wonder...

I knowoverthepasttwenty yearsorso,there'salwaysbeentalksaboutwhat'sthe differencebetweenadeveloperandanengineer. Somestatesandsome countries,youcan'tcallusengineersbecausewe'renot,technically.I'm almostcurious...ifapprenticeshipwillbeit, isthatit?

Arecompaniesgoingtomakegood apprenticeship programs?Whatisthatgoing tolooklike?Everybody's architecturesaresodifferent.Arewegoing toseethecloudofferingsstarttocollapseandnotofferasmuchstuffso wecan starttostandardizeonsomewaysofrunningsystems?

Or,like,dowestarttosee licensureactuallybeimportant?Becausewe're makingthings, security is important,seeingallthestuff.Team PCPisouttheredestroyingrightnow,left andright... like,wearemaking alotofsoftwarevery,veryquickly andlike,wegottofigureouthowto securethesethingsand thinkaboutthesesystemsbetter.

AndIfeellikethat's thepartthat...onthe othersideofthe veryexcitingpart... likethatisthglaringholethatisvery,veryconcerningtome.Ifeellikethat'sthepartinthemiddlethatpeoplearen'ttalkingabout.

It'slike,Icanseethe valueofthis.Ican'timaginegoingbacktowritingcodeeveragain.Like,Ilikethewaythatwedoitnow.IStillwritecode-it'sverymuch likeharnessesquethingstokindofcontroltheAIandAIandwhereit'sgoing.

Istilllikereadingthecodeandthinking aboutthesystem.Butliketositthereandcludgethroughthetyping-like,noway,man. I'mnothittingthattab.Autocomplet,autocompletetheentirethingforme.

Butthisotherside,Ithinkisgoingtobevery,verycritica you'veIthinkyou're right because I think... just the other day a CEO of a company called me up... and it was all a vibe coded thing. Built the entire thing with first Lovable, then Replit, then Gemini Cloud.

And he was like going through the paces. And he was like, "I just got my first customer, how should I celebrate?" And I was like, "Well, you should get pager duty set up so you have an on call rotation, you should write some Playwright tests, you make sure that thing stays up and does what you want it to do because you didn't write any of that stuff yet." And he didn't know what those things were, but obviously he taught himself all the way through to understand how these systems interact.

I was like, "Go ask Claude to explain to you what these things are." And he's like, "Oh yeah. I asked Claude and Claude said you're right about the next things." So I'm like, "Yeah, these are the things yougot got to do."

Make sure... when you're building a system and people are paying you money, they definitely want the thing to want togetlikeacallto belike,hey,thisthingbrokeandyoudidn'teven "Hey, this thing broke and get didn't even realize it." So that's where we get into testing, automation of testing, making sure. If you have good surface area coverage of testing, it's not going to matter that much under the hood.

Now, of course, I was also like, "You need to make sure that none of your keys are exposed." So you just put into your Claude prompt, "Make sure none of my keys are ever exposed." And you should run these other linters and you know, static analyzers to make sure that none of those things ever get posted. Because, at the end of the day, if you h"" Yeah.

Awesome. Well, I know we're coming up on time. Eli, thank you so much for coming on the show. Where can people find you online?

You can find us at trunk.io, I'm sure we'll have that in the hyperlink somewhere. Love to talk to everyone who's like kind of exploring this new space and seeing their outer loop being crushed by the inner loop. We love hearing those stories.

All of our starting customer calls are now always like, "Do you guys have flaky tests or backed up merge queue?" And generally they're all like, "Both." We've migrated a lot of people onto our merge queue that are just seeing a bunch of creaky pains from other companies merge queues, like GitHub. And we have a lot of fun bringing them onto a more performance solution and getting them...

you know, setting them straight. Like we just finished onboarding a couple months ago, Brex, onto our flaky test system. They just didn't... they knew they had flaky tests, they didn't know the scale of the problem.

It was, you know, grinding engineering to a halt and we've fixed that up for them. They're now like able to just fly. And we'll have a case study on that coming out soon. It's just a very exciting time to be building and enabling teams to build faster than ever before.

Like we used to have a roadmap. The roadmap is always a dream, like, "Oh, wouldn't this be so cool? But you know, we'll never get to it." And now...

like I posted the other day, I'm like, "We're going to hit like end of the roadmap. We're going to hit the end of the road and have to add a new road on because we can actually hit the things that we say we're going to do." And that makes our customers happy. And it makes me happier because the team is just delivering at a...

we're a smaller team than we ever were, we're now only, you know, sixteen people, eleven engineers total, if I include myself... but we're building two to three times the amount of code than we ever did before. And not just code... really, just product.

We're building like two to three times as much product. Big problem to solve and we're really enjoying the whole thing. Yeah, I think one of the things right there that I would love to just put a little note on for people that are still skeptical. You may have just heard that and thought they're creating two to three times the amount of code.

That means you have a lot of code to review. But again, that's also... if I spent six to eight hours of my day working on that feature and typing it, and it is 900 lines of code. Like, everybody on the team can consume that in much less than the eight hours I took to spend it.

And I think that's one of the parts that, like, people are kind of missing because they're way over here, and they're not giving it a try, they're not experimenting with it. And they're hearing the people way over here being like, "I merged 20,000 lines of code today." It's like that person isn't, first off - they're merging a thousand here or there, unless they're refactoring their node modules. But, like, even the people that are way over here that are saying, "I'm shipping stuff very, very quickly," you have to remember that they also have much more time than you do.

An so, like, you can get good quality code, you can get good quality reviews. It just, I think it's really going to become just accepting that our job's a little different and then managing our time a little different. We're going to spend more time reading code, and that may or may not be okay with you, but it's going to be very okay with your boss at some point in time, which is the freaky part. Well, it was so awesome to have you here today.

I want to have you back soon because I actually have about, like, thirty five more things I wanted to talk to you about, so we'll have to put another time on the calendar sometime soon. Love nerding out,,Corywithyou.And we'dlove togetyouonboard onboarded onto soyounever haveto seethatflakytestagain. Yeah, let's do it, for sure.

Awesome, man. Thanks again. Let's chat soon. Thanks so much.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • #291 Why Most AI Projects Fail to Deliver ROI, Sinohe Terrero, CFO and COO, EnvoyGrowCFO Show · on Claude91 / 100
  • Why a $1.2B exit felt like his biggest failure, and the customer-obsession thesis behind AgencyThe GTMnow Podcast · on Claude86 / 100
  • Is Your Business Invisible to AI Search? (And How to Fix It) ft. Ray YoungRevenue Science · on Claude85 / 100
  • How Organizations Can Thrive in the Human + AI Era with David ChestnutThe Edge of Work · on Copilot85 / 100
  • Episode 018: Season 2, the $75 Consult and the Frankenstein StackAI Tools for Practicing Lawyers · on Claude84 / 100
  • Agentic Engineering for Testers: How to Automate Your Way to the Top with Amit RawatTestGuild Automation Podcast · on Claude82 / 100

More from Platform Engineering Podcast

All episodes →
  • What Do Service Meshes Actually Solve? (William Morgan, Buoyant/Linkerd)73 / 100
  • Continuous Integration at Agentic Velocity with CircleCI’s Rob Zuber86 / 100
  • Durable Execution for Real‑World Failures with Temporal’s Cornelia Davis96 / 100
  • You Need AI Sysadmins Can Trust, With Cribl's Nikhil Mungel87 / 100
  • AI-Native Ops: Making AI Safe for Production with William Collins78 / 100
Explore the best B2B Engineering & DevTools podcasts →
All Platform Engineering Podcast episodes →