The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Talking AI
Talking AI artwork

99% Correct Is Still Failure: The Last Mile for Mission-Critical AI

Talking AI · 2026-06-09 · 42 min

0:00--:--

Key moments - from our scoring

Substance score

42 / 100

Five dimensions, 20 points each

Insight Density9 / 20
Originality8 / 20
Guest Caliber11 / 20
Specificity & Evidence7 / 20
Conversational Craft7 / 20

Codemetal raised $125 million to solve a critical gap in AI-assisted coding: while tools like Cursor, Claude, and ChatGPT excel at generating code quickly, they cannot guarantee correctness for mission-critical systems where failure means recalls, accidents, or national security incidents. Ryan Atay explains that the company's focus is verification, validation, and assurance - using formal methods, fuzzing, concholic testing, and hardware-in-the-loop testing to ensure translated or modernized code behaves identically to legacy systems. The company isn't competing with code generation tools but sitting alongside them, providing the "last mile" assurance layer. Atay draws parallels to Tableau's democratization of analytics, noting that while AI democratizes coding for general use, mission-critical domains (military, aerospace, medical devices, power grids) require mathematical guarantees. He discusses how Codemetal handles real use cases like translating over a million lines of legacy C code to Rust for defense applications while maintaining identical behavior without system downtime - work previously considered impossible. For B2B operators evaluating AI infrastructure, this episode clarifies the distinction between "good enough" AI tools and production-ready assurance systems, plus Atay's perspective on how executives should leverage AI operationally (M&A screening, recruiting, content review) while maintaining verification practices.

Key takeaways

  • →Codemetal's value proposition is not code generation speed but mathematical assurance that translated code behaves identically to legacy systems in production environments.
  • →Mission-critical industries (defense, aerospace, autonomous vehicles, medical) require formal verification methods like fuzzing and concholic testing that general code gen tools cannot provide.
  • →The "last mile" gap exists because tools like ChatGPT and Cursor will claim they can guarantee production readiness when they actually cannot, creating false confidence in high-stakes applications.
  • →Codemetal uses hardware-in-the-loop testing and formal methods to achieve what was previously impossible - translating millions of lines of legacy code without system downtime or re-validation costs.
  • →Executive leverage of AI should include internal use for operations and decision-making, but every output requires verification before deployment, especially as AI scales to impact hardware and safety-critical systems.

In this episode

  1. 1The Gap Between AI Code Generation and Mission-Critical Systems
  2. 2Ryan Atay's Journey from Tableau to Codemetal
  3. 3Understanding the Last Mile: Verification and Assurance
  4. 4Codemetal's Position in the AI Coding Tools Landscape
  5. 5Testing Methodologies: Fuzzing, Formal Methods, and Hardware in the Loop
  6. 6Leveraging AI Across Business Operations
  7. 7Real-World Use Cases: Legacy Code Modernization and Defense Applications

Mentioned

CodemetalTableauSalesforceChatGPTClaudeCursorGitHub CopilotTeslaShopifyMIT Lincoln LabRyan AtayPeter Morales

Guests

Ryan Atay

Topics in this episode

Autonomous vehiclesFormal methodsCodemetalcode generation tools (Cursor, ChatGPT, Claude, Codex)fuzzing and concholic testinglegacy code modernizationC to Rust translationmission-critical systemsdefense and military softwarehardware-in-the-loop testing

Questions this episode answers

What is the difference between code generation tools like Cursor and Codemetal's approach?

Code generation tools like Cursor and ChatGPT are excellent at producing code quickly but cannot guarantee it will work correctly in production mission-critical systems. Codemetal provides verification, validation, and assurance layers using formal methods and testing to mathematically prove translated code behaves identically to legacy systems.

Why is 99% accuracy still failure for mission-critical systems?

In systems like fighter jets, power grids, or autonomous vehicles, a 1% error rate means recalls, accidents, or national security incidents. Codemetal solves for this by providing mathematical guarantees, not probabilistic correctness.

What testing methods does Codemetal use to ensure code correctness?

Codemetal uses fuzzing, concholic testing, formal methods (mathematical proofs), and hardware-in-the-loop testing to verify every edge case and ensure translated code produces identical behavior to the original legacy system.

Can AI tools guarantee they will successfully translate legacy code to modern languages like Rust?

No - all major AI tools, when asked directly, will admit they cannot guarantee production-ready translation of large codebases. Codemetal fills this gap by providing that guarantee through formal verification methods.

What is a real use case Codemetal has worked on?

Codemetal translated over one million lines of legacy C code to Rust for a government customer while maintaining identical system behavior without downtime - work previously considered impossible.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

9 / 20

The episode surfaces a few genuinely useful concepts - provably correct vs. merely 'guaranteed,' hardware-in-the-loop testing, and fuzzing/concolic testing as verification layers - but the majority of runtime is consumed by general AI-enthusiasm talk, the host's own AI anecdotes, and vague platitudes about adoption. The core thesis is stated early and never meaningfully deepened.

99% correct is still failure when it comes to mission critical systems
it's not really just a coding problem. It's more of a, like, behavioral. Behavioral problem at scale

Originality

8 / 20

The 'last mile verification' framing for mission-critical code is a genuinely differentiated angle, and the distinction between 'guarantee' and 'provably correct' is a sharp conceptual move. Everything else - use AI or fall behind, AI makes mistakes like humans do, SaaS isn't dead - is standard 2024-era podcast filler.

it's sort of like I want to rewire a city, but I don't want the power to go out
Prove is even a stronger word than guarantee. Right. Because it's provable

Guest Caliber

11 / 20

Ryan Atay has real executive credentials - former CEO of Tableau, Salesforce background, now President/COO of a $125M-funded company in a novel niche - which is legitimate. However, he repeatedly disclaims engineering expertise and cannot explain Code Metal's technical differentiation in depth, limiting the substantive value a practitioner can extract.

I'm going to get quickly out of my realm of expertise because I'm not an engineer. We should put a little caveat on this
I had all this great experience and I'm very grateful for my time at Salesforce and Tableau and the things before that

Specificity & Evidence

7 / 20

A handful of concrete numbers appear - million lines of C, a few weeks vs. 50-100 developers over two years - but they are unverified assertions with no methodology attached. Every named customer example is deliberately anonymized, and the technical claims are stated without any supporting data or third-party validation.

I can't mention the name yet, uh, soon
This would have taken you, I don't know, 50, uh, to 100 developers for two years? Well, there's a cost to that

Conversational Craft

7 / 20

The host raises a few genuinely interesting angles - accountability/insurance markets, human-in-the-loop as a Jenga tower, token-cost pressure - but never challenges the boldest technical claims, accepts deliberately vague customer examples without pushback, and repeatedly inserts his own multi-paragraph AI usage stories that stall the conversation.

I was a very early OpenAI chatgpt, uh, adopter and I went through this period of using both and now I've kind of shifted over. I still use both, but I went through this exercise with our quarterly planning. I had it hooked up to our CRM and all the other data
Wait, you lost me at fu. Was it fuzzing? Was that. I've never heard of that term

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A65%
  • Speaker B35%

Most-used words

code29critical21tools20mission17mentioned16sure16different16world15opportunity14point12environment12human11part11value10hardware10last10

Episode notes

AI can now write code faster than any human alive, and most of the time it's more than good enough. That's the magic powering the entire vibe coding wave. But there's a category of software where "most of the time" just doesn't cut it: the code running a fighter jet, a power grid, an autonomous vehicle, a piece of medical hardware. When that code is wrong, the consequences aren't a bug. They're a recall, an accident, a national security incident. In this episode of Talking AI, Matt Paige sits down with Ryan Aytay, the former CEO of Tableau and now President and COO of CodeMetal, which just raised $125 million to close that gap. Ryan explains what he calls "the last mile" for mission-critical industries: the verification, validation, and provability layer that sits between AI-generated code and the systems where failure is catastrophic. The conversation covers why 99% correct is still failure in defense and autonomous systems, how CodeMetal translated a million lines of legacy C++ to Rust in weeks (like rewiring a city without the power going out), and why the real problem isn't code generation, it's behavioral assurance at scale.

Full transcript

42 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: It's great if you're, you know, 70, 80, 90, even 99% correct. But 99% correct is still failure when it comes to mission critical systems. A lot of these code gem tools are over here, the mission critical systems are over here, and we have to help bridge that gap.

Speaker B: Welcome to the Talking AI podcast, where we talk AI, uh, with both experts in the field and early adopters. I'm your host, Matt Page, and we're here to demystify AI for you so you can get some value from it. Let's talk some AI. AI can now write code faster than any human alive. And most of the time that's more than good enough. That's the magic powering the entire vibe coding wave, the agentic coding craziness that we're experiencing right now. But there's a category of software where most of the time just doesn't cut it. The code running a fighter jet, a power grid, an autonomous vehicle, a piece of medical hardware. When the code is wrong, the consequences aren't just a bug, they're a recall, an accident, a national security incident. And that's the gap that codemetal just raised $125 million to close. And that's why Ryan Atay, the former CEO of Tableau, joined as president and COO to run point on, uh, what he calls safely delivering the last mile for mission critical industries. We're going to get into what that actually means, where code metal sits in this crowded market of cursor clock code and a host of other tools. And what an operator who ran one of the most iconic per seat SaaS businesses thinks about this crazy world where agents are proliferating everywhere. Ryan, welcome to Talking AI.

Speaker A: It's great to be here, Matt. Thank you for having me.

Speaker B: Yeah, this is going to be an exciting one because this is a, this is a true, uh, gap, I think, in the market. Like we, we use AI at our company 24 7. Our engineers are fully on board with it, but the point that you're solving here is critical. And, uh, just on a personal note, I was an OG Tableau user back in the day, like a complete fanboy of Tableau. But what I loved about Tableau is it wasn't just like the beautiful visualizations and charts, but it completely democratized analytics and suddenly anybody in business could be, you know, their own data analyst. And now we're living through something very similar with AI, where it's democratizing coding, design, video writing. I mean, you name a function or purpose in a business and it feels like it's democratizing it or anybody can build, but you've staged your neck next chapter on the idea that there's a category where democratization breaks down a little bit. So walk me through what you saw and what attracted you to this mission that Code Metal is going on.

Speaker A: Yeah. Um, and by the way, thank you, uh, in my old capacity for being a tableau community member, because I'm still very connected to that, that experience. So look, I think, you know, coming into an environment which we're in now, which is this kind of AI world that we're in, I really just saw this large opportunity to work with really a lot of smart people, a lot of, you know, MIT Lincoln Lab engineers, um, like our co founder.

Speaker B: Ah.

Speaker A: Or sorry, founder and CEO Peter Morales. And really it was, you know, I had all this great experience and I'm very grateful for my time at Salesforce and Tableau and the things before that. But it was like, how do I take all these things that I've learned over, you know, 19, 20 years and apply them to, you know, a new industry, but also an industry in a company specifically like Codemetal, where I could make a bigger impact? What do I mean by that? Well, it was, you know, make a difference, make an impact. Because ultimately there's so many. There's a lot of AI noise. Of course, it's like, you know, your podcast is called Talking AI, so it's like, there's a lot.

Speaker B: We're part of the hype, man.

Speaker A: We're part of the hype. But the hype's exciting, right? We're exciting to be in this environment. But there's this concept of, like, how do we make AI more trustworthy? And, you know, because if you don't know enough, it could be dangerous. You know, of course there. And we'll talk probably about this today. But, you know, my opportunity I looked at, it was like, how do I make an impact? There's this opportunity to make AI, uh, more trustworthy. And ultimately, I think there are a lot of people that don't understand it. So, you know, it's critical for our nation, for our military, you know, for our government, for mission critical companies, maybe that ship automobiles or airplanes or go down the list of things that we depend on every day as, you know, citizens. And if it's not done correctly or used correctly like, that could be problematic. And there are many ways and I can explain them, but like, that. That could be problematic. And I think there's a lot of noise around, like, well, is it going to Take my job. What about is it safe for the things that I depend on every day, like when I go get in my car or when I get on an airplane? Like, this is sort of the next chapter of is it safe? Is it trustworthy? And our intent at Code Metal is to really deliver that last mile.

Speaker B: Quick break in the pod. Our State of AI 2026 report just dropped. And it breaks down what actually is changing in AI, what's hype, and what leaders need to be paying attention to this year. You can grab it right now in our show notes or at hatchworks. Com.

Speaker A: Yeah, because that's how I looked at the opportunity.

Speaker B: Yeah, but you're injecting this, uh, uncertainty into the world with this new amazing thing we have. It's AI and generative AI, but there's a lot of legacy systems that exist, these deterministic systems. So it's like, how do you graph those two together? In a sense, I feel like that's part of it. And you talk about like the last mile for mission critical Industries. I'm curious, like, and this probably, um, you know, a bleeding edge of where is the last mile? Where does it start and where does it end? Or is it almost like a holistic view of how you build and leverage AI and the coding and building process?

Speaker A: I mean, I think the last mile is relevant if we talk about things like code gen tools. And there are a lot of them today, the cloud codes, the codexes, the cursors. And they're all great, right? We know this. Um, we use them ourselves. Um, but when you get into, and you can even ask the various AI tools today, like, you know, can you, can you generate code? Let's say, can you translate C to rust and guarantee it'll be production ready and safe? And they will. All, all the tools will tell you. Well, almost, but not quite.

Speaker B: Ah.

Speaker A: Um, they can't guarantee it. And that's the, that's the tool itself. Right. And we know that the hardest part is verification. And, you know, will it be correct? Um, have we tested every use case or edge case? Do we know every requirement? Is it secure? Um, this is kind of like what I call the last mile, because there's just a lot of stuff that people forget at the end of the day that needs to happen. It's great if you're, you know, 70, 80, 90, even 99% correct. But 99% correct is still failure when it comes to mission critical systems. And so how do we, you know, our goal is to, you know, I think about like a Lot of these code gem tools are over here, the mission critical systems are over here, and we have to help bridge that gap. It's the, uh, assurance, if you will, like, how do I, how can I be assured that when I translate or modernize code that it will actually work in a mission critical environment? And I think that is, that is really what we're trying to solve and we are solving. And you know, it's, it's a lot, but it's also very, you know, it's a lot to unpack and understand, but it is a big opportunity.

Speaker B: Well, it's funny you mentioned, like, if you ask AI today if it could do, I forget the two you mentioned Rust and something else. I feel like most of the time, instead it'll be like, heck, yeah, I can do that. No big deal. In the end, you're like, you screwed XYZ up. Because you do have that sycophantic nature with AI where it will reassure you and say, yes, I got this. Even if it doesn't have full context. Which is scary, I think, when you get to some of these, uh, you know, mission critical things like you're talking about, but you mentioned it though, like your teams, they use cloud code and some of these other tools. So I think it'd be helpful. Like what. You're not necessarily in competition with these tools. It's, it's almost like, um, tertiary or uh, in support of these, these other tools. So where, if you look at the map of everything going on in this crazy world of, uh, AI coding tools and AI tools in general, like, where. Where does codemetal sit in that, uh, that. That hemisphere?

Speaker A: Yeah. Um, yeah, I think at number one, I think it's early in our. Even though it feels like we're a few years into this, we are, but it is very early in the sense of like, understanding what's possible and also, of course, identifying some of the risks and things that can be solved. We don't really see a lot of competitors, you know, at this point in terms of, you know, focused on verification, validation, assurance. The things that I've mentioned, like, making it trustworthy. Um, you know, when I think about it, let me try to like the, the thing I said before was C to Rust, right? So like, you could ask a question like, well, you know, and literally if you go in and you say, like, go into ChatGPT and say, like, can you do this in a production environment? Will you guarantee it'll work? It will say, well, yes, but like, there's how he's at the button there. And so I think the scenario would be like, yes, you could use any code generation tool. Um, you know, you could. They could, like, run an entire repository. It could be agentic. It could generate, you know, code pr, it could refactor some things. It's pretty good. Um, but the problem becomes, like, almost the behavior. So, um, it's not really just a coding problem. It's more of a, like, behavioral. Behavioral problem at scale. So if the tool succeeds, like, you know, as I said before, 70% of the time on, like, a complex refactoring, you could still fail at some level. Right. And be, you know, a scenario where I think you run into challenges. Right. So they can write code, but they cannot guarantee that, let's say a million lines of C, or in this case, you know, let's call it C legacy C, that will. It will behave exactly like it did before once it's transformed. And that is the thing that we're solving. And there's a whole bunch of work that goes into making sure that it's guaranteed, but that is really what we're delivering. We're delivering a guarantee or an assurance that it will work. And that's the difference.

Speaker B: That's interesting. I wonder, too, because you all always hear a lot of people's answer to this kind of problem. You're talking about is human in the loop. Oh, we have a human in the loop here to verify all of this. Which, you know, frankly, I think people get comfortable in that human in the loop. You know, verification decreases as time goes on. But I think that the bigger point is, though, I mean, this is a process in a system like anything else, and there's natural bottlenecks, and I think that the bottleneck is shifting in many ways to that human in the loop. So it's like, as this just completely progresses and scales like, how. How do you think about human in the loop? Because it's having that as the verification step. It almost feels like you have, you know, a game of Jenga and you're at the very end of it where it's about to topple over. In a sense.

Speaker A: Yeah. I mean, I think. And perhaps human in the loop, you know, there's. I like to think of this as, like, you know, well, let's. Let's unpack a couple things. So the first thing is like, you know, you've got AI Makes everyone's jobs easier. Yes. If used the correct way. I agree. I use it all the time. I'm happy to talk about that as well. And there's a human. And you know, let's say I've got an engineer and now they're not just writing code, let's say for said application, they can be doing other things, they can be thinking a bigger picture, they can be reviewing code, you know, is a. But what about the hardware? Right. Because we're not just rewriting software, we're trying to connect it to hardware and do it in a production way. So we talk about our kind of our North Star here at Codemetal, which is like how do we basically make things. We want to make all things programmable software, hardware, et cetera. Think of it like um, if you've ever been in a Tesla or full self driving experience where you have the ability to ship over the air updates, software impacting hardware, that's great. They can do it, but not everyone can do it. And so I think making sure that it works. Now if I say well okay, let's just translate some old code and we'll just have a human review it. Well, what about this process of like we need to analyze the code, we need to understand, you know, and this is of course some component. What we do, we want to do test generation. So we do something called fuzzing concholic testing. You go down the list of various things that we include in our suite to make sure that it checks every use case, it checks the source code. Is it behaving the same as it is in this new system? Is it checking every edge case? Is it checking for variables that aren't expected? So there's a, you know, and I'm going to get quickly out of my realm of expertise because I'm not an engineer. We should put a little caveat on this. But the point is we're going a bit deeper to make sure that you can count on it ultimately at the end of the day. And that includes in addition to human in the loop. It's hardware in the loop. So we want to be testing hardware continuously. It's a big part of like when I walked into one of our offices at first I'm like, wow, there's all this like hardware. It's a different world, right. And so how do we make things run efficiently, you know, from a um, M small device to a mission critical device and do it in a safe way.

Speaker B: Wait, you lost me at fu. Was it fuzzing? Was that. I've never heard of that term.

Speaker A: Yeah, I will. You know, think of it like um, what are the different edge cases that could be a problem. Right. And so it Tests every different way. And we use this thing called formal methods, which is another word for math. Like there's a lot of. This is my learning curve which is like, okay, what does this mean? How do I uh, you know, trying to bring it all to English. Uh, so that I like to say how can I make sure my parents can understand what I'm working on? That's my uh, my intent and my goal as we talk about these things.

Speaker B: Yeah, that's a good North Star. But like this is separate. But how amazing is it like that you have AI as this partner to like help guide you as you learn? I think just, just such a cool premise in general. But you mentioned uh, you mentioned like you're, you're leveraging AI. Like I'm curious. Yeah. Companies right now, especially newer companies are growing in a totally different way because they have AI at the beginning versus existing companies going through their own innovators dilemma of how do we apply this to our existing business. I'm curious, like how are you a, as an executive leveraging AI? I think that's super interesting for our audience but also from a business context. How is AI being leveraged? Whether it's an operations, uh, it could be anything. Sales, marketing, operations. Name your function of the business, anything that kind of sticks out in your mind either daily use or within the business.

Speaker A: Yeah, um, you know I, it's funny because we, we talk about this, we're a small company, you know, we're still under a hundred employees. But um, we, you know, we look at the opportunity, you know, so we don't have like I was in theory one of the first business people, non engineers in the company. We have a lot more now, but we can do a lot more than we could have, you know, if we had started this company in like you know, 2000 or 2010. Like even it's a, it's a different world. So I think, you know, we use, you know, whether it be you know, even things like an M and a pipeline or you know, recruiting and screening candidates more quickly, like it accelerates a lot of what we can do. Um, or hey, give me an opinion. I wrote something, you know, maybe I want to post it on LinkedIn or wherever I want to post it. I don't have a marketer yet. So hey, I need some. Can you help me just like think through or give me perspective? But again I'm not going to just take exactly what it says. I want to make sure that it's like verified. Again, it's the same premise. It's like, can I trust what is in here and how do I actually go through it? It's like, you know, in, in the past you'd ask many different people before you maybe would make a decision. It's a similar thing. It accelerates that process. Um, I do think that, you know, and of course there's, I think we found with the code generation tools is like, you know, if I'm building a simple application or a SaaS application or whatever it may be, maybe I can get started in advance really quickly if I'm where we shift from that into a world where I just say, like, let's be careful. Let's make it sure it's trustworthy as we're going to convert legacy code to modern code. Let's say it's a million lines, like I said before, it has to behave the same if it's going to, let's say, you know, impacts a, uh, you know, know an autonomous vehicle. Right.

Speaker B: Yeah.

Speaker A: Whether, especially in a military environment, like, good on the. You can just see how this can spiral out of control very quickly. Um, and so I, I think that every business right now needs to be using AI internally to run its operations. You know, like, hey, I have built a cash flow statement. Hey, does this look right? Or you maybe it can build it for you. Like, those are all great things, but you can't just like, send it out. Like, I think that's the, uh, the message I would want to tell people. Like, absolutely. If you're not using AI, the companies that aren't will lose to the companies that are. But, uh, be careful is sort of the intent, especially at this stage.

Speaker B: Yeah, yeah, there's definitely nuance about it. And I feel like the only way to get good at understanding the nuance is just by using AI a lot. Like, I always go back to the Shopify CEO talking about he has given a mandate for his team to use AI reflexively. Just like, it's just second nature. And then you do start to get, uh, it's the nuance, it's the taste element of like, okay, I can use AI here or I know to push here. Like, I, I don't know about you, but I've been completely claud. Code pilled over the past several, several months. And you know, I was a very early OpenAI chatgpt, uh, adopter and I went through this period of using both and now I've kind of shifted over. I still use both, but I went through this exercise with our quarterly planning. I had it hooked up to our CRM and all the other data and had memory for everything we do. And I just had IT build the full report for me. But to your point earlier, it wasn't just like, okay, there's the report. Go. Like, I went through it and I found some interesting insights that I hadn't even noticed or thought of, which was just this really cool experience in, like, another one we have it hooked up with, again, our CRM and Slack and whatnot. And it's going through there and identifying. Okay, what. What prospect should we be following back up with, uh, that we haven't talked with in a while, or something's happened. And it's just giving us this regular report with a, uh, you know, a score attached to it. And we now just have it, like, running and doing its thing. So it's just once you see, uh, the possibilities, it's like, oh, my God, there's so much I can do with this.

Speaker A: Yeah. I think that there we're in the inning of perhaps, you know, or the stage of seeing all the possibilities. Right. And then I think, you know, like, we've identified code is one area where it is real and relevant and very, very useful. You know, I still think, like, I'm still trying to work, whether it's, you know, Claude or pick your. I'm like, hey, schedule a meeting. It's not quite my assistant yet. It's not quite perfect. It's getting better every day. You know, schedule a meeting, do this. It's like, well, I can't access your calendar. Or, well, I can't do this. Or, well, scheduled it on the wrong day. So we're getting there. But I think it's education, just like what we're talking about here. Like, why am I doing this? It's because we. There's an opportunity to educate and make sure that, you know, not only impact the business, you know, or like, you know, we work a lot with various, uh, forms of our government. You know, it's pretty mission critical that things are done. Right. Right. It's pretty mission critical that, you know, we're making sure that if there's a drone in the air, there's a tank or whatever it is that's being operated in a safe way that is dependable. Um, and that is a big opportunity. But not everyone understands this, so you can't just assume, like, you know, if it sounds too good to be true, I always feel it probably is. Right. That's a saying from the past. So just check your work.

Speaker B: Yep. It's maybe a theme, you know, 100% and it definitely applies. I'd be curious though, like, what you can tell us, like, what are some cool use cases where you're starting to apply your approach, your methodology, your tooling with specific customers? Obviously, I'm sure some of the government ones you can't go as deep on. But are there any interesting ones that are worth noting?

Speaker A: Yeah, I mean, maybe the, um, the one I mentioned before is a customer. It's one. I can't mention the name yet, uh, soon. But the one where they basically had come to us and said, you know, we have over a million lines of legacy C code that they needed to basically, they needed to basically translate, um, and modernize, but using the same system, the same hardware that they've used. So it's kind of like. It's sort of like I want to rewire a city, but I don't want the power to go out.

Speaker B: Yeah.

Speaker A: So, okay, well, how do I do that without. I mean, it's impossible. So this is like an impossible task that all of a sudden we're like, well, we can do this in like a couple weeks. It's not that hard. Um, and we can guarantee it'll work. Right. So a million lines of C, you know, translated into Rust with zero issues in a very short time frame is basically impossible. Um, and now it is not impossible. I think maybe another example would be. This is the one I have to be even more vague with. I apologize. We can do this later when we can talk more about it. But, you know, think of like a, ah, Department of Defense or Department of War. They have bought software and, you know, various types of software that's been like, call it stovepiped or disconnected. Some of it they can use, some of it they can't. Millions and millions of dollars has been spent on this. But at the end of the day, right, they want to make sure, like, how do they modernize it? They can rewrite it, perhaps, but without spending a bunch more money to revalidate that. It's actually working. So there's, you know, let's go through a process of like, where I need to translate existing software into a modern, maybe more maintainable scenario. And I want to use that so I can do some simulation for war fighters. Right. Um, Pick your things. Right. The code needs to be translated, it needs to be provably correct, and then it needs to. We need to know that whatever goes in comes out and there's no need to revalidate it. So it's a. You know, again, I'm sorry, I'm Being vague. But like I, There's a lot of information that's sort of sitting in different places. We're bringing that all together and providing a simulation environment so they can run real world scenarios. And that, that is very powerful and really only something that we can do at this point.

Speaker B: No, that's, that's really interesting because like you've, you've been in the business world for some time now. Like how many.

Speaker A: Very long, very old.

Speaker B: Yeah, yeah. How many transformation initiatives get proposed, done and they just die on the vine because they become, they're too complex, they take too long. By the time you get ready to do something, things have changed. Like that's, that's a re. Very, very real problem. And you mentioned like Tesla uh, earlier, there is this element of like modernization that potentially is possible now because like we talked about earlier, the innovator's dilemma. A core part of that is just like there's so much built up infrastructure and nuance to things and AI can actually help accelerate. I mean that's kind of what y' all are doing in a sense I feel like in, in a lot of ways.

Speaker A: Yeah. And I was trying to think of another one. Like you know, I think robotics, you know, again in some of the use cases I give are more like perhaps military focused or defense focused. But that's just a big, big area of demand for us at the moment. But you know, if you think about maybe mid, non traditional type of environment where there's uh, you know, let's say uh, some kind of a tank or something that is autonomous, whatever it may be, right. Or drone or something like that, how do we, you know, how do they get more time to think about, you know, what is the environment like? You see the environment just like in a, in a self driving car, you see the environment. But how do we get more time to react? Um, and how do I do that in a way that is safe and accurate, like very important to be accurate in that environment. So I want to compute, let's say the scene or sort of the visual faster so that the object can have a safer ride at higher speed. Gives the machine more time to react. That can only be done, um, you know, if they trust the output, right? And so if they don't trust the output, obviously they're not going to do it. And so that it's. These are the types of things that are very like, you know, very critical. That's why we say mission critical. Um, we're not focused on the non mission critical things. These are like things that are very, very important. They, you know, failure can be catastrophic and potentially any of these examples.

Speaker B: Yeah, and you mentioned a few times the word like guarantee, where you can guarantee something. Like, how do you, that word carries like a lot of weight. Like, how do you think about that in the sense of your, your business and whatnot? Especially, I mean, when the stakes are this high. You, you almost need that, uh, in

Speaker A: a sense, I mean, we think about it as, as assurance. Uh, you know, we try to talk to our customers about outcomes. You know, how do you. In a world where, you know, and, and this, I think is something that we could talk about for, for a long time, which is how do you price these things? Like, how do you go through the process of, you know, making sure the customer sees value? Well, when I can take a million lines of, uh, you know, do the impossible, right, Which I said before is take a million lines of C and convert it to rust without changing behavior, and it runs in the same system, it's highly valuable. Uh, and we can guarantee your, prove that it will work. Right. Prove is even a stronger word than guarantee. Right. Because it's provable.

Speaker B: So it's that verifiable proof that you're doing in these test environments and things like that. It's not just like, uh, satisfaction guaranteed. I put the stamp on it, it's good to go. Like there's something tangible behind it.

Speaker A: Correct. And we can show those results. Right. So we can say, hey, we will show you that it works. And that is where this becomes quite valuable. And if I can say, hey, this would have taken you, I don't know, 50, uh, to 100 developers for two years? Well, there's a cost to that. And then what are they not working on that they could be working on? We can just do this for you in a few weeks. It's a different world. And that, that's where I think that's why, like co generation and that's why I really like all these, you know, the OpenAI's, the anthropics, et cetera, the cursors, they're all part of the same thing. They're just doing it at a different, like a different stage, if you will. Uh, we're very focused on provability assurance, verification, validation and delivering that last mile, which I think is pretty important. It's just maybe not on the top of everyone's mind because we're still in a mode of like, oh, I can, I can write a blog post faster and I can write, you know, I can build a Time and expense app, like really quickly. Those are all great examples too. Um, we're sort of over on this other spectrum which is like we want to make sure that we're safe and we want to make sure that there are no catastrophic issues that happen. That's why it's attractive to me.

Speaker B: I'd be curious your take on this. It's like the accountability aspect uh, of it because AI is not like a human to where it can be responsible and accountable but as things as it becomes more mission critical and foundational, um, where do you see that accountability element going? I could see like entire insurance markets starting to emerge around this and there's obviously business insurance that already exists. But I feel like there's this new element that's, that's going to be introduced where okay, well the AI is at fault in some way. What happens then? I'm just curious, like how do you think about accountability? Uh, and this may not necessarily specific to Codemetal, but I think it's going to be a bigger and bigger topic as we keep progressing.

Speaker A: Yes, um, it's a really interesting question. I think that um, I feel like they're, you know, there's two ways to just to, to look at this which is like there's probably a group and again there's a lot of media attention on everything we talk about with AI. There's probably a group that will say it's more dangerous. Right. Because it could make a mistake. Well, humans can make mistakes too by the way. Right. We know this. Humans can hallucinate. Maybe we wouldn't call it that, but we know that you may get the wrong answer. But if we can prove that something works. Right. Because that's not really talked about. It's just like well, AI is amazing and it is. But when we can prove it works with um, in our case we do it for mission critical environments. But in let's say any environment, if we can prove that something can work using maybe it's what we do or some other technology, you should be able to assign accountability to it. It's just going to take time. It's going to take education back to. My earlier point is like we need to educate the risks, how it works, how you can prove it works and where you shouldn't use it. Like you know, we, we know we shouldn't use it in certain scenarios. Like for example, it's not great necessarily at uh. Well, I don't want to go down that road because uh, we'll get in other discussions. But there are things it's better at than things that it, uh, is maybe not as good at is maybe the way to say that coding is one of them.

Speaker B: It's great, totally. And I think one of the best analogies I've heard, because it's probabilistic at its core. And this, this uh, variability, uh, very, um, being able to verify it is a whole nother element I haven't considered. So my mind's kind of like opening to this new concept.

Speaker A: But think of this, think of this like we call it verification and validation. V and V. You can't really, it's easy to remember V and V. Like is it verified? Is it validated? Like, this is how I look at a lot of these things.

Speaker B: I like that, I like that. But the analogy that I've heard is, okay, if I have a calculator and I put in two plus two and it gives me, I throw it away because it's broken. It obviously is not working. But with AI, like if you get an incorrect answer, it doesn't mean it's broken. It may not have sufficient context or there's some other variable there. And the point you mentioned of Guess who else is not always right? Humans. Because we're kind of probabilistic beings in a sense too. We're working off the context, our nature, all these different factors that play into it. And I think it's just important that people view AI in that way. And ah, it gets back to the point you mentioned, like it's going to make mistakes. Like you have to go into it knowing that in a sense.

Speaker A: Yeah. I mean I feel like we can all give a hundred examples of where it wasn't quite perfect, but it was still pretty darn good.

Speaker B: Yeah, exactly, exactly. And so one other point I want to get your take on. You mentioned kind of token usage earlier, but the Uber AIM cto, I saw him the other day, he basically said they blew through their entire AI budget. Uh, and for context, for those listening right now, we're in April. So if you're listening around Christmas time and thinking, oh, that's not too bad, we're in April. And he said they've blown through their entire AI budget and they have, you know, 5,000 engineers, 92% are now using AI and agents in the coding process. I just feel like this concept of tokens is going to be such a critical part of the business conversation. The, the, you know, cost of goods or however you want to think about it. Um, but what, what's your thought thoughts on token usage, how it applies where it's going to go. Any, any thoughts on that either from your business and use of AI or whatnot?

Speaker A: I mean, um, I like to, I come. I've seen like many different spectrums. Right. So whether it's the per seat or usage model or tokenization and tokens, et cetera, I think is great. I think there's an opportunity and I think there's an evolution around um, whether it's called outcomes based or uh, value delivered. And I say that because like, let's just say that I, you know, I'm CEO, I'm thinking about our usage of AI, uh, internally. Do we, you know, should we have a budget? I think it's more like, well, if you can demonstrate a business case or a use case, that shows me that you can, you know, do more with less and be more, you, uh, know, have a better output. That's a hundred percent a win. Right. Because it's, it's not just about, it's like budget is one angle, but it's also hard if you're running a large company having come from one to know like, what should I allocate and where do I want to spend?

Speaker B: Yeah.

Speaker A: I think in a world where you have to look at these AI initiatives as almost like. And I, I'm um, I don't. Never thought I would use the word again, but like a business transformational scenario which is like, okay, what the C thing I mentioned before, if I'm going to like save, let's say you know, five to $10 million, but also continue to be able to sell my hardware and because I basically been able to modernize it, that has a lot of sort of variables included in it that would impact my business. And I go back to like, I like to say what we try to do is we help companies be compliant or safe is maybe another word. We help them go faster and then we help them ultimately be more open. Right. Because. And those are things that you can't just say like, oh well, what's my budget to be safe, compliant, fast and open. Yeah, I mean hopefully that's everything you're thinking about. Um, and I think that that's really like what we're trying to unlock is AI can do a lot of these things, but if done incorrectly it can destroy a company too. We haven't seen enough of those examples yet, but I'm pretty sure they're coming. She's like, oh, I bet my, my bet the house on this particular thing using AI and oh, it didn't work. Or oh, there was Some catastrophic thing that happened. We don't want that. We want to do those three things I mentioned. And we want to be able to hopefully deliver in what we're seeing. I mean we have customers paying us based on the outcomes that we provide. Um, but it's a, it's a different world, right? It's like a different way to think about what is the value, how do you calculate or you know, sort of quantify what they would have spent otherwise and what is the value of being open and extensible. Right. Think about chip manufacturers, right. Think about various languages to one particular vendor who I won't name that's dominant. Uh, and others like how do we open the supply chain? Right. That's a big opportunity as well, which we're working on.

Speaker B: Totally. We got another V for the uh, the two V's, Value. So we got But Verify, uh, validate and value. So we got the trifecta now. Exactly, exactly. Last thing, uh, I got for you. So obviously you ran Salesforce massive company. Like you're very adept in the whole SaaS ecosystem. But I'm curious, like, how do you see that world evolving as it's not just humans using the tools. Like agents are becoming a big part of the users of the tools. And I think a big part of it too is like, okay, well how do we tailor our products and services to this new class of user, which is AI in a sense. Mhm.

Speaker A: Yeah. Um, I mean like I mentioned with the other thing we talked about is there's definitely a lot of media attention and I think gen, you know, story generation, it's like SaaS is dead or SaaS is this or that. I still use these tools. Uh, and you know, it's not just because I worked at Salesforce, but like we just got off a forecast call. We're using Salesforce. You know, there's use Tableau. Like they use all these tools on a regular basis or maybe you know, workday or other tools. Like they're not going away, they're in fact getting better in my mind. Um, now does the market value them differently? I think they're trying to because they don't exactly know how to value them in a new world. Right. There are new, new, new things coming at them. You know, what's the long term cash flow scenario like? Of course, like, you know, there's, there's questions. But I look at these things as they're critical systems that are also evolving. I mean having come from Tableau and seeing things like Tableau next or Slackbot or other things which are great technologies, they're evolving too. Right. So I think the companies, and there are a lot of them that are all evolving, those are going to be great. The ones that aren't evolving, those ones I might be a little worried about. Um, you know, as it relates to, you know, I think many of them are trying to switch their pricing models, et cetera. How the public markets value them, I guess, is it's hard to know what the public markets are going to do, in my opinion.

Speaker B: Yeah, well, I think your point's spot on because at the end of the day, stock price and valuations based on the prediction of or certainty of future cash flows, I think the market's just uncertain right now. It's not that it's not going to be there. There's just a lot of uncertainty. And you see that playing out in the stock price, for example. But like, the example, not to go deep on Salesforce, but you mentioned Slack, like, it's getting almost, uh, a new rebirth because, like, this is this new conduit to talk to AI and agents. So there's this completely new use case that never existed before that's like, oh, and I think every business. I don't know if you have any thoughts on this, but I feel like every business needs to go through that exercise of, okay, what. What is enabled for us or how many users or AI, uh, use our products in a different way now that we have this new thing. I don't know if you have any thoughts or frameworks or approaches or is it just, you know, get everybody together and see. See how they're using tools in unique ways.

Speaker A: You mean internally? Is that. I just want to make sure I understand the question.

Speaker B: It could be internally or I think just your products and services in general. Like, how could they be used in unique ways for, you know, existing companies? We have a lot of, uh, existing enterprise listeners and whatnot, and they're thinking like, okay, how do I evolve my product and where does my strategy go? And I think part of it's how do you tailor it to this new, new way of working? In a sense.

Speaker A: Yeah. I mean, I really. I think it just depends on what your goals are at the end of the day. I mean, you know, we. We use, uh, like, we use Salesforce, not just for sales, but we also, of course, think about it like, where are we tracking our M and A opportunities or wherever it may be. So, um, what I, what you said about, like, the Slack scenario, I think is interesting because it now becomes like that interface for all of these, you know, whether it's CRM, um, related things or things that are outside, you know, a lot of different use cases.

Speaker B: So last thing for you, if a business leader, you know, taking one thing away from this conversation, uh, between like, you know, AI hype and what's real. What. What's that one piece of advice you're given to leaders listening right now that are worried, excited, confused, whatever, whatever analogy you want to give it.

Speaker A: Yeah. Um, you know, I think, um, I'm going to go back to kind of a little bit of a thing that I think about a lot, which is, and I try to talk about with our team, which is there's a real opportunity to leverage AI in a way that can accelerate not just your company, but your personal goals at the end of the day. But you just have to know where it's best to use it and where it isn't. I think the biggest risk for people and, or companies and, or nations ultimately is to not do anything, um, or be afraid of it. A lot of this. The best way to learn these tools is to use them. Try them. All right? Compare the results. I mean, I was doing this the other night where I was like, okay, I, you know, I haven't hired all of the team members that I want yet. I'm trying to, but I need help with, you know, scheduling certain things and setting certain things up. And I'm, you know, got my cloud code and that, uh, or, uh, Claude for work. I've got these other things running and I'm testing them all. OpenAI, you know, good. On the list of various tools, some do things better than others. Right. Well, and then I'm wondering, well, why doesn't this work? It should work, right? Or things that aren't quite perfect yet, but it evolves so quickly, which is great because if you check back in a week, it'll probably be doing a better job than it did last week. And so be aware, educate yourself, talk about these things. You know, listen to more stuff like your podcast here and others I think are all important things like the learning curve is steep, um, and if you take a break from it or, like, I'm good, I don't need to use this stuff right now. I think those are the, those are the people and, or the companies that I think are going to be perhaps behind, uh, and maybe not in position to, uh, take the lead.

Speaker B: Great advice. So listen to the podcast. If something doesn't work, come back in a week and try it again, which I think is very true. That's what I do in today's age. Ryan, thanks for being on Talking AI. So where can people find you? Learn more about Codemetal or I guess maybe it, uh, sounds like you have some openings as well. You're looking for some, some people.

Speaker A: We're hiring a lot of people. Uh, yes, we are rapidly growing and um, you know, if you're out, if you're really thinking about the opportunity to make AI more trustworthy, I think we're a great place to be. Um, I guess codemetal.com is kind of where we are. We don't yet have a conference or anything, but we'll, uh, we'll be communicating more soon once we get our marketer on board.

Speaker B: Awesome. Brian, thank you for Talking AI.

Speaker A: All right, thank you. Have a great one.

Speaker B: Matt, thanks for listening to the Talking AI Podcast. If you enjoy the show, give us a follow or subscribe on your favorite PODC podcast. And don't forget to leave us a review. We love those. For more info on talking AI, visit talking aipodcast.com Quick break in the pod if you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high impact tailored AI use cases based on your business, your goals, your pain points and and your industry. No fluff, no generic use cases, just real ideas that fit your business and the rank by ROI potential. It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com AI opportunity finder.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Formal methods with Hillel WayneThe Pragmatic Engineer · on Formal methods95 / 100
  • The Evolution of Crash Test Dummies: Ensuring Road Safety with Chris O’ConnorAVL's Reimagine Mobility Podcast · on Autonomous vehicles86 / 100
  • The Robot Is Waiting on Your Data.AI Proving Ground Podcast · on Autonomous vehicles78 / 100
  • 3D HMI in automotive: opportunities, challenges and the future of the industryShine: a podcast by Star · on Autonomous vehicles75 / 100
  • Pricing, Data, and the Future of Insurance: A Conversation with Michael NadelBanking on Information · on Autonomous vehicles74 / 100
  • Data Storage Steers AI StrategiesTech Barometer · on Autonomous vehicles73 / 100

More from Talking AI

All episodes →
  • More Agents Than Employees: How Zapier Disrupted Itself Before AI Could90 / 100
  • The State of AI 2026 Mid-Year Reality Check
  • Context, Control, Collaboration: Why Capability Was Never the Bottleneck
  • Past the Productivity Ceiling: Rebuilding the Enterprise from First Principles
  • The VC's Lens: How AI Is Rewriting the Rules of Defensibility
Explore the best B2B AI & Data podcasts →
All Talking AI episodes →