
Software Unscripted · 2026-06-19 · 1h 18m
Key moments - from our scoring
Substance score
68 / 100
Five dimensions, 20 points each
Tiger Beetle, an open-source distributed financial database, achieved what few databases have: a glowing Jepsen report from Kyle Kingsbury, the industry's most feared database auditor. Jeroen Greef, the founder, explains how they accomplished this through meticulous engineering in Zig rather than more fashionable languages. Kyle's audit was particularly comprehensive - he built an independent reference implementation of Tiger Beetle's entire API, added novel storage fault injectors to corrupt data files during testing, and even leveraged Antithesis's deterministic simulation testing. Despite all this, Kyle found only one minor bug on the read path that didn't affect core durability or linearizability guarantees. Greef emphasizes that Tiger Beetle's success stems not from language choice alone but from understanding the difference between memory safety, correctness, and safety. He argues that memory safety (which Rust provides) addresses only 1 of roughly 1,000 correctness invariants in a distributed system, and that true safety means graceful failure when bugs occur. The company operates a managed service for smaller organizations while keeping the database open-source, addressing the operational burden startups face when running complex infrastructure.
Kyle found only one bug on the read path related to pagination that didn't always return all results, and didn't break Tiger Beetle's linearizability, consensus, or durability guarantees. He also tested data corruption scenarios but couldn't break the core system.
Zig provides checked arithmetic by default in safe builds and enforces static allocation with no global allocator, preventing out-of-memory crashes in mission-critical systems - capabilities that Rust and JavaScript ecosystems actively work against by design.
Protocol-aware recovery allows Tiger Beetle to recover from storage faults and data corruption through its replication protocol, enabling the database to survive scenarios where Kyle deliberately corrupted or moved data files during testing.
Tiger Beetle operates a managed service for smaller startups that lack the resources to run complex distributed infrastructure themselves, while keeping the core database open-source.
Memory safety prevents use-after-free bugs (one of ~1,000 correctness invariants); correctness means the system behaves as specified; safety means detecting violations and shutting down gracefully rather than corrupting state, which is critical for financial systems.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains several genuinely non-obvious claims per the memory-safety/correctness/safety tripartite distinction, checked arithmetic being off by default in Rust safe builds, and the economics of linear developer salary vs. exponential software asset value. However, roughly a third of the runtime is occupied by mutual Zig appreciation, wandering LLM commentary, and soft landing-pad questions that dilute overall density.
memory safety is one of a thousand in terms of correctness. And yet our view of correctness I think is too small. It needs to be bigger as an industry.
by default in a safe build, it's disabled. And what people don't know is in Zig, uh, and positively speaking, I think I'm sure that Rust will change this default, and they really should. In safe builds, it must be on.
The physical-vs-logical memory safety distinction is articulated more sharply here than in most industry discourse, and the supply-chain-attack argument inverting the 'memory-safe language = secure' assumption is genuinely contrarian with a practitioner basis. The Nabokov/Hemingway language analogy for correctness being a design property rather than a language property is an original framing; the LLM section recycles broadly circulating takes and drags the originality score down.
it's kind of like saying, I'm going to choose the language I grew up speaking as a baby from my parents... if I speak Russian, I'll become Nabokov, and if I speak English, I can become Hemingway. So I'm expecting English to make me a great writer.
most Rust developers I chat to uh, don't realize it doesn't give you logical memory safety, it gives you only physical memory safety
Joran Greef is an authentic practitioner who built a production distributed financial database from scratch in three and a half years, operates it for major financial institutions, passed a rigorous Jepsen audit, and ran his own in-house DST fleet - not a conference-circuit thought leader. His deep specificity about systems-level tradeoffs reflects genuine hands-on experience rather than borrowed frameworks.
we built everything in Tiger Beetle to very high standard. We didn't take anything off the shelf. So we built every line of code is handcrafted, very high standards, tight, strict tolerances. And we got to production in three and a half years.
he actually could also use antithesis. So he could use their DST against us. Um, and he didn't find anything like that.
The episode delivers several concrete, verifiable specifics - the 1,024-core Finland DST fleet, 700x speedup yielding ~2,000 simulated years per day, 60ms SLA for multi-regional deployment, $256K TigerBeetle donation matched to $500K total - which anchor the technical claims solidly. The main drag is that customers are described only by category ('second biggest brokerage in a country') and there are 10,000 assertions claimed without any breakdown or external citation.
there's a fleet of 1024 cores dedicated Ah, they're in Finland... There's a speed up factor of around 700. So one second on one core is worth 700 seconds and we've got a thousand cores so it's like 2000 years a day of simulated runtime
some of the SLAs are like 60 milliseconds even for multi regional deployment which is like incredibly tight
Richard Feldman is a technically literate host who adds genuine substance - notably the Rust global-allocator ecosystem point and the out-of-memory attack vector hypothesis - but the conversation tips heavily toward mutual agreement, especially in the LLM and Zig sections, with almost no productive pushback or probing of unchallenged assertions. Follow-up questions are present but tend to extend rather than interrogate the guest's claims.
I would bet that if you had used Rust or you had used JavaScript that Kyle would have been able to find ways to break Tiger Beetle or cause it to go down by exploiting out of memory errors
I often wonder if there's all this talk about large language models making people able to get stuff off the ground faster and, and a lot of the rhetoric around that kind of reminds me of the NPM stuff
Computed from the transcript - who did the talking, and the words that came up most.
Richard talks with TigerBeetle founder Joran Greef about how their team's database achieved a wildly successful Jepsen Report. Where other databases have been famously skewered in the past, TigerBeetle passed with flying colors, and Richard and Joran discuss the story of how they got there, how TigerBeetle makes money when their database is open source, and many other topics around software quality in general. This episode was sponsored by mailtrap.io - modern email delivery for developers. Try Mailtrap for free: Patreon supporters get ad-free episodes! TigerBeetle's Jepsen Report - MongoDB's 2020 Jepsen Report - Zig - TigerStyle software development methodology - Deterministic Simulation Testing (DST) - Antithesis -
Transcribed and scored by The B2B Podcast Index.
Speaker A: That was the one bug that he found. But otherwise he didn't break Tiger Beel's trixurizability. And also what's interesting is what was a first he added new um, nemeses or storage fault injectors into Jepsen. If we wrote to a data file on one machine, um, he would take that write and move it to another place in the data file or move it onto another machine's data file. So he was literally playing with the cups. So this was just a normal Jepson test. And then now he was also playing with the cups and Tiger Beetle had to survive and it did. Not many people know this, but he actually could also use antithesis. So he could use their DST against us. Um, and he didn't find anything like that. Uh, which is the interesting thing. Um, it was the human, it was Kyle that found all the interesting findings in his report was pure Kyle. Kyle really gave it everything he could. But again, skill of the human, uh, he was phenomenal.
Speaker B: Welcome to Software Unscripted. I'm your host, Richard Feldman. Today I'm talking with Jeroen Greef, founder of Tiger Beetle, which is known for having developed a legendarily reliable open source distributed database for financial transactions. We talk about how they achieved such a high level of quality using things like ZIG Deterministic simulation testing or DST for short. Their unorthodox Tiger style programming style. And even subjecting their database to the also legendary scrutiny of Kyle Kingsbury and Jepson, which has absolutely dismantled the quality claims of other popular databases in the past, but which earned Tiger Beetle a glowing review. We also get into broader topics of software quality, the economics of open source and where the real value comes from for software in general. I want to give a massive thank you to everyone who's been supporting software Unscripted on Patreon. If you enjoy these episodes and you'd like to become a supporter too, you can get ad free episodes by Signing up@patreon.com SoftwareUnscripted thank you to Mailtrap for sponsoring this episode of Software Unscripted. If you don't know Mailtrap or you only know them for email testing, they are a modern email delivery system for developers. They support things like straightforward integration into your code with native SDKs or they have a security compliant API and also SMTP access in their free tier. You get 4000 emails monthly so it's pretty easy to try out. And they also have 24. 7 support where you talk to real humans, not AI chatbots if that's the type of thing you're interested in, because you do email delivering at your business. Check out Mailtrap IO to learn more. And now here's Jeroen Grief. All right, Joran, welcome back.
Speaker A: Hey Richard. So good to be back. I think. Um, it's been about two years.
Speaker B: Yeah. And something that happened in the last two years since we talked was the Jepson report coming out about Tiger Beetle, which was incredibly impressive because I always thought of Jepsen reports as being something that would just skewer databases. I think about the famous mongodb example just left and right like here's this problem and that problem and this guarantee didn't hold up and that correctness issue and this security flaw. Uh, but yours was not like that at all. I mean they did their usual deep dive and uh, I was really impressed by the result which. How did you feel about getting that result?
Speaker A: Yeah, thanks. I really appreciate it. So I was also really happy. Um, I should say I could tell you how I felt during the audit. Um, that was quite something because you do feel like you skewered. Carl is so good. Uh, he has an instinct for where to go in the African safari to find the lion, to find the bugs. And uh, he, he was just so good in building a reference implementation of our state machine. He did almost the most comprehensive implementation that he's ever done for any report. Um, because he didn't just try to break Tiger Beetle's strict sizeability just to break the consensus. Normally he achieves that for almost everything. Um, because that's one of the hardest problems in a distributed system is um, is the distributed system so the consensus protocol and you just have to find a way to break that and you break everything. Um, so he checks linearizability and all these things. M But he actually didn't only try to check that. He tested almost Tiger Beel's entire product surface area or entire API. Uh, and he re implemented everything in a second independent implementation and then differential fuzz both. Um, so he did actually find one bug on the read path. So it didn't affect durability. Our uh, query eng and didn't always return all the results. Maybe the tail would be less and then it paginated so you'd go back and then you'd get it. Um, so that was bad enough, that was the one bug that he found. But otherwise he didn't break Tiger Deal's tricks risability. And also what's interesting is what was a first he added new um, nemeses or storage fault injectors into Jepsen. So while he was running the distributed database, like these three cups, he was taking the data files and he was moving them around, around underneath Tiger Beetle. So he was corrupting Tiger. Uh, Beetle's designed one of the first databases to actually, uh, not only make real time backups, but test them and heal them. So in other words, it can survive from corruption or recover from corruption. And so Karl was actively. If we wrote to a data file on one machine, he um, would take that write and move it to another place in the data file or move it onto another machine's data file. So he was literally playing with the cups. So this was just a normal Jepson test. And then now he was also playing with the cups. And Tiger Beetle had to survive and it did. Um, so he found a very easy bug in another part of the system. But, uh, the core foundations held firm. And I think what's interesting is, um, that we built everything in Tiger Beetle to very high standard. We didn't take anything off the shelf. So we built every line of code is handcrafted, very high standards, tight, strict tolerances. And we got to production in three and a half years. It was about four years by the time Kyle started auditing. Um, which is also one of the shortest, you know, durations for shipping something to production. Um, and it passed as well. So he, you know, you couldn't really break it in the ways that count. And so I think, yeah, that's just, um, we weren't lucky. You know, that's just tribute to tigerstyle. The, like the philosophy, philosophy behind or the methodology behind how we, you know, actually built Tiger Beetle. So we didn't just engineer Tiger Beetle, we engineered the way we would engineer so that we could pass Jepson, which was, that was the goal. But yeah, it felt pretty nerve wracking during the audit. There were times when, you know, he thought he found a terrible bug and we realized that he hadn't. Uh, and so you have these moments where you're like, um, yeah, yeah. And not many people know this, but he actually could also use antithesis. So he could use their DST against us. Um, and he didn't find anything like that. Uh, which is the interesting thing. Um, it was the human, it was Kyle that found all the interesting findings in his report was pure Kyle. And I think it's one thing to do dst, but the skill is really. How good is your verification? How good is your test harness? Um, because no DST harness will give you that you have to write your verifiers. So that was Karl's skill. Um, yeah, there was one that antithesis found where they said, oh, data loss across the cluster. And we realized actually on their side, the docker they were doing. We don't recommend docker, but it was doing like a RM F when you would restart the. And of course that will lead to data loss if you delete the data files in the test harness. Um, but otherwise, I think that was the scariest moment. Data loss in dst and we realized no. So even using an advanced dst, he couldn't break Tiger Beetle. Um, and we obviously have our own DST that is all in house. So all our own DST from the beginning has been also our own. Uh, but Kyle really gave it everything he could. But again, skill of the human, he, uh, was phenomenal.
Speaker B: Yeah, that's why, I mean, and certainly that particular human has a lot of experience breaking databases in all sorts of fun and exciting ways. So it's really impressive. Especially because, uh, as I understand it, one of his main things that he goes for is testing your claims. Like you say that it can do all these impressive things. Well, let's see if that's true or not. And it seemed like my conclusion from reading the report was he was like, yeah, it actually does. I mean, you're claiming some really, really strong things and you actually are backing them up in practice.
Speaker A: Yes. And what I really loved about his report is not what he said about Tiger Beetle, but what he said about view stamp replication and protocol aware recovery. When you put these two great protocols together, they actually work, like, correct. Um, and they give you more availability because you're taking your redundancy through protocol aware recovery and you're now able to recover from storage faults as well. And so Kyle says, look, all these protocols, they work. So I was so overjoyed to see a distributed database for the first time. A storage fault model. But Kyle's testing this and then he was saying, wow, I can't wait to move the cups around other databases. But I mean, that breaks everything. So. But yeah, all tribute to the protocols that we used. Um,
Speaker B: well, so I'm obviously very impressed by this result. Uh, but I have to wonder, do people ever come up to you and say, like, hey, maybe you've made one of the most resilient, fault tolerant distributed databases in history, but couldn't you have been even better if you wrote it in rust? Wouldn't it have been more memory safe? People must say things like that too, right?
Speaker A: Or if we wrote it in a memory Safe language like JavaScript or TypeScript wouldn't have been more safe. Or Ruby.
Speaker B: I think think of that. Much better. Much better. Yeah. I mean what do you say when people bring those things up?
Speaker A: I would say they're all pretty much memory safe. Okay. One is going to also test concurrency for you, um, which is a hard problem but you know, would writing Tiger Beetle in a memory safe language like Typescript make it correct and. Yeah. What do you think, Richard?
Speaker B: Well, I mean my thought would be that the types of things that you just described, like for example Kyle's um, simulating, uh, I guess. Well, not simulating just like actually creating data errors by moving parts of files around and things like that, that's not really related to memory safety. Um, but I can actually think of just taking Your example of JavaScript, I can think of examples of where switching From Zig to JavaScript would cause problems because now you have potentially out of memory errors that you can't recover from in the way that you can with Zig because I mean I happen to know that you have the static never allocate philosophy. So you're in very, very complete control of all of your allocations and in Rust you can be. But everything down to the standard library in Rust assumes that you have a global allocator that can allocate whenever and
Speaker A: you would have to leading you down the path of you not going to, you know, so yes you can, but you're going to. The ecosystem is going to lead you down the other path. So.
Speaker B: Right. I mean that's. Yeah. So like in, if you were using Rust you would need to fight against the way that Rust wants you to program in order to have the like memory, uh, not just safety, but memory. I don't know. Reliability guarantees that you have around like handling out of error conditions. Like concretely I would bet that if you had used Rust or you had used JavaScript that Kyle would have been able to find ways to break Tiger Beetle or cause it to go down by exploiting out of memory errors by doing things that would intentionally run it out of memory. And I think there's a pretty good chance that he would have succeeded in that because you're just really playing whack a mole if you're using uh, kind of anything other than Zig or maybe
Speaker A: C. No, I would agree. I actually hadn't considered that out of memory because we did explicitly set out that we don't want to have out of memory. That is not okay for a mission critical database that it could crash because it can't statically allocate resources. And I think what's also interesting is um, there's a few things. So uh, um, I think people put too much pressure on a language for correctness and separately they conflate memory safety with security with correctness. The three are totally different. So I think the most horrible security exploits today are in the NPM ecosystem. So I did some hacking in uh, White Hat in a uh, past life and if someone gave me the choice, here's some C code, here's some JavaScript, find a P0 bounty. I'm going for JavaScript, I'm going for the memory safe language that has npm culture 600 dependencies. Show me that one. I will find the supply chain attack. Well I'll try but that would be my bet. I would invest my time there and today I think that's where most of the security risk is. Things have moved on and supply chain attacks are going to be the next swing of the pendulum. And how do we solve those? And I think the other thing is people put too much pressure on a language for correctness. Uh, they conflate that with memory safety and they conflate memory safety with security. Correctness is far bigger than memory safety. So in a distributed system like Tiger Beetle, local memory safety actually what you're caring about is end to end correctness and global safety of the global system as a whole. So you have multiple computers. So yes, you could have something that can verify that one local component it doesn't use after free. But how do you verify that the global distributed system is not doing the equivalent of use after free and split braining the log, um, reusing a prepare slot for something else, um, that's a much harder problem and Rust can't protect you from that. It doesn't guarantee distributed memory safety uh, in the logical sense of the system as a whole. Because really what memory safety is is my state machine. The state will not be corrupted and yes other properties like we're not going to be able to do a physical exploit. But I think the other thing that people conflate m when they say memory safety is they're not clear. Do they mean physical memory safety or logical or both? And most Rust developers I chat to uh, don't realize it doesn't give you logical memory safety, it gives you only physical memory safety. So what do we mean by this? Um, so what was some of the most expensive security exploits in history?
Speaker B: Um, um, I'm going to guess off the top of my head. The log 4J1, uh, from a couple years ago was one.
Speaker A: Okay, okay, I think you win. But I was thinking of cloudbleed. Hotbleed. So logical buffer bleeds where, where a hacker is trying to read sensitive data and your server is saying to them, well, you can read my sensitive data. You don't even need a physical memory exploit. You can just do a logical memory exploit by telling my software logically to read from a buffer from the wrong offset. And Rust would have prevented, I believe it would have prevented heartbleed, but it wouldn't prevent all buffer bleeds because many of them are just logical reuse of a buffer once initialized. So, uh, think of a file format decoder where the hacker can force arithmetic overflow to get the decoder to read from the wrong part of the buffer and expose sensitive. So that's often what happens. But a lot of people don't realize that. And I don't mean to pick on Rust here, I just mean to elevate our, uh, understanding of these things because these are unknown unknowns. So we're all chanting security, security, but nobody is talking about buffer bleeds or are we talking about physical or logical memory safety or checked arithmetic? And checked arithmetic is so, so, so critical if we really. You have to have it enabled by default in safe builds. And every time I asked a Rust developer this, you know, do you have, you know, do your safe builds enable ticked roof? They say, I'm sure you know, or some of them will know. But actually, by default in a safe build, it's disabled. And what people don't know is in Zig, uh, and positively speaking, I think I'm sure that Rust will change this default, and they really should. In safe builds, it must be on. In Zig, checked arithmetic is enabled because it's highly dangerous that you just let integers wrap around with no consequences, because hackers will always abuse that. I would just love to. Some of the work I used to do was all about, how do you detect this in zero days? You can detect a whole lot of zero days if you just look for checked arithmetic exploits. It's actually quite easy. And that'll actually. Static analysis will help you find memory exploits too. So these are like my. This is kind of what I've learned. I love Rust too, and our team love Rust. So Matt clad on the will say, you know, Zig and Rust. Um, and he also tried hard mode Rust. You know, how can you use Rust for static allocation? Of course you can. Uh, and shortly after that, he joined Tiger Beetle he loved Tiger style static allocation. So just to be clear, both languages, very positive. But I do think as an industry we need to stop chanting and start thinking. Checked arithmetic, yes. Enabled by default and safe builds. Let's not conflate safety, correctness, security, they're all different fields. And also safety is more than correctness. So correctness just means that software is correct. Um, it works. Uh, safety is about what happens when it isn't correct. Does it kill people or does it shut down safely? Um, in my mind some people will have different definitions, but I see safety as what happens precisely when there are bugs. What does the software do? Um, and that's so important when you're in a mission critical domain. So there was a Knights Capital, um, I don't know if we want to go there, but the whole company went bankrupt because their software, it was correct most of the time. And when it wasn't, it wasn't safe and it spent all their money. They um, went bankrupt. Millions of dollars, uh, because the system couldn't autonomously shut down when it entered, um, when uh, it violated invariants. So these are kind of getting to all the Tiger style things. Um, but I do think in a distributed system you look at all the invariants, there's thousands of invariants that all have to hold or Kyle Kingsbury will break strict sibility. Like all a thousand of these invariants, we have about 10,000 in Tiger Beetle. If one of them breaks Richard Karl's report, the finding would have been different. And uh, I think for me, what I learned through this is that only one of those 1,000 is memory safety. So he only has to break any one of a thousand and the hacker is through. And memory safety is one of a thousand in terms of correctness. And yet our view of correctness I think is too small. It needs to be bigger as an industry. And safety is far more than correctness. So Carl actually found quite a few bugs in Tiger Beetle. They were awesome where he could get Tiger Beetle to shut down. And the interesting uh, thing there is that if Tiger Beetle hadn't shut down, he would have had himself a correctness bug. He would have broken our strict sizeability. But because Tiger Beetle realized someone's doing something weird, they're running the database on its head upside down or something, like they've turned the cups, uh, Tiger Beetle could pick that up and actually shut down safely. So he could get it to do that in about five different ways. And most of them the next thing he would have had is P0, but we stopped it There. So that's what I kind of love about Tigerstyle is there's memory safety. A thousand times bigger is correctness and about 10 times bigger, that big circle is safety. And that's kind of uh, how I see things.
Speaker B: But back to you, I mean it's so interesting how different projects have different tolerances for different levels of safety. Like you mentioned, the NPM ecosystem is full of exploits. People are actively pursuing them. They've found you can make a lot of money doing this. But of course a lot of projects are choosing to do that anyway. A lot of greenfield projects knowing that that's an issue are just like, yeah, well we're not as worried about that as we are about not shipping on time or something. And we think this ecosystem is going to be what gets us to ship fastest. I might personally disagree but uh, I can understand them making different trade offs compared to you making a database that's literally for financial transactions. The costs of uh. First of all, attackers are going to be highly motivated to try to figure out some way to exploit the system because they like money. And uh, if they're successful then the consequences are very serious. So if you want to become a critical piece of infrastructure for lots of institutions, uh m. You have to take things that seriously if you want to do a good job. That's what it means.
Speaker A: Um, even as a company, even a few years ago when people were joining for an internship ship on day zero, people are trying to hack them. They're getting highly targeted phishing, um, from me it's quite something. So now when people join the company we have to warn them. Like look, you're going to be fished on day zero or even before. So it's, it was pretty interesting. Like that would not have expected so soon, you know, but this was even just going into production and already that happened. So um, yeah, yeah, so I guess that's the question is like you know, um, does quality take longer? Or you know, maybe we've got different trade offs around time, you know, to delivery of projects and so therefore we use npm. What are your thoughts?
Speaker B: Well, I mean for me personally I actually have a personal no NPM policy in general. Like if I'm making a website I'm just like the first rule is just don't even install NPM on your system. You know, it's uh, I don't, I don't think it's worth it. I'm sure a lot of people disagree with me about that. Um, and that's, I'm happy to agree
Speaker A: with you, but standing next to you, they can disagree with both of us.
Speaker B: Um, but mostly just because like you mentioned, I mean there is this whole cultural thing about the ecosystem that I think just causes more bugs and performance problems than it pays for. Uh, you know, what you get out the box has a quality problem that to me is more serious than the benefit of having so much in the box. Um, absolutely. But others may disagree. Uh, I'm curious about, uh, like, so you know, Tiger Beetle is open source and you mentioned, you know, you have interns, you're hiring and things like that. Uh, so like, what is the, what pays for all that? Uh, is it since you know it is open source, is that you helping companies install and adopt it? Like consulting or how does that work?
Speaker A: So we, we operate Tiger below for people, um, smaller startups. It's, you know, open source is too expensive for them. It's like Goldilocks. So they love open source. They want to be able to run it locally on their developer machine, but you know, to spin it up, to upgrade it, to set up monitoring operations, et cetera. You know, one of multiple systems and you want to get to market and build a product and iterate and like, so do you really want a, you know, 24, 7 SRE team as a, you know, as a young company? Uh, it's too expensive, you know, with special skills just for one database. Um, you're going to use like Supabase and we have a startup program and we'll just talk to us, half an hour later, it's up and running. And now your engineers, which, uh, cost a lot, they can focus on product. So what differentiates your company? M most companies, we're not incorporated to have SRE teams to run Tiger Beetle. That's not the mission of the company. We're going to be Tiger Beetle operators. That's our mission as a company. So we do that and that's what we specialize in. So our customers, uh, work directly with our engineers. Our engineers understand hundreds of thousands of lines of code and they understand the system of a whole in our customers environments, their business, how that works, their architecture around Tiger Beetle. So usually it's not Tiger Beetle going down, it's Kafka or somewhere else. And our engineers can jump in and help with that. But primarily we're operating so cloud, um, usage based pricing. And then what we do for enterprise is they say, well yes, we do have SRE teams, but open source, I need to pay you because I need 24, 7 priority support if Tiger Beetle goes down and we're trying to page you. It's too late to escalate. We actually want you to page us. That's sort of what our customers say. How can you page our engineering team to let us know that you found a problem? Maybe beyond Tiger Beetle? That sort ah of blew me away when one of our customers asked for that. They want us to page them, you know, uh, um, yeah, so that and then enterprise also needs scale because it's one thing to do like 500,000 transactions a second. But where are you storing all of that? You can't store it on replicated mme. So they want to archive to object storage. And we do, you know, enterprise connector to object storage. Our principle there is we're a company first. Our uh, technical contribution to the world is Tiger Beetle is open source. Um, that's different from our product. So our product is Tiger Beetle operated at scale and you can operate Tiger beetle great on NVMe local replicated. But how are you going to write a file system driver to do tiering to object storage? That's what we do as well. So we have Tiger Wheel is the tip of the iceberg. And then we have a lot of code um, beneath that that connects you for petabyte scale. It's really hard to run. You know, people do, you know, hundreds of transactions a second which is actually very high scale, um, or thousands or 10,000 a second and all of them like they can't trip the SLAs in terms of latency. Um, some of the SLAs are like 60 milliseconds even for multi regional deployment which is like incredibly tight. So these are the things that our team work on and solve. Um, yeah, we don't talk publicly about customers but uh, yeah, we do have customers and we have a nice little business uh, that is growing.
Speaker B: Yeah, really cool because I mean, so it sounds like to summarize, you have a mix of open source. So Tiger Beetle itself is open source. You do have some proprietary stuff that supports that for specific use cases where a customer says I could write this myself as an add on to Tiger Beetle but I don't want to. I uh, would rather pay you to do that. Um, and then separately you also have uh, yeah, like ah, SLAs and being on call for them and stuff like that and setting things up and monitoring them. I like the example you gave at the beginning of uh, you know, a startup comes to you and they're like, well we could run this ourselves but we'd rather pay you. And I think Something you're maybe underselling a little bit is if I'm starting a startup where I, for whatever my startup is, I want Tiger Beetle. That's what makes sense for me on a technical level. Um, I would be scared that it's like, okay, I know Tiger Beetle itself has been vetted by Kyle Kingsbury and Jeffson and all this, and it's reliable. What if I set it up wrong? I make one little mistake in my AWS config, and now a hacker gets in, not because of Tidy Beetle, but because I'm m. Like, I don't want that. I want to go to the experts who made this thing and pay them, um, the system as a whole.
Speaker A: And that's. So we, you know, for, for big companies, we actually fly on site and we sit with engineers for a week and we accelerate them. Um, because that's cheaper for them than to try and like, read through the docks of like, you know, we actually just like, you know, Rafael or Federico or whoever, like, you know, all our team do this and they literally fly all over the world, always going, uh, I mean, some of them fly to, you know, North America and then immediately to Europe and they do the world tour. Uh, um, so. But that's what our team do is precisely to go. And like, we do a lot of review because these are big companies, some of the biggest in the world, and they're doing migrations. Like, they literally are migrating because they have to, because there's nothing that can power their scale but Tiger Beetle. So they have to migrate. And they, you know, it's been 30 years of, okay, we can't migrate. You know, uh, this is core. But the world is increasing and it's becoming more transactional. Um, even humans are doing more transactions. We can forget AI and autonomous stuff, but there's just. Everything is becoming faster and so, um, they, they have to migrate. And this is what our team help with and take them through that journey. And some of these projects take a year and then we get them to production. Like some of the biggest, um, you know, um, brokerages, you know, in. In. In a country, you know, like the second biggest, you know, and then, and then they're on. On Tiger Beetle. So it's pretty cool. Or, you know, some of the biggest wealth management companies for a country. So the whole country has their savings with this company. And, and all of those numbers are, you know, being tracked with Tiger Beetle. So it's, It's. It's a lot of responsibility. Um, that's, um. And yeah, so. But we, we love that. That's our duty, you know.
Speaker B: So, uh, yeah, and I think it's interesting that if you look at your sort of success story of making a company that can be very easily, it sounds like self sustaining in terms of financials and being able to pay for people. And you moved into a new office space. Congratulations. And you know, you're hiring interns and
Speaker A: it's a small office. This, you, you. Because we're a remote global company, so you, you're seeing here, uh, space for three people, just this little. The glass is not ours, you know.
Speaker B: Uh, sure, but I mean like, in contrast, I've also heard some stories of, you know, companies built around open source products that are shutting down or laying people off and you're going in the opposite direction. And one observation I would make is that everything you just told me about how the business side of this open source enterprise works is that you're doing something that's very, very hard. And it's not something that people can just pick up off the shelf and say like, oh, I'll just use this. Not because you've made it hard, but because the stakes are very high. Like you're, you're doing something that's very valuable. You've open sourced it, which is great, but it's not like just the fact that it exists as an open source thing, you know, means that the rest is easy for people. Um, whereas a lot of other open source projects, people look at that and say how is it that there can't be a self sustaining business around this? It's so popular. But to me it actually seems like that's maybe looking at the wrong dimension. It's not so much that it's popular as, as it is like, well, if it's used by a lot of people, but it's very easy to use and there's, there's no, you know, benefit to uh, having an expert help you out with it. Well, in some sense it's like how could there be a business around that? I'm not saying it's impossible, just rather that it's, it's uh, it's not as natural a fit as it is for something like this where the stakes are really high, the cost of a mistake is really high and uh, and the expertise required to do it, do it really, really well is also very high.
Speaker A: Right? That's right. And yeah. So you actually want to create value for people. Um, you know, I never like to use, some people say capture the market. I hate that phrase. You want to serve the market serve the community honorably at a profit. And that's great. Like business is wonderful thing on the right basis. And I think the best business has this basis. Serve honorably at a profit and that's sustainable. And then you can invest and things come out in open source. And a lot of our code is not open source. So people see the Apache 2 target beetle. Um, but some of the code that isn't open source is if an enterprise were to need it from us, it's kind of like the equivalent of let's write our own zfs. But that's the kind of decision you would definitely get fired for. If you're going to write your own object storage driver, you will get fired. Um, absolutely. So don't, you know, so I think it's pretty. This is the value that we bring is we bring incredibly. Um, you know, durability is sacred for enterprise. If you mess this up, you know, if you're writing your own, um, NVME drivers, um, you know, let us do that for you. Don't try and have your SRE team now get so excited that they're going to write this themselves. Um, because you will get fired. And that is our values. You know, this is our job. Like that is what we do. Um, so I kind of always love that saying, nobody got fired for IBM, but people certainly got fired for writing their own file system drivers. And that is part of our value. So, um, we had. Yeah, yeah. Um, yeah. So I think that's the thing with business and open source. But if you have business motivating you and propelling your team and your team love to serve honorably at a profit, all three, then what is great is the, that the business motive actually drives the engineering motive. And now you've got a healthy business motive and now your engineering motive is. Well, absolutely. We're not using NPM because we care. We want to give quality. And I should also say I did grow up in JavaScript. I'm really thankful. I learned a lot. Uh, I was writing um, Rhino on the JVM before Node came out because I was absolutely convinced the world would move to server side JavaScript. I thought this is inevitable. So I was doing that. And then Node JS came out on Hacker News and on day one I was in the community. But over about 10 years. Richard then I realized is actually too hard for me. My skills are not good enough to write a production API in Node js. It is impossible for me. Other people can do it, but I'm not good enough. There's just no ways I can write an API with explicit limits that can handle overload DDoS and stuff. You. I know, I mean, I was doing stuff where you're like patching the VAGC to try and make an API node JS production grade. And um, I was doing stuff like Martin Thompson in Java where you're not even using pointers in JavaScript, but you have one big gigabyte buffers and you're doing pointer arithmetic in the buffer and at some point you just realized, I would love a language with pointers. Uh, please, can I just have first class pointers. We have first class functions. But now, please, just first class. It's just going to be easier for writing stuff at scale. And so, yeah, and I think now, I think a lot of the reason why people went into NPM was because the tooling in C was so terrible. Like, how do you compile a C program? And then, okay, we're going to rather spend the next 10 years of our life writing JavaScript than trying to figure that problem out. Um, but that problem isn't a problem anymore. Zig is so easy. You download the compiler and you write some code. You can cross compile a binary for Windows or whatever obscure architecture. Um, it's wonderful. So these days I think the best language to be learning, it's so simple and powerful. Great path weight ratio, I would say Zig, you know, I love rust, I love JavaScript at the time, but Zig for me is really like the sweet spot of power to wait. Um, and I think it's a healthy start for young engineers to learn. Zig, you're going to get the, you're not going to be learning half of programming, but without pointers, you're going to be learning the whole of it. Um, so yeah, I'm excited for Zig for how it can teach the next generation of coders. It's so readable, like Python or Typescript. Um, uh, but I think we need to build these mental models anyway. I can wax lyrical about this.
Speaker B: Yeah, well, this calls to mind, I mean, going back to our, our previous discussion about sort of there being like levels of needs and things around security, but also maybe scale. I'm willing to bet that pretty much nobody who has used Node JS and is listening to this has ever tried. Hey, for my server side API, I'm going to allocate a giant gigabyte array and manually do memory management inside of that, like you, uh, just mentioned earlier. Uh, but I would assume that the reason that like what drove you to that is that, that you were doing things the normal way and you just ran into production problems. Is that right?
Speaker A: Absolutely. Huge production problems, like two day outages. And they were not the fault of my code. It was VHGC that was pausing for two minutes, uh, literally. And that was the hardest thing because you think, but the platform must be, you know, it's like kind of like you're doing mathematics and suddenly you realize that all the first principled axioms are actually wrong because you always assume the axioms are correct. Uh, and suddenly you discover, oh, the GC doesn't work. And now there's nothing my code can do. Like you have to like hide stuff from the gc. So. Yeah, so, um, that was my experience. I ran so far. You know, even in Node JS you can have memory fragmentation. That's the real problem. If you just use Zlib. I don't know if it's been fixed, but um, yeah, I found some wonderful performance issues in Node js. Yeah, I believe the CPU pool is still four threads or uh, that the thread pool is still four threads. And so if you're doing DNS lookups, those are being dropped into that four thread pool. It's so easy for an attacker, uh, just to get your system to resolve some poisoned DNS resolver, um, that will be a tarpit and just consume all four threads of your application and now you DDoS. And people would never know that this is happening. You just got some tardy DNS um, blocking your whole thread pool which is only sized to 4. I don't know if that has changed but I mean there were so many issues. Please let's a thread pool. We should be able to size it for CPUs and let's separate IO DNS from CPUs because the runtime, the latencies are. It's like racing Formula one cars and trucks on the same racing track. It's not a good idea. Let's have two racing tracks, two thread pools with different performance characteristics. But the whole even like Fedor and Duttny, he was so awesome. I mean they were great people in Node, so I am thankful to it. It was a phenomenal learning experience of how not to write production software. Uh, and many of us benefited and we had jobs and great, but I think it's a cul de sac. Um, and for example Fedor Dutney, he showed you could double network performance for Node js. Uh, and that pr, I think it didn't go anywhere and it should have. He was doing great and all it Was was bring your own buffer. If you want to read from the kernel's TCP receive buffer in Node js, you have great buffers. The API should allow you to bring your buffer receive into it. And he did a PR to show that it could double performance for the world's Node JS servers. And it never happened. Um, and I mean, there's reasons, because the project is not going to change. These are fundamental changes. So it's too late now. Um, but, yeah, so I think, yeah, it's just, um, that I ran away from that experience and I looked for a language where can we just do something as simple as bring a buffer to receive into intrusive memory, an intrusive memory pattern. But it's bring your own buffer byob. And Zig was perfect for this because you could not only bring your own buffer, but bring your own, uh, allocator to the standard lib. The standard lib isn't going to use globals. And you think, think, wow, like a standard lib that doesn't have. I mean, surely this is like programming 101. The standard lib shouldn't be reaching out to global somewhere. But almost all standard libs are broken like this. And finally, Andrew is like, he really cares. He's like, well, this is a very simple problem, but it's very important to fix. Let's bring our own allocator. Now let's bring you an I O. So I was so happy for Zig and so thankful. And I loved sea as well. Uh, and C, you know, you had all these knife edges, like walking blindfolded on the cliffs of Dover and you're going to fall off. Um, and Zig started adding guardrails. Um, obviously you can still jump over the guardrail if you like. Um, the borrow checker would stop you, but if you really want to jump over, you can. But Zig, you pretty much have guardrails. Um, and there's ways that you can design. I think this is the other thing. Um, the problem where all the memory issues are really a design problem. So this is coming back to where we began, is that people put too much on a language for correctness. It's kind of like saying, I'm going to choose the language I grew up speaking as a baby from my parents. I'm going to choose whether I speak English or Russian, because if I speak Russian, I'll become Nabokov, and if I speak English, I can become Hemingway. So I'm expecting English to make me a great writer. Um, and obviously that logic doesn't hold. And as programmers, yet we do the same thing. I'm expecting to program in JavaScript memory safe because it's going to make me secure and correct. Um, and JavaScript absolutely is not going to do that because correctness I think is not a language property. It's a design property of the system as a whole, end to end. It's uh, in a distributed system, the language doesn't even span the machines, uh, the compiler doesn't know about them. So, so there's no way it can help you. So it's like it's really a design problem and you have to think of methodology. So I think the way we get to correctness is yes, language being readable, explicit, like sig checked, arithmetic enabled by default. Yes, great bounds checking. Rust too has got phenomenal. It does bring to the table, but it doesn't stop there. And I think that's where we make the mistake. As the industries, we have these language wars and actually we're forgetting that correctness is a systems thinking problem. End to end design. Are we writing fuzzes, are we doing differential testing, are we doing deterministic simulation testing? Do we have assertions? Um, all these things. So that's the stuff that really makes for safety, uh, and correctness.
Speaker B: Yeah, I think, I mean as someone who likes languages a lot, I totally agree with what you're saying. I mean it's definitely not the case that you can just pick a language off the shelf and be like, oh, this is the correctness language. Well now I will have correct code. Um, at best a language can help you with that.
Speaker A: You like languages a lot, Richard? I've never met anyone who knows more languages than you.
Speaker B: Oh, I have. I'll name Dietsch Aditya, uh, Sirom. Off the top of my head. Uh, he definitely got me into languages and knows more than I do. Um, Hillel Wayne also probably knows more, uh, like a longer tale of languages
Speaker A: than I know, but he was telling me about, he's been going into all the history of GOTO and all the variants of that. Um, um.
Speaker B: But I mean I think an interesting parallel to your comment about English and Russian is um, you talked about how like in Zig, you have in the standard library this design invariant where, where you don't allocate, you just say if I need to allocate memory on the heap, then give me an allocator, it'll allocate stack memory. That's it. Uh, uh, and also now it also is the same with IO, it's like hey, Tell me how to do I o. I'm not even going to assume that I know what the file system is or that there even is a file system, things like that, uh, both of which are really cool, useful invariants. And uh, it reminds me of functional programming in the same way where there's this culture around, even in languages that don't have it as a first class language, guarantee they'll still say like pure functions in the standard library. That's as much as possible, really, really avoid any kind of side effect. Um, and one way you can look at those as sort of two sides of the same coin is that uh, when we talk about pure functions, um, there's a little bit of hand waving that has to happen because CPUs don't have a concept of functions, let alone pure functions. What CPUs do is they read registers and mutate registers. It's just side effects for days. So there's, you know, this all has to be sort of an abstraction on top of that it's a conceptual pure function. Um, and similarly when you think about, you know, allocating memory and things like that, you can pretend that allocation just always works and always succeeds and is always fine. But of course in practice that's not the case. Now uh, there's a similar limit when it comes to stack memory memory because even in Zig, like, you know, for practical purposes you do kind of assume that you're not going to run out of stack memory. But you could not, not in Tiger Beetle.
Speaker A: But well we, you know, we, we're very careful with recursion in that, but we wish we could be more disciplined about this in how we, you know, actively measure our stack usage.
Speaker B: But um, yeah, I was actually wondering if that uh, if you have any special tools around that to try to create it to go from something where it's like, well, well this shouldn't happen. But it theoretically could because I mean in theory if someone could find a long enough chain of calls, even if there's no recursion involved, where the functions are not getting inlined and you don't have enough registers, even if there is inlining where you have to go to the stack, and the only possible way that this code can compile is if it compiles to something which at runtime could potentially, theoretically blow the stack. Uh, what do you think?
Speaker A: Uh, you got me there. Now we're with you. We have the same like, desire and I think Zig is working towards this because um, you know, the early Async version, it was pretty Cool. How you could get a frame, you know and see oh this is the, you know the size of the stack and obviously we don't use that in Tiger Beetle. We didn't at the time. And the new one we don't. We do our own callback style because we like to be very explicit about the memory that survives um, through the lifetime of the asynchronous function. Um but we also use very um, our code is very simple so the control flow is by design is meant to be simple and quite flat. And so we do know if there are places where we could get a bit of a chain or a graph, then those places we rewrite the style so we will use a state machine where the code also becomes more readable and so that you're not um, so we approach it with simplicity, discipline. We're actively aware of this but um, we also don't do recursion following um, NASA's power of 10 rules for safety critical code. So we don't do recursion as well. Um and so it's not really a problem for us. It um, would be cool if we could measure and sort of monitor stack usage and be decreasing that just as a waste thing. But um, it's not a. We, we also run our ah, simulators. We have a fleet of 1024 cores dedicated Ah, they're in Finland so they've got you know renewable uh, it's very cold place and so the, the cores are running 24, seven and each one of them runs the Wapper. Our, our own in house deterministic simulator that we've always used. That, that's the thing that really got us to um, production in three and a half years. And this simulator runs across a thousand cores all the time. And a second there, there's a speed up factor of around 700. So one second on one core is worth 700 seconds and we've got a thousand cores so it's like 2000 years a day of simulated runtime in all different ways. And so again if we did have a stack thing this would probably catch it. Um, I should actually ask the team if we've ever had one. Um, uh yeah but uh, give me one second. Uh, Fred, have we ever had, has WAPPA ever caught a stat overflow?
Speaker B: Don't think Wappa has.
Speaker A: No. No. Okay, thanks. So yeah, we've never, I mean we've found hundreds or thousands of bugs but we've never, the fleet has never caught her. Which actually could also mean that our uh, testing is just not good Enough that we should be finding it, but, uh, we aren't. You never know. But.
Speaker B: So yeah, I do remember in your blog post about the Jepsen report, you did note that the one correctness bug, uh, was actually something that the fuzzer missed. So it is possible.
Speaker A: Right, Exactly. And that one, we wrote that whole blog post. So that was that bug that I referenced in our query engine where we, if you were querying data out, we wouldn't return in a page, we wouldn't return all the data. If you went back for the next page, you would get it. Um, but that is terrible. So I told Karl, well done. You've broken out correctness on the read path. And if you query a database to get data out and it doesn't give you the data, well, it's almost as if you never wrote it in the first place. So it's very bad. But we did fix it immediately and it's fixed and it didn't affect durability. Um, Karl thought it wasn't so bad. I said to him, no, this is horrible. So, um, yeah. And that we had I think three fuzzes and it made it through all of them. But that was why we employed Kyle, because we want an independent auditor, um, to find. So you know, I think you and I, we were in person in New york in uh, 24. Kyle was there as well. And we went for coffee after that with uh, Oscar Wickstrom, who's working on Bombadil, uh, now with Antithesis Us. But um, Kyle was there at coffee and I actually said to Kyle, like, you know, what can we pay you and maybe you can help us find bugs. So we're very happy, like all the bugs he found and we'll do it again with him because he's good at finding bugs. Um,
Speaker B: yeah. I often wonder if there's all this talk about large language models making people able to get stuff off the ground faster and, and a lot of the rhetoric around that kind of reminds me of the NPM stuff where it's like, um, well, you know, you can get stuff off the ground faster as long as you're okay with the downsides being quality performance, security, et cetera. Um, and I'm not really excited about that personally. Obviously a lot of other people are and I'm sure they will make money doing it, but uh, and probably lose some money. Some catastrophic outages too. However, uh, be some big market crashes,
Speaker A: you know, as the, as the valuations of some of these companies. Correct. Like, um, some. I, I, I kind of think of these things as like, you know, additive increase, multiplicative decrease, like TCP congestion control. Capital, uh, is flooding into the market and some point someone, you know, people start paying twice as much for engineers as they really should be, like half a million dollars. And then the market realizes this and then it crashes and then goes up again. And eventually we realize like, oh, LLMs are a thing, but maybe, you know, we're a little bit too hasty.
Speaker B: Tripe.
Speaker A: Uh, it says don't, don't be hasty. Uh, but so I'm sure there's value. But like any market, it will have to correct and so that'll happen too. But yeah. Sorry. Back to you, Richard.
Speaker B: I get frustrated, uh, with what I see as, um, mismatches between rhetoric and reality, of which there are many right now. And one of the like, big top line ones is I just ask myself, like, of the software that I'm using that I see personally as, uh, just thinking as an end user, not as someone who makes software, but just as someone who uses software. Like, what trends in output have I seen and over the past like three years? And you know, as soon as I say this, people are like, well, anything that was before November 2025 doesn't count because that's when Opus 4.5 came out and we crossed the threshold and like, okay, we'll see. Um, but I mean, I have definitely the only trend in output that I've noticed has been a, uh, decrease in quality. I've noticed software getting buggier and that trend already existed, but it definitely feels like it's accelerated in the last few years.
Speaker A: Are we being autonomously worse at quality? We were humanly bad at quality and now we're like autonomously really bad.
Speaker B: Uh, m. It seems to be accelerating that problem. But I also have not seen this acceleration in releases. I have not noticed that new features, or at least not features that I notice or care about about coming out faster. Um, I complained about this on Twitter and someone responded, uh, with some examples that they had noticed of software, uh, that I also use. And I was like, okay. I mean, those are new features. I didn't know they. I kind of looked them up. I'm like, I don't think I'm going to use any of these. But okay, fair enough. I mean, they did ship them, I guess. Uh, so maybe that's a perception problem of mine. But I don't know, when I talk to people, I don't hear people saying things like, like, software is getting so much better. Software is getting. I'm so Much happier with my software. Um, so if that's the output of the system, how much does it matter how many PRs per day you're landing if you're not getting to the output of people being happier with the software?
Speaker A: I'm so glad you say, I was thinking about this this morning. I think as an industry we're optimizing the wrong variable and when that lands that like all this high billions of dollars investment, we're optimizing the wrong variable and it's just basic economics. Why is software valuable to the world? It's because you can write a piece of code once and sell it to the whole world. And you created um, the BASIC compiler and then DOS and Windows 3.195 and your bill Gates because you latched onto this idea of software can be a scalable business because you write something valuable for many people and they will run it for many years. So I kind of think, of course you get bespoke software. You get agencies that do WordPress, websites that are ad hoc and custom. So I'm not talking about bespoke software because that isn't really valuable. If you're creating bespoke software for people, you're going to be paid for your time always. And in that industry for sure, LLMs, they're going to take your job away. So then you won't even be paid. Um, I think I agree with that and I think that's great because now we're going to, yes, we're going to get all this bespoke software, wonderful. But the world of software that I, that you and I, I think are really interested in is software that has many users. So you number of users, um, how many users are there? There's a lot of users. It's the whole world using the software. It's like infrastructure software or Linux or whatever, you know, SQLite. Um, uh, everyone's using it. So a lot of use, um, and then for a lot of T, they're going to use it for many years because this stuff is infrastructure, it's got a long half life. Postgres 30 years old, you know, MySQL also basically my uh, SQLite not far. Um, so these things have a long half like big T, big U. Um, and performance is very important, you know, and safety is very important. And so you multiply these things and then you don't want to have production outages because that is huge cost. So as a developer I have a certain hourly rate. It, um, but if I write software with big U and big T And I break the whole Internet for a day. That cost is way more than my hourly rate. And I think this is the economics that everybody is missing is that software is valuable because it's scalable, because software is eating the world. You know that famous saying, and the world is eating software. Um, but the world actually wants the software to taste good, good. And if it isn't, that's bad. Because now software can be highly scalable in a good way. It can also be very dangerous in a bad way because you don't want to have blue screens of death or your software, your operating system gets slower and slower every day and it gets viruses and whatever and then you switch to Apple, um, or then you switch to another system, sorry, uh, to mention product. But then you, that has a cost too. And so kind of, I think with software the lesson we need to learn is that ambition bites the nails of success. So as developers we think of our own productivity and that's the wrong variable to optimize. If I take a day to write something or if I spend an hour extra and get rid of blue screens of death for the whole world, I mean that hour, what is an hour amortized over 30 years of production usage, it's free, it didn't cost anything. And this is the thing with software development costs. Developers are not expensive. They're only expensive if you're a WordPress agency, then developers are expensive and they will be replaced. But if you're in any kind of business that's creating value in the world and serving many users for a long time, Big U, big T. Developers are not expensive. They're a linear salary, uh, creating an exponential asset. And so it's not about linear cost of development because you can take 10% longer relative to the exponential value you're creating. Uh, you know, if you take 10% longer and add an exponent, it's way worth it. So this is just basic economics that everybody's saying, you know, LLMs are going to make you code faster. They forgot, like why didn't we get into software? It's not about our uh, time, it's about the user's time. Serve the community honorably at a profit. Don't, don't try. And so I think, you know, myself as a developer, I try to get more and more productive and it was like a trap. You know, I started writing higher level languages, I thought they'd make me more productive. And now it's getting to the point where it's like, well, you node Js, you're so productive. Not even node js. Now you write in English, you know, you will be very productive. You can write in English. And at a point I'm like, no, I realized that, you know, I actually want to do things in the most direct way possible. I want to be explicit. So I would rather write my prompt in zig, thank you very much. I want to be explicit on what I want. It's easier for me just to go directly, you know, um, because it, you know, Tiger Beetle didn't take long to get the design and the basic thing working took like a month. The first prototype took five hours, the next one a month, the next one six months. And in total it was like three and a half years to production. And it can power countries. So three and a half years is really like someone could do it in two years. It's not material, you know. But I think the point is that we did it faster than anybody else else also. So we didn't take longer. We got higher quality in less time with tigerstyle. But I think now, um, I can carry on. But I think LLMs are exciting when people realize the variable to optimize is quality across all the users across T. Um, that I would love to see LLMs being used to break software. So people are thinking of it as constructive. I love using them and in a destructive way. So I will write a talk, for example, and I'll say to the LLM, tell me what sucks, what's really bad, what are people going to misunderstand, what am I missing, what are my unknown unknowns? And there hallucinations don't matter because if it finds a bug, it finds a bug. I think that's the interesting thing for LLMs, um, is the increasing quality and the degree to which it can do that. Um, yeah. And maybe, you know, maybe yes, it will write better than human. But at the end of the day it's not the writing of the code that dominates distributed systems. It's not the creation of code. Also people have totally missed this writing a, uh, distributed uh, database. You can do it in a month. The time sync, you all know, Richard, is the testing, the maintenance, there's all the bugs. So if you don't have a deterministic simulator, you'll have bugs that take your tears to fix, you know. And um. So yes, LLMs can now help you there. So Jens Expo with Iearing has been using LLMs to reproduce bug reports and fix them faster. So that is now cool because it's increasing quality. But um, yeah, I've been thinking about this For a while. What makes software valuable? What do we optimize? Um, So I think LLMs are exciting, but I think the whole industry has gone down a cul de sac. English is not the most productive language for coding and coding time. Even in our software lifecycle. Most of the cost is production outages, uh, incident reports, fixing things, maintenance, testing, and I would love to see that. But those things you can already solve with dst, which is like autonomous testing, but you don't have to use an LLM.
Speaker B: Um, yeah, I mean, I would love to see a cultural change in the industry where, where, um, people are using LLMs and just in general to try to improve software quality as opposed to just getting lower and lower quality software out the door faster and faster. I am a little bit optimistic that maybe, uh, that might just naturally happen because if it becomes so cheap to get something low quality out the door, then you just are not differentiated anymore. So it's like you put your low quality thing out in the world at breakneck speeds and everyone's like, like, who cares? There's a trillion of those. And so the thing that makes you stand out and makes people choose your thing over somebody else's is it feels less frustrating to use. People want to, you know, your, your pitch is, I'm going to take your pain away. You're, you're using this low quality thing that's slow, it's buggy, it's unreliable, uh, it's confusing. And you're saying, look, this will be fast and it will not break on you, it will be reliable and, and et cetera. And at some point people are like, you know what, I will actually pay a little bit more, uh, to switch to that because I don't like feeling frustrated. And maybe at that point we start to see that, uh, what are LLMs capable of when it comes to quality? And I love your example of using it to break software because trying to figure out where the problems are not just in distributed systems, but in all sorts of different types of software is in a lot of cases a really important part of achieving quality is like if you don't know where your quality problems are until end users slam into them them, um, there's going to be a lot of frustrated users between you and actually achieving quality.
Speaker A: Oh, so well said, Richard. I like, it makes my heart burn, you know, and because I think it's also just basic business sense, like, what do people buy from us? Do they buy our development time? No, they buy the value, the quality, the experience. Like that's what we always were selling. You know, there was someone on, you know, um, someone in the venture circles, you know, online, and they said that if your cost of coding goes to zero, what are you selling? But the problem in their premise, there is an assumption that you're selling your cost of coding. And we never were. We should know this stuff, like, and these. Yeah, it's just phenomenal. Um, it's kind of scary to me. That's why I think there will be a market adjustment coming, coming, because we have really marketed the wrong variable being optimized. And actually, businesses are value based. They price according to value, not according to cost. So if your development cost goes to zero, frankly, it doesn't change the value of your product. It shouldn't. Because otherwise you've got an economics problem. I think, um, yeah, it's the quality of your ideas and your execution. However you do that, you can use an LLM, but, uh, it's the end result that matters. So how can LLMs be part of a better end result? M. But I think I'm just saying the same as you.
Speaker B: Uh, totally agree. Uh, and by the way, I know, uh, you got a hard stuff we got to wrap up, but, um, I do want to make sure I thank you because, uh, in Rock's rewrite, uh, to Zig, we actually use some of your stuff from Tiger Beetle, um, some of your testing tools and, uh, lints and stuff like that. So thank you for open sourcing all that because we're using it and we've gotten value out of it. So we appreciate it.
Speaker A: Oh, thank you. I must say thank you to you because Rocks is such an incredible project that, um, I'm also, um, flattered that you folks are using it. And thank you to you also for, uh, investing in Zig and in the ecosystem and spotting a wonderful, like, it's a big wave, you know, you're a surfer and you saw the back, you saw the swell coming and you, like, pedaled out. You didn't look for a wave. Uh, how many surfers are riding the wave? And now I will paddle out to it, but, you know, we paddle. We saw the swell in 2018, 2020 already, long ago, you know, and then we, like, paddle. We paddled it and we met each other out there and we've just had fun riding the wave, you know. And so thanks for being a fellow wave rider.
Speaker B: Absolutely. I'm excited for, uh, you know, I have to be careful about how I talk about rock. While we're not quite done with this big rewrite yet, because on the one hand, it's very exciting and we've got a lot of stuff working and, and people are starting to use it for things, uh, which is great. But on the other hand, I'm also like, yeah, if we gave this to Kyle Kingsbury, he would be like, this is not even done yet alone. Uh, you know, uh, completely bug free. So I think on the quality front, we have a long ways to go with our rewrite, but it's gone very well so far and I want to live up to that before I start bragging about how awesome it is because we got to be honest about where we are.
Speaker A: Maybe would you ask Carl if he can order the tweet you at some point?
Speaker B: That would be really cool. I mean, it's not a distributed system, so it's not necessarily in his wheelhouse, but, um, definitely I would be very excited about finding out. Are there pathological paths through the compiler that can get it to miscompile something, or are there, uh, ways that you could break some of our security guarantees around, like these functions? All must be pure. Um, we have by design, like no FFI and stuff like that. Uh, uh, yeah, I mean, hopefully we have designed our way out of that. But I'm, I don't have so much hubris to think that we, you know, could uh, possibly survive a Kyle Jepsen audit of our whole, whole code base, um, without him finding anything.
Speaker A: It's like going base jumping that moment, you know that how you feel as you jump off a mountain. Uh, it's terrifying as you engage him. But I would encourage you. I mean, I'm sure you'll do really well. So I can't wait if you could, because he does. I know it's not distributed, but he has that instinct, you know, so if you said him find functions that aren't pure, he would go, he would just be excited by that.
Speaker B: I mean, right now I'm sure he could find any number of problems. But yeah, once, once we get it to a point where we feel confident in it, I would love to. Yeah, maybe I'll reach out to him. I don't know. We don't really have budget for that also because, uh, we, we don't have a business plan at all. It's just we're, we're doing it for the love of the game. But, uh, yeah, it would be really cool if, uh, if we could get his hands on it. Um, once it's at the point where we're like, okay, we think it's good now.
Speaker A: Hopefully Carl's listening and he's like, yeah, we all three fellow ST speakers, let's pay it forward and I'll. You know, he did that interesting report recently on Galera MySQL Galera just, you know, of his own for fun because he wanted to show it broke so many, um, almost all of the things he could possibly not get right, you know, in consensus. Uh, but I think for you it would be very different, obviously.
Speaker B: So Jepson's written in Clojure, so uh. Obviously he's a fan of functional programming. So. Yeah, yeah, maybe. Um. I would actually love to talk to him. Uh, I have any number of things I would ask him. But uh. So maybe we could chat about uh, making that connection. Because that'd be really fun. He'd be a really fun person to talk to.
Speaker A: Be so refreshing because it's so lateral. Like it would be taking his thinking and applying it to a compiler, uh, language.
Speaker B: I mean, I think we're like a couple years away from that being realistic. But um, I mean we don't even have the LLVM back as of this recording. We don't even have the LLVM backend working yet. So we have like some machine code backends, but not LLVM yet. So definitely not done. Um,
Speaker A: it will be one day you'll look back and it's done and then you've made systematically the right decisions, you've invested in quality. You did the hard thing today, but you really made tomorrow easy. NPM make today easy, tomorrow a nightmare. And if we're thinking of quality, it's like, well, when we reach production, let's make that easy and let's make it maintainable, you know, and fast and safe and secure. ROK is pretty awesome. Yeah, I mean ROC is like a shining star like that.
Speaker B: The philosophy around that, that's where we're aimed. I really. You, uh, know I'm pushing hard to get there, so I don't want to claim we've gotten there before we have, but that is absolutely what we're going for. So, um. Well, thank you for being an inspiration for, you know, like shipping something that actually like not only aimed there but then got there. Uh, and I think the uh, Jefferson Report is like a really strong validation of that. And, and uh, tigerstyle, which uh, I know we kind of talked about in bits and pieces, but a lot of people have talked to me about how that is just an inspirational read about how you do software and the things that you're thinking about and it's not even that you just took NASA's rules and followed them directly. It's like that was just an ingredient into what you came up with. To be like, how can we make the highest quality system possible? Um, and, and so few organizations seem to be aiming that high. It's just really inspiring to see you aim that high and then actually deliver it. And it's not just talk and promises. It's like you implemented it, companies are using it, uh, and you've had it vetted by literally the best in the business and came back with a glowing report. Um, really, really well done.
Speaker A: Thank you. Thank you so much. I think for me, it was just, I was thinking it'll be easier to aim m high than to end low, because if you aim low, you're going to go even lower. So I was like, let's aim for like 10,000x and we just hit 1,000x, you know?
Speaker B: Um, uh, you'd be really, really happy with what you achieved. Right? Um, great. Well, uh, anything else we should chat about before we wrap up?
Speaker A: Yeah, people should donate to the Zig Software Foundation. It's, uh, replacing C. It's a great tool chain to learn systems coding. It's easier than Node js, um, to write fast, safe, know software. It's fun. You can do graphics stuff. What do you think? What else should we plug, download and run and rock?
Speaker B: Uh, be part of the rewrite Beer Beetle Jepsen. Um, yeah. And also, uh, you've made a very nice donation to the Zig Software foundation, which was really cool to see too. So you're not just, uh, taking the money and using it to hire interns. You're also directly giving back to what made it all possible.
Speaker A: Yeah, and that's Zig. Like, Zig is a huge part of Tiger Beetle's success. Not only the language we couldn't have written. Tiger Beetle as it is, is with the same fidelity in any other language. Uh, even matclad, he said he's almost pretty sure we couldn't have done it as nicely with the same integrity and elegance in Rust. Um, because we do intrusive memory and we need to program the system as a whole, which includes the kernel, which the borrow checker can't. You have to think of these things as, uh, systems, and you can't let your thinking stop at the language boundary. It's bigger than that. So we just couldn't have made Tiger Beetle in C or Rust or JavaScript. You know, zig was really what we wanted to express with tigerstyle. This was me at the Beginning, I loved C. Um, obviously there's huge problems. I really liked Rust. I was very interested and I looked into Rust again for Tiger VL and decided, no, the power to weight ratio of grammar and what we really need for safety. This is one aspect. Um, and then there was Zig. And I'd been following it since 2018. And, um, I realized, oh, uh, this is exactly what I'm trying to express. Andrew's feeling the same. And Barbara Liskov has a quote. She said, if you want to teach programmers new ideas, you have to give them new languages to think them in. And so, like the ideas of tigerstyle, you know, we needed Zig to think them in and Zig came along and so I'm so grateful. Um, but, yeah, I think that's a good way, good place to end is that, um. Yeah, it's. We, we. Zig really got things started and, um, you know, so C. We did that donation back to Zig. Tiger beetle did 256,000. C matched us and we did this over, um, over two years. So it's half a million dollars. And, um, we paid forward. And, um, I'm. I met Derek, CEO of Synadia through Zigzag. Uh, because he was also like you. He saw the swell. He saw, oh, this is great. Same as, like back in the day, rhino on the JVM. And, you know, it's inevitable. Server side JavaScript, you know, it's inevitable that someone's going to fix the C tool chain and the C language. And so you see Zig, you're like, ah, ah, this is inevitable. Paddle to the swell. And there's Derek Kalizon of Synadia. But not many people know. I think he was the person who was instrumental as a CTO at VMware, um, who hired Antirez back in the day. You know, remember, Antirez was hired by VMware and then he could work on Redis as open source. And like, behind that, you know, there was this. There was Derek and, you know, now he's doing nats and Cinea. So I met him through Zig and said to him, like, we want to do a donation. We'd already been doing it every year. We just hadn't told people. And we thought, well, maybe we should tell people that we're donating because some people need to know how much is donated before they start learning a language that you can learn in a weekend, uh, which is Zig. You know, don't even ask people, can I get a, a job, please? Just. It's a Weekend, just go and learn it like, and, and, but then, anyway, so I, I did think, well, we should maybe write about this because it will benefit Zig. And uh, so I reached out to an, to um, to Derek. And they matched us. And um, and we want to do more, you know, so. But uh, yeah, so that's, that's all credit to Andrew because you, you also want a BDFL actually. You know, you want someone with conceptual integrity. You don't want committees to ruin it language. I think that's the biggest risk for any language is that a committee will ruin it. Um, ECMAScript 5000 or whatever. And I can joke like this because it's a committee. There's no single person whose test I'm criticizing. But Andrew, uh, has always made such great design decisions and we learn from him. Whatever Andrew says, if he says no unused variables, that decision has caught bugs in Tiger Beetle. That would have been correctness bugs. So programmers think twice. If you criticize Andrew Kelly about what he says for no private fields unused variables, he's right. And just take a minute, listen to him. He's got a lot of experience and these things help us too. And we're fully on board. Uh, I love when he breaks the standard lib API because the brain is broken and he's resetting it because you want these APIs to be designed for the next 30 years. Um, and for us to update these changes. It's a small PR. We do many PRs a week. Uh, it's really easy to beat with Zig. So yeah, lots of miscellaneous thoughts. Uh, Andrew. Ah, Richard. What are your thoughts there as uh, we're coming to land?
Speaker B: I certainly agree about uh, Andrew's design taste. We've been very happy with Zig in general and also the breaking changes have not been particularly painful for us so far. Uh, and also I see you also
Speaker A: on Hacker News commenting about that when people always ask every time what about? And then you speak for rock and yeah, not mental for you.
Speaker B: Every time I talk to Andrew about our Zig experience, he's always like, just be unvarnished about it. You know, whatever's good, whatever's bad, just don't hold back. Uh, and fortunately, I mean I have a blog post in the works about our experience that I'm holding off on until we actually, actually cross the threshold of uh, getting everything working so people can come try it. Um, but uh, yeah, I definitely, overall it's been very positive and uh, full credit to Andrew for creating the language and uh, doing a good job. Running it.
Speaker A: Full credit. Let's take him out for a gelato next time in Milan.
Speaker B: Next time we're in person. All right. Well, yeah, thanks so much for taking the time to talk to me. This has been really fun. And, uh, I'm excited for the future of Tiger Beetle, continuing to be, uh, a shining beacon of quality in the world of software. So thanks so much.
Speaker A: No, thanks to you, Richard, too. And, um, yeah, thanks for having me. Such a pleasure.
Speaker B: That's it for this episode. I hope you liked it. I put links to some of the things Yaron and I talked about in the description. By the way, if you've been enjoying these episodes, I'd really appreciate it if you shared them around. I don't advertise, so word of mouth is the main way new people find out about software unscripted. And if you'd like to become a supporter of the show on Patreon and get ad free audio and video recordings, check out patreon.com softwareunsk scripted until next time.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.