
Adventures in Machine Learning · 2024-11-28 · 56 min
Key moments - from our scoring
Substance score
61 / 100
Five dimensions, 20 points each
Barzan Mozafari brings a rare academic-to-founder perspective to solving the exponential growth problem in cloud data warehousing: as data volumes outpace Moore's Law, linear optimizations (indexing, compression, parallelism) become insufficient. Rather than targeting a single pain point, Mozafari approached the problem philosophically - observing that Databricks, Snowflake, and similar platforms lowered adoption barriers so effectively that cost spiraled out of control. Kibo's solution trains reinforcement learning agents on performance telemetry and metadata (never touching customer query text or data) to automatically adjust resource allocation, query routing, and infrastructure decisions in real time. The platform uses a unique pricing model where Kibo captures a percentage of savings, aligning incentives with customer outcomes. Unlike Databricks' Predictive IO or cloud-native solutions, Kibo remains agnostic to the underlying stack, leveraging native platform primitives rather than replacing infrastructure. The conversation covers why optimization isn't a shrinking problem despite serverless trends - DBT model proliferation, data quality monitoring, and workload intelligence create new optimization surfaces. Mozafari also addresses the organizational dynamics: once optimization unlocks capacity, teams typically increase query volume rather than reducing spend, shifting the challenge from tuning knobs to teaching users to interact with optimized systems effectively.
Kibo hashes query text and trains reinforcement learning agents only on performance telemetry and metadata patterns. The agents learn which optimizations (resource allocation changes, query routing decisions) correlate with cost savings and performance, then autonomously pull those levers in real time without needing to understand the semantic content of queries.
Kibo is platform-agnostic and doesn't require data or infrastructure migration; it leverages existing Snowflake functionality rather than replacing it. Kibo also provides additional services like smart query routing and data quality alerts beyond just optimization recommendations, and uses a performance-based pricing model tied to actual savings.
No - as platforms become more serverless and hide tuning knobs, they're actually automating optimization decisions rather than eliminating them. The optimization surface expands to DBT models, multi-tool pipelines, data quality monitoring, and workload intelligence, creating new cost-reduction opportunities beyond simple resource sizing.
Optimization typically doesn't reduce spend; instead, it enables organizations to run more queries, integrate additional data sources, and execute more complex pipelines. The freed capacity gets reinvested in expanded analytics rather than returned as cost savings.
Mozafari frames Kibo as an unpaid customer success department - by helping Snowflake customers accomplish more work with their existing spend, it increases customer lifetime value and stickiness, ultimately driving more overall usage and revenue for the platform.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains solid technical insights about cost optimization, data learning, and the intersection of academia and startups, but much of the middle section devolves into anecdotal discussion about research culture and personal preferences rather than dense, actionable advice. The core insights (shrinking optimization knobs, reinforcement learning for cost reduction, query rewriting via LLMs) are present but padded with conversational meandering.
data volumes were growing growing at the fastest rate than More's law, right, So it was pretty scary because like if you're a computer scientist or you know math, you know that when you have two exponential curves, once you fall back fall behind, you're never going to catch up
we train AI models from how uses and applications interact with the data and the cloud, and then we start our agents that you know, for those of you from our reinforcement learning, which is essentially a major step in lll MS, very similar concept
While the framing of data cost optimization via reinforcement learning and the hash-based privacy approach show some novelty, much of the discussion recycles standard startup narratives (failing fast, research mindset, unlocking business value). The query rewriting via LLMs work is genuinely interesting but presented as an existing paper ('the general right') rather than novel thinking developed in this conversation.
we train AI models from how uses and applications interact with the data and the cloud, and then we start our agents
we don't have to make assumptions about what is that workload? Like is this an ETL? Is this a BI? Is it reporting? Is an ad hoc? Is a data science? Is a machine learning?
Barzan is a legitimate academic (UCLA, MIT, University of Michigan professor) who founded a real company (Kibo) with deployed products and technical depth. However, he is primarily speaking as a thought-leader and founder rather than demonstrating hands-on operational scale (e.g., managing large teams, surviving multiple business cycles, or executing at enterprise scale with measurable P&L accountability).
He studied computer science at both UCLA and MIT and then moved to the University of Michigan as a professor of Professor of Computer Science. He still teaches to this day, but recently founded a startup called Kibo, which is a fully automated cloud optimizer
I've been teaching databases, building databases, and selling databases. Right, so, like I know databases
The episode lacks concrete metrics, customer names, revenue figures, or specific case studies. Barzan mentions 'a major game baby company' and vague references to savings ('eighty percent or twenty percent') but never provides named customers, actual cost reductions, timeline specifics, or quantified outcomes. Claims about optimization effectiveness are asserted rather than evidenced.
Whether they save eighty percent or twenty percent. It just varies from customer to one customer to another, but it does actually generalize pretty well
you're a major game baby company or game development company, software software games
The hosts ask reasonable follow-up questions (e.g., about differentiation vs. Predictive I/O, handling long-tail queries) and show genuine technical curiosity. However, they rarely push back on vague claims, don't demand specifics when Barzan speaks in abstractions, and allow several answers to meander without course correction. The conversation is conversational but not sharp - more collegial than investigative.
Do you find that it's a generalizable solution to get say eighty percent of the way there
I have a really saucy question and go pronouncing around a slightly different topic. So Data Breaks has been working on this thing. Called predictive io
Computed from the transcript - who did the talking, and the words that came up most.
In today’s episode, Michael and Ben are joined by industry expert Barzan Mozafari, the CEO and co-founder at Keebo. He delves deep into the evolving landscape of data learning and cloud optimization. They explore how understanding data distribution can lead to early detection of anomalies and how optimizing data workflows can result in significant cost savings and unintended business growth. Barzan sheds light on leveraging existing cloud technologies and the role of automated tools in enhancing system interactions, while Ben talks about the intricacies of platform migration and tech debt. They dig into the challenges and strategies for optimizing complex data pipelines, the economic pressures faced by data teams, and insights into innovation stemming from academic research. The conversation also covers the importance of maintaining customer trust without compromising data security and the iterative nature of both academic and industrial approaches to problem-solving.
Transcribed and scored by The B2B Podcast Index.
Welcome back to another episode of Adventures in Machine Learning. I'm one of your hosts, Michael Burke, and I do data engineering and machine learning and other stuff at Data Bricks, and I'm drum by my wonderful co host Ben Wilson. I investigate cerper column like issues at data Ricks. Today we are speaking with Barzan.
He studied computer science at both UCLA and MIT and then moved to the University of Michigan as a professor of Professor of Computer Science. He still teaches to this day, but recently founded a startup called Kibo, which is a fully automated cloud optimizer, and they're most famous for their Snowflake integrations. So Barzon, as Data. Bricks employees myself and Ben, we understand how powerful it is to automate back end infrastructure, cluster provisioning, that type of thing.
But I'm curious as an academic, how did you enter this world? Were you using. Snowflake and head pain points or well, what was the origin story? That's a great question.
So I think it's it's probably easier to just start from like the word ebo. So it actually our first ideas where we're all about how we're going to speed up quits right, So keyboard in Japanese means hope. So the idea was like when you've tried everything and all other focus lost, like what else can you do? So it actually is an interesting uh intersection with the data bricks founders.
We are actually with some of data bricks as founders. You're working on approximate quity engine. The idea was because of Moore's law. So for those of you enough familiar, More's law is predicting how fast hardware is price to dropping or hardware speed is improving.
And then we were seeing the data volumes were growing growing at the fastest rate than More's law, right, So it was pretty scary because like if you're a computer scientist or you know math, you know that when you have two exponential curves, once you fall back fall behind, you're never going to catch up. So if the rate of data growth has already surpassed Moore's law, it means if you're happy with your performance database performance, you're going to be sad next year. And if you're sad this year, next year you're going to be depressed.
So the idea was like, okay, you know what, what is it that we're going to do to close that gap. So there's a lot that's been done in the computer science community and in the industry database industry, like whether it's indexing, data compression, you know, paralelism, all of that stuff, and that's all great, and you have to do all those things, but the idea is that all those optimizations are actually your linear speed up. So if you compress your data by ten x, you're only getting ten speed up.
None of these linear speed ups is going to basically help you eventually catch up or get ahead of that exponential curve. So that's where we start looking to statistical solutions to this problem, and then very quickly we realize actually the problem is not just about More's law or speed you know, you the bikes of data bricks or snowflake and other players in the state in the space of that amazing job of lowering the adoption barrier to sort of analyzing your data and getting insights.
But what's happened is that because it's so much easier to you know, get you know, get up and running and start analyzing data. Now there's a lot more users and applications that's happening into data and happening into a lot more data, and they're basically combining a lot more data sources, so now the cost of this infrastructure is going through them, and it's just wasn't humanly possible for anyone or eave to this day, Like it's not possible for you to look at you know, squint your eyes and stare at like in a one million quazy day and say, you know what, I think, here's how I'm going to reduce the overall cost.
Right, So that's where the story originated. We saw a real problem and we're academics and we thought about like how we're going to create a solution. And I was always an outlier, even academic, to be honest with you, because a lot of academics are just excited about coming up with theoretical solutions that complicated and they can publish it. But for me, it was less satisfying.
It was more about how can we create a solution that also gets a widespread adoption. So that's how Keyboard started. We started actually creating this data learning platform where we train AI models from how uses and applications interact with the data and the cloud, and then we start our agents that you know, for those of you from our reinforcement learning, which is essentially a major step in lll MS, very similar concept. We learn from how those and actions happen, and we start actually pulling different levers and real time and start optimizing it.
And you know, we came up with this pricing model off which I'll talk about it and the presentation. You guys are interested. But that's how the whole story started. Like we said, you know what, whatever money we save the customer, we take a small percentage of that so that the incentives online.
So that's how the whole story started. Well, that's that's a super interesting sort of paradigm shift from founders we usually talk to because it's usually born out of a frustration in a professional space. They're like, ah, this is too slow, this is too expensive. And instead you guys went from a philosophical and like academically based law approach.
I was just wondering if you have seen other startups be founded from those sets of principles or if it's usually more of I hate this thing, I'm gonna go fix it myself. It's it's it's it's a little bit of a you know, it's a little bit of both. To be honest with you. I think what happens is you know, it happens both ways sometimes, Like you know, to your point, someone been working in the travel industry for twenty years, Like this thing is way too complicated.
I'm going to just solve this industry, right, So, like I call them like founders on a mission where like, hey, you know, I just want to start a company. Here's the space. I understand. Let me work on some cool ideas, right.
I think in our you know, our studio was like we're talking to a lot of customers, like I, you know, I wish I could tell you a fancy studio of one morning I woke up I had this epiphany. But the truth is actually a lot of interesting solutions come the other way around. Like you basically are looking at the really important album in tried to figure out like what is it what you know, what's it gonna take to bring this solution to the market. Like you know, people tell you, oh, I have these ten problems, and then you can't just go and solve it and then hope that when you come back they're gonna pay for it, right, So you're gonna figure out what is it that drives them?
Like what are the characteristicscept that problem or the solution that will be acceptable to them? So short answer is no, Actually, but data breaks has a very similar story, right, So your founder's right, which I personally know, like they were seeing that map produce was really slow and it didn't make any sense. They can all this really cool idea of hey, what if we kept the data that we running erative competition in memory? And you know, there you go.
That's how spark was born and got rapid adoption and people went from there. Right, Okay, cool, that's a very interesting version story. Now how much of the internals can you disclose? H you know?
And now all of it? Well, I mean, we have patents in this space, we actually publishing coplications in this space. I can you know, I won't be able to get into any grady details of how we you know, train those models and whatnot that I can tell me, like you know, the high level workflow, you know how the whole system works and to the design principles and whatnot. And we have you know, a dozen of different algorithms under me.
Even if I want it, I have enough time to get in too details of every single one. Yeah, Having personally seen the source code for a que and spark it would take several weeks, I think. Exactly. So do you find that it's a generalizable solution to get say eighty percent of the way there based on the types of operations that different customers views are doing.
Do you see, okay, eighty percent of people who are adhering to you know, utilizing CTEs when querying data that structure the data and the lazily evaluated instruction set that's submitted, that they can say, okay, we can optimize that really well and it works pretty darn good. What do you do with the long tail of like somebody writing something almost you look at it and you're like, are you intentionally trying to break this? And what do the optimizers do with that? I think that's a that's a good, really good question.
Actually, you know, like I all I've done in my entire career pretty much is like I've been teaching databases, building databases, and selling databases. Right, so, like I know databases, but like you know, when you're saying someone's writing a really bad ct like there are quits, we see it, Like I'm looking at that quid and like I've spent all my career looking you know, writing seql quits. I can't optimize this myself, right, So you know that's our inside joke is like, you know, we want to be that infinitely competent, infinitely patient DBA.
But the short answer to your question is yes, actually, and the interesting part is we don't even see the customers queries. And that was a very intentional decision we made from early on, is that we wanted to a lot of people think that the hardest part about AI is that technology would used to be like a decade ago, but now I think we are as a field at the place where the technology is not the barrier. In many cases, sometimes it still is, but in many cases it's not. It's the adoption barriers that are basically stopping us.
Right. People worried about paid implementation, the autoi, privacy slash security, maintenance, tuning, you know, hallucination, all of that stuff. So one of those decisions that we made intentionally than was that because you know, I was involved in on the startup before Tebow, and I was seeing how difficult it is to convince. I mean, think about it like a cloud data ware housed is where you're keeping the most precious digital asset up an enterprise.
Now you're a startup, you're going in and say, hey, have this really cool solution. I'm going to slash your bill by fifty percent, which is a lot of money. And I'm sure you guys are aware of, Like you know for Croudata warehousing, it's a very expensive solution. But you know they're not going to trust you with the data.
So one of the decisions we made was that it has to be a no brainer from a security perspective. And what it meant was that our models can only learn and train on performance telemachine metadata. So not only do we not store any customer that we don't even see it, including the quit text, we hash the quit text. And the beauty of machine learning is that it can actually we don't have to make assumptions about what is that workload?
Like is this an ETL? Is this a BI? Is it reporting? Is an ad hoc?
Is a data science? Is a machine learning? It's it's just a bunch of numbers. Machine learning looks at this and says, hey, whenever, whenever I see this kind of pattern, I see this kind of behavior, I see this kind of cost, and like any you know, human, clever human, the agent's letter, they pull a lever right and if it basically managed to save the customer money without causing a slow down, the agent gets rewarded and learns from that.
And whatever it does something that doesn't lead to cost saving, it gets penalized and learns from that. So the answer to your question is surprisingly yes, Actually we don't know. We have not seen to a single customer for which we've not been able to save some money. But what's that percentage?
It depends on a bunch of problems, how underprovision they are, how optimize their workloud is in the first place, how open they are to you know, blooding. The models get more aggressives, some of them ask the agent with a slider until the agent where it needs to be conservative or aggressive, Whether they save eighty percent or twenty percent. It just varies from customer to one customer to another, but it does actually generalize pretty well. Cool I have a really saucy question and go pronouncing around a slightly different topic.
So Data Breaks has been working on this thing. Called predictive io, and it seems similar to what you guys do. And all these just cloud things have a bunch of data and a bunch of resources to build something similar. How do you guys differentiate and how do you guys avoid becoming super surpassed by a cloud specific solution like predictive iiom.
No, that's a very good question. So look like we basically what we do like we like super Laser focused on just being a data learning platform. We're not trying to replace Snowflake. We don't go to a Snowflake customers say you know what, you should go to database, and we don't go to your customers and say you need to migrate to Snowflake if you want X, y Z smart.
We're telling people as whatever exactly be on this call. We basically tell people whatever data stack that you've already invested in, that's great, keep that that what's keyboard is orders of magnitudes faster and significant cheaper than that thing that you're already using without keyble. So one of our other design principles has been like we should not require any data migration, any infrastructure migration. So whenever the cloud provider or the cloud data warehouse has certain functionality, we actually leverage that.
So Snowflake, for example, also had a bunch of really clever internal mechanisms. What we do is that we never reinvented with because of predictive by oh, people will actually try to leverage that to some extent, and it's not just the optimization we actually provide finops. We have a new technology on the same point platform called smart quity valery, right, so you could potentially use your own predictive I ought to figure out where to route those quadities, right. So we use that to decouple the application the customers application logic from the application performance, so the you know, the user, the customer can just focus on the use case without worrying about costs, without worrying about performance, and just decouple those decisions.
So you just send those quoties to the smart quid out and they will decide, hey, maybe you know I need a small, the square needs a large, just one needs a medium and so on. So short answer to your question is, we don't invent the wheel. We're not trying to replace the underneath, the technology underneath. We take advantage of whatever primitive and functionality that's in there, whether it's for better insight, better recommendations, or better actions.
But it seems like a sort of shrinking pie at for instance, data Bricks has invested heavily and serverless and they don't really expose knobs, so there's less that can be tuned. And so what are the sort of sticking points that you anticipate There will be optimizations for the next five and ten years. But that's a very good point. Like, look, you know, when you're thinking about people are trying to So if it's sort of just like go back and look at for example, Biitquity, another player in this space.
Right, you can just you know, you spot instances, or you can go completely like you know, here's a flat trade, or you can say I'm just gonna send you the quod you figure out what you're gonna haunt it and whatnot. There's always a cost performance trade off, right you can you know when you're going with several as someone else is making that decision, you're hiding the knobs and you're automating those knobs. Right. So I don't think I think the idea of a shrinking pie for knobs is a valid question.
But I don't think it's just data learning is not just about knobs, because at the end of the day, I'm sitting next to the customer's most valuable digital asset. I am understanding the data distribution because that's just the first app that doesn't see the data. We have additional apps that wants the customers in the platform. They actually see the data, they see the quait text, they see all of that stuff.
If I'm sitting next to your cloud data warehouse, I actually understand your data distribution more intimately than any single individual organization. So when something's out of the ordinary when it comes to data, I'm actually the one like me, meaning the agent right, is the one that actually finds out first that, heyst this column never had null value. Suddenly you have a lot of non values. You're working with a pretty large customer, and tend out that like one of the really important columns had become null for several months and no one will be noticed.
So we understand those drastic changes in the data distribution. We can actually see certain KPIs. So like you know, warehouse organization, which is what you're referring to, is just one use case. But even that use case, I don't think it's going to go away because you're actually creating something several less.
Now people are creating more that basically that just like moves the bar a little bit. Like people are not, you know, worried about what's the sever size I'm going to be using, but they're going to tisch it with five other data tools and then build a more complex pipeline. And now the question is how do optimize that pipeline. One of the major drivers of costs in the cloud these days is DBT models.
Right Like, you know, you created like one hundred and eighty five hundred the DVD models. You know, good luck optimizing that. You could be several less, Like you can remove all the knobs you want. At the end of the day, the question is the customer has to pay x dollars.
What can you do to reduce that cost? Sometimes the solution is changing the knobs. Sometimes the solution is changing equity. Sometimes the solution has change the way that you're actually quitting your data.
But there's a lot more like you know, we're as optimization, workload intelligence, SmartWare, routing, data quality alerts. There's a lot of different ways you can expose that. You can leverage the understanding that you have of the customers usage behavior and expose it to them at different parts and that data stime. Heard.
Yeah, I couldn't agree more with your perception of that as somebody who many many years ago, back when I was doing data science work and like data engineering and work several times at companies I worked with or customers I was working at when I was in the field that data breaks. You always hit that point where you've migrated to a new platform, people start using it, and sort of bad processes have propagated to the new platform, and you open it up and like, all right, it's ga.
Everybody can use it. And the issue you first query and you're like, man, this is slow. It's faster than it was on our old platform, but it's still slow. Yeah, And then you're like, all right, we need to take an entire quarter or two quarters and redo the data model properly and get rid of all that old tech debt.
And every time that I've been a part of a team that's done that, you open up a whole different problem. Right after you get all that fixed, you're like, Okay, the query that used to take four hours to run now executes in ten seconds because we actually put the data where it should be and optimize it, put in nexces. On stuff, everything exactly. But you still hit that there's a finite resource limit that's placed in any business, which is the CTO gives you a budget for like, there's so much money you can spend on this stuff.
And when you fix all those problems, people just start issuing more. Queries exactly, They're doing more and more exactly. True. So then you have to like, Okay, the queries aren't optimized, how do we tea And then like, what you're tackling is the thing that's every time I've tried to do it or been part of an organization that's tried to do it, it's the hardest thing to fix, which is how do you teach people how to interact with an optimized system properly?
And no matter how much effort you put into it, you're never going to be as good as an automated service that can do that. That's that's hundred percent. And that's actually one of the common questions that sometimes people ask us is like, aren't you afraid that, like, you know, the snowflocks of the world like feel like you know, you're reducing their revenue, And I'm like, no, actually, I'm just there. We're just their unpaid customer.
Success department because yeah, we're just letting you know, those customers get more work down with less money. So at the end of the day, people like to your point, when you optimize their workload, they end up actually doing more. You know, it's not that they go back to the CTO. Sometimes they do more often than not, you know, the bar just shift someone else, like now they're going to send more quid, they're going to stitch it with five out of data sources.
And that story seems to be you know, repeating everyone. Yeah, that's what we see with data breaks customers all the time. It's like they start off with with ETL, they get all their data in their warehouse, and then they didn't move on to BI and they had all the query suck because of tech that and they fix all that and then it unlocks the mL side. They'll hire a data science team.
Somebody knows what they're doing on that team, and then they'll they'll get some stuff into a you know, maybe staging and validate it and eventually to production. You look at the account usage over time, they're like, hang on a second, like, yeah, they've increased ten x as our customers, but they weren't doing any of this stuff before, and if they're a publicly traded company, you can kind of look at them like, jeez, for the last four years, like they've doubled in revenue.
Is that because of us? You know, you will you kind of want to take credit a little bit for that, like maybe that was two percent US. And sometimes I'll say that like, yeah, this unlocked our business insights and you can now compete against our competitors. No, that's that's that's so true.
Actually, you know, one of the I wouldn't say sadisting, but only at least the most interesting things that we're seeing is that you're just still a lot of data teams that basically are constantly like spinning their wheels, trying to reinvent the wheel. They're trying to sort of they're too too because like they haven't finding box right, and I understand that they're trying to sort of manually pass things up like hey, you know, I'm going to do X y Z, I'm going to reduce the cost.
Like just imagine how much value would be unlocked if if you actually shifted those resources into growing your business your point, right, Like you know, those all those smart engas, like for example, you're a you're a major game baby company or game development company, software software games, and and you know your data team is just spaying. They will starting to sort of figure out how to leverage snowflake more efficiently, or how to leverage how to optimize the data backs workload, whereas there you know, games game development company, they should be focused on like bringing the nuts and better version of that game and increase the top line instead of being so focused which is very common easy because of the economy obviously, But you're right, like I think when when you free up t you know, the resources from being consumed by all these you know, cost saving and things that are not the core business of that that that customer.
To your point, you go back and look at it and say, I'm glad that those engineers are actually focused on their growing their business started. And how do we pay you know a little less of this particular tool that was supposed to free our time up instead of tending us into you know, optimizers for this attitude. So another question back to sort of your origins from academia, what are some of the skills and concepts that have been essential and founding Keebo, Like, what are the things that you learned in academia that translate really well to being a startup founder, specifically in such a technical space.
That's a good question. I think I don't know how much of this would generalize every startup, but I can talk about like the kinds of startups that look like kebo I. You know, sometimes jokingly say the listen, the reason why he was being successful like we've been going out of pretty quickly, is not because you have really charging sales reps. The reality is like our product or you know, I shouldn't take her for it, but our products built a product that's out smarting of the other solution out right.
So the reason why we can do it is because it's just we were not just looking at hey. I usually give the example of key value stores, right, like there was a there was an era where every other week there will be a new key value store out there. At some point they run out of they run out of names for these companies, right, because it was very easy to build a new key value store. You would just and the nice thing about it is like there's one hundred plus different key value stores, so you never get stuck on anything you don't know how to implement X y Z.
That's what. There's ninety nine other products you can look at. But whatever, you're trying to create something for the first time, and it is truly innovative. Like now it's just a better user interface.
It's just a slightly more optimized version of what everyone else has been doing for the past twenty years. That requires research skills, right. So one of the nice things about academia is that you know, in the industry, right, like, if you want to pitch an idea to your boss or to the company, they think about risk. So oftentimes they try to serve out put all those apples in one basket.
They say, you know what, uh, that's hygd reward high risk, which is usually shorthand we're saying we're not gonna do it, right. But you hear this a lot at you know uh in in UH in the industry. But like in acadia, it's the opposite. You get rewarded for taking on hairy, big problems and considering solutions that no one else has developed.
Because even when you fail, you learn from it and you go do something. That's because that's what academy has made for right, like for for for people to go and freely innovate and push the boundaries and things of that sort. So I think research skills, which doesn't mean you need to have a PhD, but like the ability to take on an open ended problem, I think outside the box, come up with a solution that maybe no one else has has thought about and and kind of execute done. I think that's definitely one area.
And the other idea is this, like this whole thing about failing fast. We keep talking about failing fast, but that's pretty much what happens in academia, right like, so the still cycle in the industry unit if you're thinking about for example, B two B software, right like, you have to come up with a you know, usually an MVP takes at least two quarters. Right after that you're working with data customer that's not a quarter. And then then we talk about a really fast like product to market kind of cycle.
And then you know, you have to chain the sales team and you start selling some like every customers getting traction and whatnot. Nagadimia, you're write, you live life writing one paper at the time. So if you have an idea, you submit it to a conference, you know, as soon as like you have some you know, compelling results. You write up a paper.
Your code could be complete crap, but you just have a proof of concept. You write up a paper, you run a bunch of experiments to see if it works or not. You don't have to go higher sales people. You don't have to go, you know, spend millions of dollars on marketing.
You just basically go out there and and that paper and then and then get peer reviewed. And if it's a bad idea, you'll find out. And like most conferences and computer science, you hear back within two months three months laters, right, So you have to fail fast, this idea of being scrappy, you know, and making sure that you know you see somebodys elf before you invest too much into it. I think those two things from our company DNA really did help help us out a lot.
Keebo must have been something in that lab that you are in, because that exact approach has actually carried over into data ricks R and D. It is pretty common. This is just but I've heard from other people that have come from fank companies into data bricks and their remarks are like, I can't believe we're allowed to do a Spike and like, yeah, we have to do design talks and stuff, but we get time to do a prototype, and sometimes somebody will give us like, hey, go see if you can figure this out, Like take these like you five people from all these different teams, just just take six weeks and play jazz, figure out what you can come up with.
And sometimes it's a failure, like an abject failure. We'll even release it the private preview, get like twenty customers trying it out, and the response is like, we don't know about this, And then four months later we have version two point zero that's in public preview and people are like, this is amazing. Where was this all my life? But yeah, that that iterative process of just failing, like failing really hard.
Sometimes it is critical to like how we release products the way that we do, but a lot of companies in the tech space just don't do it. They don't do it. No, you're spotted. And I think it's just also a little bit about like getting people with a research mindset because like you know, like as someone I've been writing code from an early age, right, like I was a program before I was a researcher.
But like if I had to confess. Like researchers usually doing write the best quality code, right, Sometimes we write crappy code because we're just trying to prove that concept and walking down and that drives solid engineers and experience, you know, techniques sometimes crazy you know, how can you or something like this? Right? But like I think if you can create an environment where people like researchers can being the research skills, solid architects can bring their expertise and like ten help like transition once those ideas are devers or tried out, help transition to product that skills right, like something that's robust and production quality.
I think a lot of amazing things happen, like researchers on their own another or create something that actually you know works at scale. But like if you can pair them with with engineering teams that you know are are solid and can take those ideas and transition and like obviously that means both camps have to get out of their comfort zone a little more, right. But I think when you when you have an environment that's conduc it to that kind of collaboration, just amazing things happen to your point, but.
Some comfort zone transitions, but it's exciting like everybody gets so in used about it on both sides, because you get the researchers. A lot of people come from that we've hired, they have like ten plus years POSTCRAD, they've been doing research at Berkeley or Stanford, MIT or something, and they come in they're like, whoah, this code's complex, and engineers are like, oh, what are you working on? I want to I want to see it, And there's no, there's not Like I think there's a brief moment of panic on both sides, but then everybody's like, hey, let's work together and let's team up and let's make this awesome.
And you just see everybody grow together because you're expanding the mind of engineers to see like what is theoretically possible and it unlocks a lot more creativity on their side. And then the R and D researchers eventually they're writing like production grade code within a year or so. So you're like, yeah, it's a win win all around spot exactly. So it's it's it's exactly like how this.
Kind of which process do you both like more? Do you like research spikes or more engineering focused work. I think we do both, but I you know, I think it depends what you're trying to do, right, I think no. Personally, Like, which do you enjoy more?
Oh, I definitely enjoy research spikes. I think it's just like, you know, like I said, like you can never pay me enough to go and create another key value store. Like I'm just the kind of person like life is too short, Like if I want to do something, I want to be the first person doing it right. So research spikes usually have that kind of flavor where like, hey, you know, this is an idea.
I might come back and say, guy, that's that's not promising. You know that's not going to work. But you know, when you do come back and you come up with a you know, new solution no one else has thought about and it actually works, you get to you know, big you know spike of dopamine or and it's that makes it all worth it, at least personally for me. I enjoy three distinct points in that development process.
The first one is I love seeing all of my dumb ideas fail in the beginning because it just it shortens the path to getting something that might work. And it's also kind of fun. I like seeing, like. I think I was telling you Michael.
The other day, I was doing something late at night and getting some CI set up and a package that I'm working on, and I wrote some really terrible code because it was like twelve thirty in the morning, pushed it to get hub actions, and then I crashed the runners, like killed them all basically effectively, like a stack overflow, and I just looked at it. I was like, I'm going to bed, but I kind of chuckled to myself when it's been a while since I've broken something like that. And then the next morning I look at what I actually submitted, I'm like, yeah, don't code when you're that tired, dude, and fixed it and then it passed.
I'm like, all right, sweet. But I also love the transition from the proof of concept works and buy in has been signed off, like it's been effectively peer reviewed amongst peers of the company, and then banging out that first production grade version of it. I love that experience. So like, Okay, I know how bad my code was.
How do I make this actually usable and extensible and maintainable, and how do I just kill all of this complexity that I had to build in the script that I wrote. That's very enjoyable. And then finally the release not not the response, I don't really care about that. I actually look for people like who use it that then tell me why it's broken, because I love fixing the bugs on the like the first few iterations.
I love that experience. It's not like I know this code because I wrote this crap and I love I'm like, yeah, totally fix that's it. That's my dopamine hit. Well, I love it.
Like. I also like how you kind of like the three stages, like you're seeing the true value of each of those three stages, like and liking it for what it you know what it is, Like, Hey, I would not get from H too C if they didn't have point B in the middle. No, I I that's that's that makes a lot of sense. Yeah.
I think my response is I really like the research aspect, but it's sort of a product of my job because I don't have the opportunity to build really complex extensible frameworks that have like cool designs, Like I'm writing a thousand lines of code maybe two thousand for like a typical project, and the really fun thing is trying the art of the possible and seeing like can we make this work, Like what creative ideas for attacking a problem in it from a different direction? Can I employ to make it it's successful?
So yeah, it's interesting, but they both have their prison cons. It's interesting Barzon that your your angle is research because I feel like computer science is very fundamentally implementation optimization focus. Would you agree or do you think there's a lot. There's a lot of what or would do you think there's a.
Lot of sort of innovation and like groundbreaking like far out their ideas. It actually depends on what discipline you're looking at, right, So, like I might get into trouble for saying this, but for example, if you just look at databases as a field, like which is my own field? So I feel like I'm allowed to say things like this. I think the field has kind of plateaued.
You go to as you know, you go to the event, you go to like these places where they talk about innovation, and you're looking at this and saying, like that's really cool that like that ship now has X. Actually Oracle had that like thirty years ago, right like, Hey, I'm so glad that you guys do auto indexing here, but that happened here, or like you have this storage optimized, think here to use this compression. Well you know what, actually Verdicta had that like twenty years ago.
So it's the field is popular. It doesn't mean there's no innovation, but like if you just try to build another database, a lot of it is being tried. And I'm not saying there will never be enough innovation. I'm just saying the number of new ideas that are like radically new and actually are effective, it's we're running out of those ideas.
Like the field has matured, which is a good thing, right, It means we can go and build the next set of you know, AI enabled a AI enabling applications on top of what people are like now we're We wrote a paper a few years ago goe to my former PhD students about database learning. So the idea was like, okay, now let's see you do have a database that's optimized. But every time, you know, if I keep asking you The example I give is about cars, right. If I let's say that you're like me and you don't know anything about cars, right, then if I keep you know, if I ask you a question about this particular model of Ferrari, you like you're going to go online and look it up and give me the answer.
If I keep asking questions about cars, you're going to keep like, you know, googling it. But after two three days, you're gonna pick up a few things. You're gonna learn. It's going to take you less and less time to come up with an answer to car related cars, right, because we're humans, like we learn.
The databases don't learn, you know, asides from like very basic things like hey, I cash this data, I cash that result before the data changed you. Every time you go mediquated it, this does a bunch of work, send you the results back. For the most part, that work is lost. Afterwards you go back and it starts like the databases don't learn.
So that the vision that we basically presented and we actually built a proof of concept on it, was like, how can we build a database? It actually learns over time, It becomes smarter every time that you quit it. You can think about it like if I ask you, hey, what's the average sales for this particular region her department, and then tomorrow asks another question that kind of overlap, like maybe said, hey, what's the total number of transactions in the region for the entire country.
The fact that I know something about that region should help me come up with an answer to the second question a little bit faster. Right, So, I think there is still innovation, but it's very build specific. Certain sub disciplines within computer science are world researched. People either have moved on or they need to move on.
They have people who still haven't moved on, and they still like, you know, cip iterating over similar ideas. Hey, actually I found this corner case where I can make the indexing like five percent more efficient. But there's a lot of interesting things, especially like the time we're living and with other lens, with machine learning, with you know, hardware acceleration that we can we can still actually come up with pretty cool ideas, like human mind doesn't not other cool ideas.
That's a nice thing. It's just like, you know, maybe you fix something, you go and create a new discipline. Curious, for both of your guys' opinion, what are the frontiers that you're excited about? Are the new piece tech algy.
That in the database and data querying space that you think are going to be game changing? Then wasn't just talk. I think for data querying, the ability to map to an entire data warehouse or entire system of rdbms like implementations that exist in an organization and for you to be able to talk to an agent and ask a very complex question and you get the accurate response from all of that without you having to build all of the interfaces to that, because today you can theoretically do that, right.
You can create a bunch of tools that all issue all of these different queries to all of these different platforms, or you can you know, have like basically fine tune the model on the metadata of your table in your databases. And I don't think that anybody's gonna pick that up to get to, you know, a high ninety percent accuracy response rate. Like we offer something called Genie right at Data Bricks, and that's Language Model interface to query Unity Catalog tables and in demos, it's incredible, like amazing.
I've played around with it, I'm doing integrations with it. I'm like, man, this is so cool the fact that I can, you know, put one hundred column table with a million rows and I can ask it just plain language questions and it figures it out. And I can do this with five different tables and it'll generate those queries for me, and it's pretty performant because of that optimized engine in the background. But then I point it to our internal tables or data that I had written to years ago in the Unity catalog during the demo days of that, and I usually the same query and it loses its mind.
And then I'm like looking at it, like why why does it work so well on these tables that I created, you know, last month, and that my old data. It's it's just not good. And then I just go into the UI and I'm like, oh, yeah, there's no metadata here, Like there's no comments anywhere explaining what this table is, what's in it, or the conditions for the ETL that is actually putting the data in. And then the column names are almost intentionally obfuscated because I was just doing shorthand nonsense and I have no parameter comments anywhere of like what this column contains, so it's making guesses it's inferring from what metadata it actually has, and I'm just like, Okay, there's got to be a better way to do this.
So I think that the golden goose out there is for a parson on this team to figure out how do I do that with the table? How do I generate the metadata? It is highly accurate, that is contextually relevant to this business in a way that you know interfacing with an agent will work properly. Yeah, just real quick before you jump in.
I have so much beef with Genie right now. The account teams that Data Bricks have sold the proof of concepts like five different customers and then the customers are like, oh great, so now you're gonna build me this agent. I'm on three of those projects right now, and we just have to like lower the expectations three orders of magnitude because it's just not there yet. It's a really cool technology and it will be there soon, but the demo is not what it is in reality.
Yeah, over to bar Zone. No. I think the explanation is one of those years I'm so really excited about. But like if I'm kind of zooming out, like it's very easy, Like if you ask me, what's like if I had the magic one, I could solve any problem.
I would obviously say world hunger and like cancel right like, but also very realistic about my own skill set, right, so I think the most important thing, like the way I'm looking at it as someone who's excited about innovation, but also like I want to make sure it's practical and gets adoption right. And part of adoption, like to me, has four legs and one of it that has to work right, not to just be on the demo to Ben's point, right, like it has to work otherwise you get the rest people really excited and to get really frustrated, which I think is a big one of the barriers to some extent with AI is like people if they if the level of excitement doesn't match that expectation, then they get burned out and I don't know when the next time that the CIO was going to sign off on something at the word eleven min it right.
So if I'm looking at it from that perspective, I think I think the key to success would be to focus on what the intersection of what can be optimate up automated and what should be automated. Sometimes people try to automate things that shouldn't be automated and or the things that should be automated but cannot be automated with today's technology, and because they're inaccurate, they're inefficient, unreliable, all those reasons right. So if I'm looking at that intersection of what can and should be automated, one thing that's actually working on that thing is very exciting is like with elms, we've seen massive success with actually quity rewriting.
Like as someone who's been like just dealing with quities for the past twenty years. It can actually rewrite quitties that we never thought possible. But it's not just like hey, chat GPT, can you please rewrite this quity for me into a more efficient form, because actually four out of five times, or I should say eight out of ten times it actually gets equated either doesn't even compel, or it actually compels, but it gives an incorrect answer, or it compels gives the correct answer, but that's actually slower than the one I started.
Like, we've created this framework around and like the papers out there for those who are interested in the audience is called the general right. We actually we've created this really cool cycle where we basically get that we were interacting with the LM actually come up with what we call human readable rewrite rules. So like when we ReLit it, we actually turn it once we valve it turned into a rule, and then when the quit comes and we use those rules as actually as hints to the l ELM.
So now we basically get pretty accurate, like ninety plus percent accurate with in the sense that we can actually whenever we rewrite the quity, we have pretty high confidence that's actually correct and it's actually more efficient than the original quity. And the nice thing is that this database of human I forgot what we call it in the paper. I think it's human understandable or human rewrite rules, something like this HR to L something like that. That database actually keeps growing.
So it's like more along this vision of creating the database that keeps getting smarter over time, like chat GBT. The more people are interacting with it, it's also getting smarter and smarter. So like creating a system that gets smarter over time the more we use it, I think is also super exciting. But I think we will get to a place like when that we will be able to explain a lot of interesting things like hey, why did my sales negotiate department of this particular Walmart store you know go down last month compared to the you know, other stores, comparable stores, right, And we will never be able to fully automate experimentation and causality, but at least we will be able to show them most likely causes to the domain expert, who will then have that domain expertise which should not be automated or cannot be automated at least today, to tell us, hey, you know what, these are the top series.
And I think it's because we have too many people, you know, out of office, or there was local event that this out of place that was not here. So I think that's that's the line that I'm really excited about just working on things that can and should be automated. Yeah, that example brought to mind an old example that I used to use when when teaching new data scientists to teams at past companies about the difference between correlation and costality and intelligence systems. And like, here's this model, and I had this data set that I would always use that was it was basically like year round temperature at a park in New York City, and then another column was like amount of ice cream sold.
And you build a very simple model, a regression model, and then use explainability tools and costality tools on that data. And of course it it's like, hey, I want to optimize sales and what does it come up with. It's like the thing that you need to change is just increase the temperature, and that's going to teach people like, hey, be careful of how you interpret things that come out of, you know, an algorithm exactly. I think that the thing with like the explosion of jen Ai and its popularity and it's democratization.
The only thing that I see as potentially disillusioning in that as these these capabilities become greater and greater over time, I'm like, hey, I can query all my data and I can ask whatever question I want, and I can bolt onto this tool that's going to do this causality analysis for me, and somebody's like, inevitably a system is going to be built that has those features that can do these sorts of things and it can query the right data. And then somebody's going to say, how do I make my sales go up?
And they're going to ask that to the system, and the system's going to go and it's not going to say increase the temperature of the planet Earth in January, but it'll could do something similar to that in their business and they might not know like, oh, maybe if I yeah, focus my efforts here. It turns out you're cannibalizing from another part of your business and you know, creating chaos or whatever. I think with incredibly intelligent and reliable systems, it could create trust issues with people with those systems.
Is that something that in academia people are thinking about. I think so, not as much and not as many as but there are some actually, Like you published a paper called dB Sherlock a few years ago. It wasn't using l lens, but the idea was like, how we can actually incorporate cause on models into a system that can show the most likely causes and then use it the cause on model to actually help use there so that we don't tell people what caused the reign was that you know, your wife took down Brella, but we say, hey, these are most correlated with each other, and then can use consolity models, so the system learners over time.
But I think there's some people we are looking into it, but not as many, to be honest with you, not as many as I you know, wish these days. So I've got a silly question for you. You've been in the space for a while and I've been doing research for a very long time, and we're likely exposed to the things that everybody thinks is pure magic nowadays about like, oh my gosh, chat GBT is the best thing ever. It's it's so smart.
Anybody who's been in AA like dealing with advanced computer science for decades is going to look at that be like, yeah, we had these like a while ago. They've been around a while. Maybe not transformers models, they're they're slightly more advanced, but they're growing off of the shoulders of giants that they came before. Were you doing like a table slap or knee slap with a bunch of other professors saying I called it?
I knew it was going to. Happen this year. Where my grandma knows the name of something that is involved with artificial intelligence. That's a really good question.
I think when you spend a lot of time in a space, you actually see like certain things that become trivial, like or look become certain to you, become clear to you, right, but like to outsiders because that's all you know, right, Like, if all you've done all your career is like this very narrow area, which is you know, sadly, the situation with a lot of us in academia is like we know everything about a very little narrow topic, right, so it becomes pretty clear, but to outside it looks like magic.
So yeah, I would say, like, you know, I mean I had students who work on like transformer models and whatnot, so like we were seeing the advances that are coming. But you know, I think what surprised all of us is how quickly the public kind of was impressed with it. Right, Like we go to a conference, we say, hey, we improve this accuracy that like you know, half a percent, and we clapp for each other we get right, but like eventually when it becomes good enough that everyone else's also gets excited about it because they're not there in the journey where like it were just growing a little by little, little by little, where it's harder to see it like they saw like hey there was like sci fi movies and now this is actually here, so.
Right, yeah, we even got to see that over the last you know, eight years or so at data bricks with even traditional IML, where you look and the first couple of months or probably the first year the mflow is out and you're looking at the statistics of like how many people are saving what types of models? You're like, oh yeah, we've got like one hundred users that saved sk learned models and deployed them and a bunch of people doing extra boost Like this is exciting. And then you look now and you're like, how many millions of these were were saved in the last week alone?
And it's become so commonplace. Every business has these things, and I just but you look an account might they might be hitting that API for logging that thing five hundred thousand times a week. It's like, wow, that's crazy. It's so commonplace.
But ten years ago that would have been like whoa, this is state of the art. And then people that have been doing that stuff for a long time, like when I came in and would talk to to like new accounts that we got on, they're like, we want to learn more about this this new thing called data science and we want to like understand it, like new thing. This has been around for a long time, like but way before I was born. They're like, what, no, we just heard about this thing that you can do.
I'm like, yeah, the paper for that was written like before one before computing, so that speed like what you talked about It surprised me a little bit. I didn't think it would hit psyitchgeist level of like everybody knows this thing and everybody's got an account on this thing. It's exciting, but it's also very surprising. No, exactly.
Cool. So I know we're coming up on time. I'll quickly summarize really interesting conversation. Some things that stood out to me.
Our research skills are very valuable for innovation, and in academia you can typically learn the fundamentals of research, at least one would hope. And then also fast failure is essential. Sort of at a macro level, a lot of organizations are turning off knobs, so there's less configuration and customization you can do. But despite that, there will always be additional layers of infrastructure to optimize.
People will start using those as discrete blocks in more complex systems. And then some future areas of innovation that we're excited about. Our agentic querying and then query rewriting, And if you guys are curious about the paper, it's called query rewriting via large language models. So barz on, if you want to learn more about you or your work, where should they go?
If you google my name? Or go to keyboard dot AI. That's that's our company. You can get live devels of what we do.
You're using you know, Snowflake or in your cloud to doaberhouse in any capacity, and you're interested in auto optimizing it, you know, diverting some of that manual effort or infrastructure bill to some other areas of your business. It's sound a keybod dot a I, or google my name and look at my academic homepage much papers, or reach out to me by elected cool. Thanks so much. All right, well until next time, it's been Michael Burke and my co host and also and have a good day everybody.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.