
Adventures in Machine Learning · 2025-04-04 · 56 min
Key moments - from our scoring
Substance score
53 / 100
Five dimensions, 20 points each
Barzan's entry into the startup world came from academic research on query optimization and the fundamental problem of data growth outpacing Moore's Law. Rather than starting from workplace frustration, his team built Kibo - a fully automated cloud optimizer that trains AI models on performance telemetry and metadata to reduce infrastructure costs. The company deliberately avoids accessing customer query text or data itself, addressing enterprise security concerns while still achieving meaningful savings. Kibo uses reinforcement learning agents that learn which configuration changes reduce costs without degrading query performance, with a pricing model tied to actual savings delivered. The platform is designed to integrate with existing data stacks like Snowflake without requiring migration, leveraging the cloud warehouse's own optimization mechanisms rather than reinventing them. Unlike competitors like Databricks' Predictive IO, Kibo focuses narrowly on cost optimization as a complementary tool rather than trying to replace the underlying warehouse. Their secret insight is that machine learning doesn't need to understand query semantics - it can optimize based purely on performance patterns and cost signals, making the solution generalizable across diverse workloads from ETL to ad hoc analytics.
Kibo trains machine learning agents on performance telemetry metadata only - never accessing actual queries or data. The agents learn patterns from numerical signals like execution time, resource usage, and costs, then pull optimization levers in real time, learning from whether each action saved money without degrading performance.
Kibo takes a small percentage of the actual costs saved for each customer, aligning the vendor's incentives with genuine savings delivery. This differs from fixed-fee pricing and ensures Kibo only makes money when customers benefit measurably.
Yes, Kibo has not found a single customer for which it couldn't save some money. The percentage varies based on how underprovisioned the infrastructure is, how optimized the workload already is, and customer risk tolerance, but the approach generalizes across ETL, BI, ad hoc, and machine learning workloads.
Kibo is narrowly focused on cost optimization as a complementary tool rather than replacing the underlying data warehouse. It leverages Snowflake's existing optimization mechanisms rather than reinventing them, requires no data migration, and works with customers' existing stacks without asking them to switch platforms.
Barzan observed that data growth rates had surpassed Moore's Law, meaning traditional linear optimizations (compression, indexing, parallelism) would never catch exponential cost problems. This led him to explore statistical and machine learning solutions that could provide exponential speedups rather than linear ones.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains genuinely useful technical insights - RL-based optimization without seeing query text, the exponential gap between data growth and Moore's Law, and an LLM query-rewriting validation loop - but roughly half the runtime is consumed by Ben Wilson's personal anecdotes (crashing GitHub Actions at midnight) and extended mutual validation between host and guest, significantly diluting the signal-to-noise ratio.
if the rate of data growth has already surpassed More's law, it means if you're happy with your performance database performance, you're going to be sad next year
our Models can only learn and train on performance telemachine metadata. So not only do we not store any customer that we don't even see it, including the quait text we hashually quit text
A few genuinely fresh framings appear - positioning Kibo as 'Snowflake's unpaid customer success department' and the 'databases don't learn' thesis underpinning their product vision - but most of the episode recycles familiar startup tropes (fail fast, academia vs. industry risk appetite) and the Moore's Law observation is well-worn in data circles.
we're just their unpaid customer success department
The databases don't learn, you know, asides from like very basic things like hey, I cash this data
Barzan is a credible practitioner-researcher hybrid - CS professor at Michigan with UCLA/MIT training, patents filed, papers published, and an operating startup with paying Snowflake customers - making him a legitimate subject-matter expert rather than a thought-leader type, though he is not a widely recognised industry name and the startup appears early-stage.
I've been teaching databases, building databases, and selling databases
we've created this framework around and like the papers out there for those who are interested in audience is called the general light
The episode offers concrete numbers and named artefacts - 90%+ accuracy on validated query rewrites, 20 - 80% cost savings ranges, 180 DBT models as a cost example, and references to specific papers and competitors (Vertica, Oracle, BigQuery) - but lacks hard customer case studies with real dollar figures or before/after timelines that would make the claims verifiable.
we basically get pretty accurate, like ninety plus percent accurate with in the sense that we can actually whenever we relte the quity, we have pretty high confidence that's actually correct
One of the major drivers of costs in the cloud these days is DBT models, right, Like, you know, you created like one hundred and eighty private on the DVD models
Michael Burke lands a few sharp questions - the 'shrinking pie of knobs' competitive challenge and the Databricks-employee-asking-about-Databricks tension are good moves - but Ben Wilson repeatedly hijacks the conversation with extended personal stories that go nowhere, and neither host pushes the guest for hard evidence behind the savings claims or customer-count specifics.
But it seems like a sort of shrinking pie. At For instance, data Bricks has invested heavily and serverless, and they don't really expose knobs, so there's less that can be tuned
I have so much beef with Genie right now. The account teams that Data Bricks have sold the proof of concepts like five different customers
Computed from the transcript - who did the talking, and the words that came up most.
In this episode, we dive deep into the evolving landscape of digital marketing and brand storytelling. We explore how the intersection of authenticity, community, and technology is reshaping how brands
Transcribed and scored by The B2B Podcast Index.
1 - >
Speaker 1: Welcome back to another episode of Adventures in Machine Learning. 2 - > I'm one of your hosts, Michael Burke, and I do 3 - > data engineering and machine learning and other stuff at data Bricks, 4 - > and I'm drum by my wonderful co host Ben Wilson. 5 - >
Speaker 2: I investigate cerper column length issues at data Ricks. 6 - >
Speaker 1: Today we are speaking with Barzan. He studied computer science 7 - > at both UCLA and MIT and then moved to the 8 - > University of Michigan as a professor of Professor of Computer Science. 9 - > He still teaches to this day, but recently founded a 10 - > startup called Kibo, which is a fully automated cloud optimizer, 11 - > and they're most famous for their Snowflake integrations. So Barzon, 12 - > as Data Ricks employees myself and Ben, we understand how 13 - > powerful it is to automate back end infrastructure, cluster provisioning, 14 - > that type of thing. But I'm curious as an academic, 15 - > how did you enter this world? Were you using Snowflake 16 - > and head pain points or what was the origin story? 17 - >
Speaker 3: That's a great question. 18 - >
Speaker 4: So I think it's it's probably easier to just start 19 - > from like work Keybo. So it actually our first idea 20 - > is where we're all about how we're going to speed 21 - > up quakes, right, so keyboard and Japanese needs hope. So 22 - > the idea was like when you've tried everything else and 23 - > all other hoopus lasts, like what else can you do? 24 - > So it actually is an interesting uh intersection with the 25 - > data bricks founders. We are actually with some of the 26 - > data bricks as founders. You're working on approximate quity engine. 27 - > The idea was because of Moore's law. So for those 28 - > of you are not familiar, More's law is predicting how 29 - > fast hardware is priced dropping or hardware speed is improving. 30 - >
Speaker 3: And then we were seeing the data volumes. 31 - >
Speaker 4: Were growing growing at the fastest rate than More's law, right. 32 - > So it was pretty scary because like if you're a 33 - > computer scientist or you know math, you know that when 34 - > you have two exponential curves, once you follow that fall behind, 35 - > you're never going to catch up. So if the rate 36 - > of data growth has already surpassed More's law, it means 37 - > if you're happy with your performance database performance, you're going 38 - > to be sad next year. 39 - >
Speaker 3: And if you're sad this year, next year you're going 40 - > to be depressed. 41 - >
Speaker 4: So the idea was like, okay, you know what, what 42 - > is it that we're going to do to close that gap. 43 - > So there's a lot that's been done in the computer 44 - > science community and in the industry database industry, like whether 45 - > it's indexing, data compression, you know, panelism, all of that stuff, 46 - > and that's all great, and you have to do all 47 - > those things, but the idea is that all those optimizations 48 - > are actually your linear speed up. So if you compress 49 - > your data by ten x, you're only getting speed up. 50 - > None of these linear speed ups is going to basically 51 - > help you eventually catch up or get ahead of that 52 - > exponential curve. So that's where we start looking to statistical 53 - > solutions to this problem, and then very quickly we realize 54 - > actually the problem is not just about More's law or 55 - > speed you know, you likes of data bricks or Snowflake 56 - > and other players in the state in the space of 57 - > that amazing job of lowering the adoption barrier to sort 58 - > of analyzing your data and getting insights. But what's happened 59 - > is that because it's so much easier to you know, 60 - > get you know, get up and running and start analyzing data, 61 - > now there's a lot more users and appetitions that's happening 62 - > into data and then happening into a lot more data 63 - > and they're basically combining a lot more data sources. So 64 - > now the cost of this infrastructure is going to know 65 - > and it's just wasn't humanly possible for anyone or each 66 - > to this day, Like it's not possible for you to 67 - > look at you know, squint your eyes and stare at 68 - > like in a one million crazy day and say, you 69 - > know what, I think, here's how I'm going to reduce 70 - > the overall cost. 71 - >
Speaker 3: Right, So that's where the story originated. 72 - >
Speaker 4: We saw a real problem and we're academics and we 73 - > thought about like how we're going to create a solution. 74 - >
Speaker 3: And I was always an outlier even academica. 75 - >
Speaker 4: To be honest with you, because a lot of academics 76 - > are just excited about coming up with theoretical solutions that 77 - > are complicated and they can publish it. 78 - >
Speaker 3: But for me, it was less satisfying. 79 - >
Speaker 4: It was more about how can we create a solution 80 - > that also gets a widespread adoption. 81 - >
Speaker 3: So that's how Keyboard started. We started actually creating this 82 - > data learning. 83 - >
Speaker 4: Platform where we train AI models from how uses and 84 - > applications interact with the data and the cloud, and then 85 - > we start our agents, you know, for those of you 86 - > from our reinforcement learning, which is essentially a major step 87 - > in lll MS, very similar concept. 88 - >
Speaker 3: We learn from how those and actions happen, and. 89 - >
Speaker 4: We start actually pulling different levers in real time and 90 - > start optimizing it. And you know, we came up with 91 - > this pricing model off which I'll talk about it and 92 - > the presentation that you guys are interested. But that's how 93 - > the whole story started. Like we said, you know what, 94 - > whatever money we save the customer, we take a small 95 - > percentage of that so that the incentives outline. So that's 96 - > how the whole story started. 97 - >
Speaker 1: Well, that's that's a super interesting sort of paradigm shift 98 - > from founders we usually talk to because it's usually born 99 - > out of of a frustration in a professional space. They're like, ah, 100 - > this is too slow, this is too expensive. And instead 101 - > you guys went from a philosophical and like academically based 102 - > law approach. I was just wondering if you have seen 103 - > other startups be founded from those sets of principles or 104 - > if it's usually more of I hate this thing, I'm 105 - > gonna go fix it myself. 106 - >
Speaker 3: It's it's a little bit of a you know, it's 107 - > a little bit of both. 108 - >
Speaker 4: To be honest with you, I think what happens is, 109 - > you know, it happens both ways sometimes, like you know, 110 - > to your point, someone been working in the travel industry 111 - > or twenty years, Like this thing. 112 - >
Speaker 3: Is way too complicated. 113 - >
Speaker 4: I'm going to solve this industry, right, So, like I 114 - > call them like founders on a mission where like hey, 115 - > you know, I just want to start a company. Here's 116 - > the space, I understand. Let me work on some cool ideas, 117 - > I think, and our you know, our study with like 118 - > we're talking to a lot of customers like I, you know, 119 - > I wish I could tell you a fancy studio. 120 - >
Speaker 3: One morning I woke up I had this emphany. 121 - >
Speaker 4: But the truth is actually a lot of interesting solutions 122 - > come the other way around. Like you basically are looking 123 - > at the vening the problem in trying to figure out 124 - > like what is it that what you know, what's it 125 - > going to take to bring this solution to the market. 126 - > Like you know, people tell you, oh, I have these 127 - > ten problems, and then you can't just go and solve 128 - > it and then hope that when you come back they're 129 - > going to pay for it. 130 - >
Speaker 3: So you're gonna figure out what is it that drives them? 131 - >
Speaker 4: Like what are the characteristicscept that problem or the solution 132 - > that will be acceptable to them. So short answer is no, Actually, 133 - > but data breaks has a very similar story. 134 - >
Speaker 3: Right. 135 - >
Speaker 4: So your founder's sight, which I personally know, like they 136 - > were seeing that map produce, was really slow. 137 - >
Speaker 3: It didn't make any sense. 138 - >
Speaker 4: They can all this really cool idea of hey, what 139 - > if we kept the data that we running erative competition 140 - > in memory? 141 - >
Speaker 3: And there you go. 142 - >
Speaker 4: That's how Spark was born and got rapid adoption and 143 - > people went from there. 144 - >
Speaker 1: Right, Okay, cool, that's a very interesting version story. Now, 145 - > how much of the internals can you disclose? 146 - >
Speaker 4: Uh? 147 - >
Speaker 3: You know all of it? 148 - >
Speaker 4: No, I mean, we have patents in this space, we 149 - > actually publishing plications in the space. I can you know, 150 - > I won't be able to get into any great details 151 - > of how we you know, train those models and whatnot 152 - > that I can tell me, like, you know, the high 153 - > level workflof you know, how the whole system works and 154 - > to end learn the design principles and whatnot. 155 - >
Speaker 3: And we have you know, a doesen of different. 156 - >
Speaker 4: Algorithms unders even if I want it, I have enough 157 - > time to get to details of everything a lot of Yeah. 158 - >
Speaker 2: Having personally seen the source code for a QUE and Spark, 159 - > it would take several weeks. 160 - >
Speaker 5: I think that exactly. 161 - >
Speaker 2: So do you find that it's a generalizable solution to 162 - > get say eighty percent of the way there based on 163 - > the types of operations that different customers or views are doing. 164 - >
Speaker 5: Do you see, okay, eighty percent. 165 - >
Speaker 2: Of people who are adhering to you know, utilizing CTEs 166 - > when querrying data that structure the data and the lazily 167 - > evaluated instruction set that's submitted, that they can say, Okay, 168 - > we can optimize that really well and it works pretty 169 - > darn good. What do you do with the long tail 170 - > of like somebody writing something almost you look at it 171 - > and like, are. 172 - >
Speaker 5: You intentionally trying to break this? And what do the 173 - > optimizers do with that? 174 - >
Speaker 3: I think that's a that's a good, really good question. 175 - >
Speaker 4: Actually, you know, like I all I've done in my 176 - > entire career pretty much is like I've been teaching databases, 177 - > building databases, and selling databases. Right, so, like I know databases, 178 - > but like you know, when you're saying someone's writing a 179 - > really bad ct like there are quits, we see it, 180 - > Like I'm looking at that quoi and like I've spent 181 - > all my career looking you know, writing sequl quaits. I 182 - > can't optimize this myself, right, So you know that's our 183 - > inside joke is like, you know, we want to be 184 - > that infinitely competent, instantively patient DBA. 185 - >
Speaker 3: Right. But the short answer to your question is yes. 186 - >
Speaker 4: Actually, And the interesting part is we don't even see 187 - > the customers queries. And that was a very intentional decision 188 - > we made from early on, is that we wanted to 189 - > a lot of people think that the hardest part about 190 - > AI is the technology would used to be like a 191 - > decade ago, but now I think we are as a 192 - > field at the place where the technology is not the 193 - > very In many cases, sometimes it's still is, but in 194 - > many cases it's not. It's the adoption barriers that are 195 - > basically stopping us. Right. People worried about paid implementation, the autoi, 196 - > privacy slash security, maintenance, tuning, you know, hallucination, all of 197 - > that stuff. So one of those decisions that we made 198 - > intentionally than was that because you know, I was involved 199 - > in on the startup before Tibo, and I was seeing 200 - > how difficult it is to convince. I mean, think about 201 - > it like a cloud data ware housed is where you're 202 - > keeping the most precious digital app sets up an enterprise. 203 - > Now you're a startup, you're going in and say, hey, 204 - > I have this really cool solution. I'm going to slash 205 - > your bill by fifty percent, which is a lot of money, 206 - > and I'm sure you guys are aware of, Like you know, 207 - > for Cludata warehousing, it's a very expensive solution. 208 - >
Speaker 3: But you know they're not going to trust you with 209 - > the data. 210 - >
Speaker 4: So one of the decisions we made was that it 211 - > has to be a no brainer from a security perspective. 212 - >
Speaker 3: And what it meant was that our. 213 - >
Speaker 4: Models can only learn and train on performance telemachine metadata. 214 - > So not only do we not store any customer that 215 - > we don't even see it, including the quait text we 216 - > hashually quit text. And the beauty of machine learning is 217 - > that it can actually we don't have to make assumptions 218 - > about what is that workload? Like is this an ETL? 219 - >
Speaker 3: Is this a BI? Is it reporting? 220 - >
Speaker 4: Is it ad hoc? Is a data science? Is a 221 - > machine learning? It's it's just a bunch of numbers. Machine 222 - > learning looks at this and says, hey, whenever, whenever I 223 - > see this kind of pattern, I see this kind of behavior, 224 - > I see this kind of cost, and like any other 225 - > you know, human clever, human the agent slurn, they pull 226 - > a lever right, and if it basically managed to save 227 - > the customer money without casting a slow down, the agent 228 - > gets rewarded and learns from that. Whatever it does something 229 - > that doesn't lead to cost saving, it get penalized and 230 - > blends from that. So the answer to your question is 231 - > surprisingly yes, Actually we don't know. We have not seen 232 - > to a single customer for which we've not been able 233 - > to save some money. But what's that percentage? It depends 234 - > on a bunch of problems, how under provision they are, 235 - > how optimize their workout is in the first place, how 236 - > open they are. 237 - >
Speaker 3: To you know, bloodying. 238 - >
Speaker 4: The models get more aggressive, some of them ask the 239 - > agent with a slider, until the agent where it needs 240 - > to be conservative or aggressive, Whether they save eighty percent 241 - > or twenty percent. It just varies from customer to one 242 - > customer to another, but it does actually generalize pretty well. 243 - >
Speaker 1: Well, I have a really saucy que question, go pronouncing 244 - > around a slightly different topic. So Data Breaks has been 245 - > working on this thing called predictive io and it seems 246 - > similar to what you guys do. And all these just 247 - > cloud things have a bunch of data and a bunch 248 - > of resources, to build something similar. How do you guys 249 - > differentiate and how do you guys avoid becoming super surpassed 250 - > by a cloud specific solution like predictive biom No. 251 - >
Speaker 3: That's a very good question. 252 - >
Speaker 4: So look like we basically what we do, like we're 253 - > like super Laser focused on just being a data learning 254 - > We're not trying to replace Snowflake. We don't go to 255 - > a Snowflake customer say you know what, you should go 256 - > to database, And we don't go to your customers and 257 - > say me to miget to Snowflake if you want X, 258 - > y Z. We we're telling people as one of the 259 - > exactly on this call, we basically tell people, whatever data 260 - > stack that you've already invested in, that's great, that that 261 - > what's key is orders of magnitude faster and significantly cheaper 262 - > than that thing that you're already using without keyboard. So 263 - > one of our other design principles has been like we 264 - > should not acquire any data migration, any infrastructure migration. So 265 - > whenever the cloud provider or the cloud data warehouse has 266 - > certain functionality, we actually leverage that. So Snowflake, for example, 267 - > also has a bunch of really clever internal mechanisms. What 268 - > we do is that we never reinvented with because a 269 - > predictive I oh, people will actually try to leverage that 270 - > to some extent, and it's not just the optimization we 271 - > actually provide finops. We have a new technology on the 272 - > same point platform called smart quity rowdy, right, so you 273 - > could potentially use your own predictive I ought to figure 274 - > out where to route those querities, right. So we use 275 - > that to decouple the application the customers application logic from 276 - > the application performance, so the you know, the user, the 277 - > customer can just focus on their use case without worrying 278 - > about costs, without worrying about performance, and just decouple those decisions. 279 - > So you just send those quodies to the smart quid 280 - > outer and they will decide, hey, maybe you know I 281 - > need a small, the square needs a large, just want 282 - > needs a medium and so on. So short answer to 283 - > your question is, we don't invent the wheel. We're not 284 - > trying to replace the underneath, the technology underneath. We take 285 - > advantage of whatever primitive and functionality that's in there, whether 286 - > it's for better insight, better recommendations, or better actions. 287 - >
Speaker 1: But it seems like a sort of shrinking pie. At 288 - > For instance, data Bricks has invested heavily and serverless, and 289 - > they don't really expose knobs, so there's less that can 290 - > be tuned. And so what are the sort of sticking 291 - > points that you anticipate there will be optimizations for the 292 - > next five and ten years. 293 - >
Speaker 4: But that's a very good point, Like, look, you know, 294 - > when you're thinking about people are trying to So if 295 - > it's sort of just like go back and look at 296 - > for example, Big Quiy, another player in the space. Right, 297 - > you can just use spot instances, or you can go 298 - > completely like you know, here's a flat trade, or you 299 - > can say I'm just going to send you the quod 300 - > you figure out what you're going to harun it and whatnot. 301 - > There's always a cost performance trade off, right, you can 302 - > you know when you're going with several as someone else 303 - > is making that decision, you're hiding the knobs and you're 304 - > automating those knobs. 305 - >
Speaker 5: Right. 306 - >
Speaker 3: So I don't think I think the idea of a 307 - > shrinking pie for knobs is a valid question. 308 - >
Speaker 4: But I don't think it's just data learning is not 309 - > just about knobs, because at the end of the day, 310 - > I'm sitting next to the customer's most valuable digital asset. 311 - > I am understanding the data distribution, because that's just the 312 - > first app that doesn't see the data. We have additional 313 - > apps that wants the customers in the platform. They actually 314 - > see the data, They see the querity texts, they see 315 - > all of that stuff. If you if I'm sitting next 316 - > to your cloud data warehouse, I actually understand your data 317 - > distribution more intimately than any single individual organization. So when 318 - > something's out of the ordinary, when it comes to the 319 - > the quality, I'm actually the one. 320 - >
Speaker 3: Like me, meaning the agent right, is the one. 321 - >
Speaker 4: That actually finds out first that, hey, this column never 322 - > had null value. 323 - >
Speaker 3: Suddenly you have a lot of non values. 324 - >
Speaker 4: You're working with a pretty large customer, and tend out 325 - > that like one of the really important columns become null 326 - > for several months and no one will be noticed. And 327 - > so we understand those drastic changes in the data distribution. 328 - > We can actually see certain KPIs. So like you know, 329 - > warehouse optimization, which is what you're referring to, is just 330 - > one use case. Even that use case, I don't think 331 - > it's going to go away because you're actually creating something 332 - > severalless now people are creating more that basically that just 333 - > like moves the bar a little bit like people are not, 334 - > you know, worried about what's the sever size I'm going 335 - > to be using. But they're going to stitch it with 336 - > five other data tools and then build a more complex pipeline. 337 - > And now the question is how they optimize that pipeline. 338 - > One of the major drivers of costs in the cloud 339 - > these days is DBT models, right, Like, you know, you 340 - > created like one hundred and eighty private on the DVD models. 341 - > You know, good luck optimizing that. You could be several less. 342 - > Like you can remove all the knobs you want. At 343 - > the end of the day, the question is the customer 344 - > has to pay X dollars. What can you do to 345 - > reduce that cost? Sometimes the solution is changing the knobs. 346 - > Sometimes the solution is changing equity. Sometimes the solution has 347 - > change the way that you're actually quitting your data. But 348 - > there's a lot more, like you know, we're on optimization, 349 - > workload intelligence, SmartWare, routing, data quality alerts. There's a lot 350 - > of different ways you can expose that. You can leverage 351 - > the understanding that you have of the customers usage behavior 352 - > and expose it to them at different parts in that 353 - > and that data stime. 354 - >
Speaker 5: Hard. 355 - >
Speaker 2: Yeah, I couldn't agree more with your perception of that 356 - > as somebody who many many years ago, back when I 357 - > was doing data science work and like data engineering work 358 - > several times at companies I worked with or customers I 359 - > was working when I was in the field of data breaks. 360 - > You always hit that point where you've migrated to a 361 - > new platform, people start using it, and. 362 - >
Speaker 5: Sort of bad processes have propagated to the new platform, 363 - > and you open it up and like, all right, it's ga. 364 - >
Speaker 2: Everybody can use it, and the issue you first query 365 - > and you're like, man, this is slow. It's faster than 366 - > it was on our old platform, but it's still slow. 367 - > And then you're like, all right, we need to take 368 - > an entire quarter or two quarters and redo the data 369 - > model properly and get rid of all that old tech debt. 370 - > And every time that I've been a part of a 371 - > team that's done that, you open up a whole different problem. 372 - > Right after you get all that fixed, you're like, Okay, 373 - > the query that used to take four hours to run 374 - > now executes in ten seconds because we actually put the 375 - > data where it should be and optimize it, put in 376 - > nexces on stuff, everything exactly. But you still hit that 377 - > there's a finite resource limit that's placed in any business, 378 - > which is the CTO gives you a budget, like there's 379 - > so much money you can spend on this stuff. And 380 - > when you fix all those problems, people just start issuing 381 - > more queries. 382 - >
Speaker 3: Exactly, They're doing more and more exactly. True. 383 - >
Speaker 2: So then you have to like, Okay, the queries aren't optimized, 384 - > how do we teat And then like, what you're tackling 385 - > is the thing that's every time I've tried to do 386 - > it or been part of an organization that's tried to 387 - > do it, it's the hardest thing to fix, which is 388 - > how do you teach people how to interact with an 389 - > optimized system properly? And no matter how much effort you 390 - > put into it, you're never going to be as good 391 - > as an automated service they can do that. 392 - >
Speaker 3: That's that's on hundred percent truth. And that's actually one 393 - > of the. 394 - >
Speaker 4: Common questions that sometimes people ask us is like, aren't 395 - > you afraid that like the snowflocks of the world like 396 - > feel like you know, you're reducing their revenue And I'm like, no, actually, 397 - > I'm just there. 398 - >
Speaker 3: We're just their unpaid customer success department, because. 399 - >
Speaker 4: Yeah, we're just letting you know, those customers get more 400 - > work down less money. So at the end of the day, 401 - > people like to your point, when you optimize their workload, 402 - > they end up actually doing more. You know, it's not 403 - > that they go back to the CTO. Sometimes they do 404 - > more often than not, you know, the bar just shift 405 - > somewhere else. Like now they're going to send more quids, 406 - > they're going to stitch it with five out of data sources. 407 - >
Speaker 3: And that story seems to be you know, repeating everyone. 408 - >
Speaker 2: Yeah, that's what we see with data breaks customers all 409 - > the time. It's like they start off with with ETM, 410 - > they get all their data in the warehouse, and then 411 - > they didn't move on to BI and they had all 412 - > the querry sack because of tech that and they fix 413 - > all that and then it unlocks the mL side. 414 - >
Speaker 5: They'll hire a data science team. 415 - >
Speaker 2: Somebody knows what they're doing on that team, and then 416 - > they'll they'll get some stuff into a you know, maybe 417 - > staging and validate it and eventually to production. You look 418 - > at the account usage over time, like hang on a second, Yeah, 419 - > they've increased ten x as our customers, but they weren't 420 - > doing any stuff before. And if they're a publicly traded company, 421 - > you can kind of look at them and be like, see, 422 - > so for the last four years, like they've doubled in revenue. 423 - >
Speaker 5: Is that because of us? You know, you will you. 424 - >
Speaker 2: Kind of want to take credit a little bit for that, 425 - > like maybe that was two percent us. And sometimes I'll 426 - > say that like, yeah, this unlocked our business insights and 427 - > you can now compete against our competitors. 428 - >
Speaker 3: No, that's that's that's so true. Actually, you know, one 429 - > of the. 430 - >
Speaker 4: I wouldn't say saddest thing, but at least the most 431 - > interesting things that we're seeing is that they're just still 432 - > a lot of data teams that basically are constantly like 433 - > spinning their wheels, trying to reinvent the wheel. They're trying 434 - > to sort of they're too too but because like they 435 - > haven't finding box right and understand that, they're trying to 436 - > sort of manually patch things up, like hey, you know, 437 - > I'm going to do X Y Z, I'm going to 438 - > reduce the costs. Like just imagine how much value would 439 - > be unlocked if if you actually shifted those resources into 440 - > growing your business to your point, right, like you know, 441 - > those all those smart and gas Like for example, you're 442 - > a you're a major game game company or game development company, software, software, games, 443 - > and you know, your. 444 - >
Speaker 3: Data team is just spaying. 445 - >
Speaker 4: They will trying to sort of figure out how to 446 - > leverage snowflake more efficiently, or how to leverage how to 447 - > optimize the data backs workload. Whereas there you know, games 448 - > game development company, they should be focused on like bringing 449 - > the nuts and better version of that game and increase 450 - > the top line instead of being so focused which is 451 - > very common theaday because of the economy obviously, But you're right, 452 - > like I think when when you free up t you know, 453 - > the resources from being consumed by all these you know, 454 - > cost saving and things that are not the core business 455 - > of that that that customer to your point, you go 456 - > back and look at it and say, I'm glad that 457 - > those engineers are actually focused on the growing their business started. 458 - > And how do we pay you know, a little less 459 - > to this particular tool that was supposed to free our 460 - > time up instead of just tending us into you know, 461 - > optimizers for this attitude. 462 - >
Speaker 1: So another question back to sort of your origins from academia, 463 - > What are some of the skills and concepts that have 464 - > been essential in founding kebo, Like what are the things 465 - > that you learned in academia that translate really well to 466 - > being a startup founder, specifically in such a technical space. 467 - >
Speaker 3: It's a good question. I think. 468 - >
Speaker 4: I don't know how much of this would generalize every startup, 469 - > but I can talk about like the kinds of startups 470 - > that look like kebo. 471 - >
Speaker 3: I. 472 - >
Speaker 4: You know, sometimes jokingly say the listen, the reason why 473 - > he was being successful, like we've been going out revenuequick 474 - > pretty quickly, is not because we have really charming sales reps. 475 - > The reality is like our product or you know, I 476 - > shouldn't take her for it, but our product team has 477 - > built a product that's out smarting of the other solution 478 - > out right. So the reason why we can do it 479 - > is because it's just we were not just looking at hey. 480 - > I usually give the example of key value stores, right, 481 - > like there was a there was an era where every 482 - > other week there will be a new key value store 483 - > out there. At some point they run out of they 484 - > run out of names for these companies, right, because then 485 - > it was very easy to build a new key value store. 486 - > You would just and the nice thing about it is 487 - > like there's one hundred plus different key value stores, so 488 - > you never get stuck on anything you don't know how to. 489 - >
Speaker 3: Implement X y Z that's what. 490 - >
Speaker 4: There's ninety nine other products you can look at, but whatever, 491 - > you're trying to create something for the first time and 492 - > it is truly innovative. 493 - >
Speaker 3: Like now it's just a better user interface. 494 - >
Speaker 4: It's just a slightly more optimized version of what everyone 495 - > else has been doing for the past twenty years. 496 - >
Speaker 3: That requires research skills, right. 497 - >
Speaker 4: So one of the nice things about academia is that 498 - > you know in the industry, right, like, if you want 499 - > to pitch an idea to your boss or to the company, 500 - > they think about risk. So oftentimes they try to serve 501 - > out with all those apples in one basket. They say, 502 - > you know what, that's hydro art high risk, which is 503 - > usually shorthand we're saying we're not gonna do it right. 504 - > But you hear this a lot at you know UH 505 - > in in UH in the industry. But like in Acadinia, 506 - > it's the opposite. You get rewarded for taking on hairy, 507 - > big problems and considering solutions that no one else has 508 - > doubled because even when you fail, you learn from it 509 - > and you go do something. That's because that's what academy 510 - > has made for, right, like for for for people to 511 - > go and freely innovate and push the boundaries and things 512 - > like that sort. So I think research skills, which doesn't 513 - > mean you need to have a PhD, but like the 514 - > ability to take on an open ended problem, I think 515 - > outside the box, come up with a solution that maybe 516 - > no one else has thought about and kind executely done. 517 - > I think that's definitely one area. And the other idea 518 - > is this, like this whole thing about failing fast. We 519 - > keep talking about failing fast, but that's pretty much what 520 - > happens in academia, right Like, so the still cycle in 521 - > the industry, right if you're thinking about for example, B 522 - > two B so right like, you have to come up 523 - > with a you know, usually an MVP takes at least 524 - > two quarters. Right after that you're working with beta customer 525 - > that's not a quarter. And then then we talk about 526 - > a really fast like product to market kind of cycle. 527 - >
Speaker 3: And then you know, you have to. 528 - >
Speaker 4: Chain the sales team and you start selling some like 529 - > every customers getting fraction and whatnot. Nagadimia, you're write, you 530 - > live life writing one paper at the time. So if 531 - > you have an idea, you submitted to a conference, you know, 532 - > as soon as like you have some you know, compelling results, 533 - > You write up a paper. Your code could be complete crap, 534 - > but you just have a proof of concept. You write 535 - > up a paper, you run a bunch of experiments to 536 - > see if it works or not. You don't have to 537 - > go higher sales people. You don't have to go, you know, 538 - > spend millions of dollars on marketing. You just basically go 539 - > out there and and that paper. And then I didn't 540 - > get peer reviewed. And if it's a bad idea, you'll 541 - > find out in like most conferences and computer science, you 542 - > hear back within two months three months laters, right, So 543 - > you have to fail fast, this idea of being scrappy, 544 - > you know, and making sure that you know you see 545 - > somebody's off before you invest too much into it. I 546 - > think those two things from our company DNA really did 547 - > help help us out a lot. 548 - >
Speaker 2: A keyboard must have been something in that lab that 549 - > you are in, because that exact approach has actually carried 550 - > over into data bricks R and D. 551 - >
Speaker 3: It's pretty common. 552 - >
Speaker 2: This is, but I've heard from other people that have 553 - > come from fank companies into data bricks and their remarks 554 - > are like, I can't believe we're allowed to do a spike, 555 - > and like, yeah, we have to do design talks and stuff, 556 - > but we get time to do a prototype, and sometimes 557 - > you'll give us like, hey, go see if you can 558 - > figure this out. Like take these like you five people 559 - > from all these different teams, just just take six weeks 560 - > and play jazz. Figure out what you can come up with. 561 - > And sometimes it's a failure, like an abject failure. Well 562 - > we'll even release it the private preview, get like twenty 563 - > customers trying it out, and the response. 564 - >
Speaker 5: Is like, we don't know about this. 565 - >
Speaker 2: And then four months later we have version two point 566 - > zero that's in public preview. 567 - >
Speaker 5: And people are like, this is amazing. Where was this 568 - > all my life? 569 - >
Speaker 2: But yeah, that iterative process of just failing, like failing 570 - > really hard. Sometimes it is critical to like how we 571 - > release products the way that we do, but a lot 572 - > of companies in the tech space just don't do it. 573 - >
Speaker 3: They don't do it. 574 - >
Speaker 4: Know you're spot And I think it's just also a 575 - > little bit about like getting people with a research mindset 576 - > because like you know, like as someone I've been writing 577 - > code from an early age, right, like I was a 578 - > program before I was a researcher. But like if I 579 - > had to confess like researchers usually doing the best quality 580 - > code and sometimes we like crappy code because we're just 581 - > trying to prove that concept and walking out and that 582 - > drives solid engineers and experience, you know, technically sometimes crazy, 583 - > you know, how can you put something. 584 - >
Speaker 3: Like this right? 585 - >
Speaker 4: But like I think if you can, you need an 586 - > environment where people like researchers can being the research skills, 587 - > solid architects can being their expertise and like can help 588 - > like transition once those ideas are divers or tried out, 589 - > help transition to product that skills right, like something that's 590 - > robust and production quality. I think a lot of amazing 591 - > things happen, like researchers on their own another or create 592 - > something that actually you know works at scale. But like 593 - > if you can pair them with with engineering teams that 594 - > you know are are solid and can take those ideas 595 - > and transition and like obviously that means both camps have 596 - > to get out of their comfort zone a little more, right, 597 - > But I think when you when you have an environment 598 - > that's conducive to that kind of collaboration, just amazing things 599 - > happen to your point, but. 600 - >
Speaker 2: There have been some comfort zone transitions, but it's exciting, 601 - > Like everybody gets so enthused about it on both sides, 602 - > because you get the researchers. A lot of people come 603 - > from that we've hired, they have like ten plus years POSTCRAD, 604 - > they've been doing research at Berkeley or Stanford, MIT or something, 605 - > and they come in they're like, Oh, this code's complex, 606 - > and engineers are. 607 - >
Speaker 5: Like, oh, what are you working on? 608 - >
Speaker 2: I want to want to see it, and there's no, 609 - > there's not like like I think there's a brief moment 610 - > of panic on both sides, but then everybody's like, hey, 611 - > let's work together and like let's team up and let's 612 - > make this awesome. And you just see everybody grow together 613 - > because you're expanding the mind of engineers to see like 614 - > what is theoretically possible and it unlocks a lot more 615 - > creativity on their side. And then the R and D 616 - > researchers eventually they're writing like production grade code within a 617 - > year or so. 618 - >
Speaker 5: So you're like, yeah, it's a win win all around 619 - > spot exactly. 620 - >
Speaker 3: So it's it's it's exactly like, how this kind. 621 - >
Speaker 1: Of which process do you both like more? Do you 622 - > like research spikes or more engineering focused work. 623 - >
Speaker 4: I think we do both, but I you know, I 624 - > think it depends what you're trying to do right, I 625 - > think now. 626 - >
Speaker 1: But personally, like, which do you enjoy more? 627 - >
Speaker 3: Oh? I definitely enjoy research spikes. 628 - >
Speaker 4: I think it's just like, you know, like I said, 629 - > like you can never pay me enough to go and 630 - > create another key value store. 631 - >
Speaker 3: Like I'm just the kind of. 632 - >
Speaker 4: Person like life is too short, Like if I want 633 - > to do something, I want to be the first person 634 - > doing right. So research spikes usually have that kind of 635 - > flavor where like, you know, this is an idea. I 636 - > might come back and say, guy, that's that's not promising. 637 - > You know that's not going to work. But you know, 638 - > when you do come back and you come over with 639 - > a you know, new solution no one else has thought 640 - > about and actually works, you get to you know, big, 641 - > you know spike of dopamin and it's that makes it 642 - > all worth it, at least personally for me. 643 - >
Speaker 2: I enjoy three distinct points in that development process. The 644 - > first one is I love seeing all of my dumb 645 - > ideas fail in the beginning because it just it shortens 646 - > the path to getting something that might work. 647 - >
Speaker 5: And it's also kind of fun. I like seeing like 648 - > I think I was telling you, Michael, the other day, I. 649 - >
Speaker 2: Was doing something late at night and getting some CI 650 - > set up and a package that I'm working on, and 651 - > I wrote some really terrible code because it was like 652 - > twelve thirty in the morning, pushed it to get hub actions, 653 - > and then I crashed the runners, like killed them all 654 - > basically effectively, like a stack overflow, and. 655 - >
Speaker 5: I just looked at it. 656 - >
Speaker 2: I was like, I'm going to bed, but I kind 657 - > of chuckled to myself when it's been a while since 658 - > I've broken something like that. And then the next morning 659 - > I look at what I actually submitted, I'm like, yeah, 660 - > don't code any of that, tired dude, and fixed it 661 - > and then it passed. 662 - >
Speaker 5: I'm like, all right, sweet. 663 - >
Speaker 2: But I also love the transition from the proof of 664 - > concept works and buy in has been signed off, like 665 - > it's been effectively peer reviewed amongst peers at the company, 666 - > and then banging out that first production grade version of it. 667 - > I love that experience. So, Okay, I know how bad 668 - > my code was. How do I make this actually usable 669 - > and extensible and maintainable, and how do I just kill 670 - > all of this complexity that I had to build in 671 - > the script. 672 - >
Speaker 5: That I wrote that's very enjoyable. 673 - >
Speaker 2: And then finally the release not not the response, I 674 - > don't really care about that. I actually look for people 675 - > like who use it that then tell me why it's broken, 676 - > because I love fixing the bugs on the like the 677 - > first few iterations. 678 - >
Speaker 5: I love that experience. It's not like I. 679 - >
Speaker 2: Know this code because I wrote this crap and I 680 - > love I'm like, yeah, I'm going to totally fix that 681 - > as that's my dope. 682 - >
Speaker 4: I mean, well, I love it, Like I also like 683 - > how you kind of like the three stages, like you're 684 - > seeing the true value of each of those three stages, 685 - > like and liking it for what it you know, what 686 - > it is, like, Hey, I would not get from A 687 - > to C if they didn't. 688 - >
Speaker 3: Have put B in the middle. No, I I that's 689 - > that's that makes a lot of sense. Yeah. 690 - >
Speaker 1: I think my response is I really like the research aspect, 691 - > but it's sort of a product of my job because 692 - > I don't have the opportunity to build really complex extensible 693 - > frameworks that have like cool designs, Like I'm writing a 694 - > thousand lines of code maybe two thousand for like a 695 - > typical project, and the really fun thing is trying the 696 - > art of the possible and seeing like can we make 697 - > this work? Like what creative ideas for attacking a problem 698 - > in it from a different direction, can I employ it 699 - > to make it it's successful? So yeah, it's interesting, but 700 - > they both have their prison cons And it's interesting Barzon 701 - > that your your angle is research because the computer science 702 - > is very fundamentally implementation and optimization focus. Would you agree 703 - > or do you think. 704 - >
Speaker 3: There's a. 705 - >
Speaker 1: There's a lot of what or would do you think 706 - > there's a lot of sort of innovation and like groundbreaking 707 - > like far out their ideas. 708 - >
Speaker 4: It actually depends on what discipline you're looking at, right, So, 709 - > like I might get into trouble for saying this, but 710 - > for example, if you just look at databases as a field, 711 - > like which is my own field? So I feel like 712 - > I'm allowed to say things like this. I think the 713 - > field has kind of plateaued. You go to as you know, 714 - > you go to the events, you go to like these 715 - > places where they talk about innovation, and you're looking at 716 - > this and saying, like that's really cool that like that 717 - > ship now has X Actually Oracle had that like thirty 718 - > years ago, right, like, Hey, I'm so glad that you 719 - > guys do auto indexing here, but that happened here, or 720 - > like you have this storage optimized think here to use 721 - > this compression. Well you know what, actually Verdica had that 722 - > like twenty years ago. So it's the field is popular. 723 - > It doesn't mean there's no innovation, but like if you 724 - > just trying to build another database, a lot of it 725 - > is being tried and I I'm not saying there will 726 - > never be enough innovation. I'm just saying the number of 727 - > new ideas that are like radically new and actually are effective, 728 - > it's we're running out of those ideas. 729 - >
Speaker 3: Like the field has matured, which is a good thing, right. 730 - >
Speaker 4: It means we can go and build the next set 731 - > of you know, AI enabled AI AI enabling. 732 - >
Speaker 3: Applications on top of what we've learned. Like now we're 733 - > we wrote a. 734 - >
Speaker 4: Paper a few years ago go to my former PhD 735 - > students about database learning. So the idea was like, okay, 736 - > now let's see you do have a database that's optimized. 737 - > But every time, you know, if I keep asking you 738 - > The example I give is about cars, right If I 739 - > let's say that you're like me and you don't know 740 - > anything about cars, right, then if I keep you know, 741 - > if I ask you a question about this particular model 742 - > of Ferrari, like you're going to go online and look 743 - > it up and give me the answer. If I keep 744 - > asking you questions about cars, you're going to keep like, 745 - > you know, googling it. But after two three days, you're 746 - > gonna pick up a few things you're gonna learn. It's 747 - > gonna take you less and less time to come up 748 - > with an answer to car related cars, right because we're humans, 749 - > like we learn. The databases don't learn, you know, asides 750 - > from like very basic things like hey, I cash this data, 751 - > I cash that result before the data changed you. Every 752 - > time you go mediquated it, this does a bunch of work, 753 - > send you the results back. For the most part, that 754 - > work is lost. Afterwards you go back and it starts 755 - > like the databases don't learn. So that the vision that 756 - > we basically presented and we actually built a proof of 757 - > concept on it, was like, how can we build a database? 758 - > It actually learns over time, It becomes smarter every time 759 - > that you quit it. You can think about it like 760 - > if I asked you, hey, what's the average sales for 761 - > this particular region her department, and then tomorrow asked another 762 - > question kind of overlapt like maybe said, hey, what's a 763 - > total number of transactions per region for the entire country? 764 - > The fact that I know something about that region should 765 - > help me come up with an answer to the second 766 - > question a little bit faster. Right, So, I think there 767 - > is still innovation, but it's very build specific. Certain sub 768 - > disciplines within computer science are world research. People either have 769 - > moved on or they need to move on. They have 770 - > people who still haven't moved on, and they still like, 771 - > you know, cibierating over similar ideas. Hey, actually I found 772 - > this corner case where I can make the indexing like 773 - > five percent more efficient. But there's a lot of interesting things, 774 - > especially like the time you're living and with other learns, 775 - > with machine. 776 - >
Speaker 3: Learning, with. 777 - >
Speaker 4: You know, hardware acceleration that we can we can still 778 - > actually come up with pretty cool ideas, Like human mind 779 - > doesn't not other cool ideas. 780 - >
Speaker 3: That's a nice thing. 781 - >
Speaker 4: It's just like, you know, maybe you fix something, you 782 - > don't create a new discipline. 783 - >
Speaker 1: Curious, for both of your guys' opinion, what are the 784 - > frontiers that you're excited about or the new piece of 785 - > technology that in the database and data querrying space that 786 - > you think are going to be game changing? 787 - >
Speaker 3: Then wasn't just that. 788 - >
Speaker 2: I think for data querying, the ability to map to 789 - > an entire data warehouse or entire system of rdbms like 790 - > implementations that exist in an organization and for you to 791 - > be able to talk to an agent and ask a 792 - > very complex question and you get the accurate response from 793 - > all of that without you having to build all of. 794 - >
Speaker 5: The interfaces to that, because. 795 - >
Speaker 2: Today you can theoretically do that, right, you can create 796 - > a bunch of tools that all issue all of these 797 - > different queries to all of these different platforms, or you 798 - > can you know, have like basically fine tune the model 799 - > on the metadata of. 800 - >
Speaker 5: Your tables and your databases. 801 - >
Speaker 2: And I don't think that anybody's gonna pick that up 802 - > to get to, you know, a hind ninety percent accuracy 803 - > response rate. 804 - >
Speaker 5: Like we offer something. 805 - >
Speaker 2: Called Genie right at Data Bricks, and that's language Model 806 - > interface to query Unity Catalog tables and in demos. It's incredible, 807 - > like amazing. I've played around with it, I'm doing integrations 808 - > with it. I'm like, man, this is so cool the 809 - > fact that I can, you know, put one hundred column 810 - > table with a million rows and I can ask it 811 - > just plain language questions and it figures it out. And 812 - > I can do this with five different tables and it'll 813 - > generate those queries for me, and it's pretty performant because 814 - > of that optimized engine in the background. But then I 815 - > point it to our internal tables or data that I 816 - > had written two years ago into in the Unity catalog 817 - > during the demo days of. 818 - >
Speaker 5: That, and I usually the same query and it loses 819 - > its mind. 820 - >
Speaker 2: And then I'm like looking at it, like why why 821 - > does it work so well on these tables that I created, 822 - > you know, last month, and that my old data. 823 - >
Speaker 5: It's it's just not good. 824 - >
Speaker 2: And then I just go into the UI and I'm like, oh, yeah, 825 - > there's no metadata here, Like there's no comments anywhere explaining 826 - > what this table is, what's in it, or the conditions 827 - > for the ETL that is actually putting the data in. 828 - > And then the column names are almost intentionally obfuscated because 829 - > I was just doing shorthand nonsense and I have no 830 - > parameter comments anywhere of like what this column contains, so 831 - > he's making guesses it's inferring from what metadata it actually has, 832 - > and I'm just like, Okay, there's. 833 - >
Speaker 5: Got to be a better way to do this. 834 - >
Speaker 2: So I think that the golden goose out there is 835 - > for a parson on this team to figure out how 836 - > do I do that with the table. How do I 837 - > generate the metadata that is highly accurate, that is contextually 838 - > relevant to this business in a way that you know 839 - > interfacing with an agent will work properly. 840 - >
Speaker 1: Yeah, we just real quick before you jump in. I 841 - > have so much beef with Genie right now. The account 842 - > teams that Data Bricks have sold the proof of concepts 843 - > like five different customers and then the customers are like, 844 - > oh great, so now you're going to build me this agent. 845 - > I'm on three of those projects right now, and we 846 - > just have to like lower the expectations three order of 847 - > magnitude because it's just not there yet. It's a really 848 - > cool technology and it will be there soon, but the 849 - > demo is not what it is in reality. Yeah, over 850 - > to bar Zone. No. 851 - >
Speaker 4: I think the explanation is one of those years I'm 852 - > as really excited about. But like if I'm kind of 853 - > zooming out, like it's very easy. Like if you ask me, 854 - > what's like if I had a magic wand and that 855 - > could solve any problem, I would obviously say world hunger, 856 - > and I can't like that. Also realistic about my own 857 - > skill set, right, So I think the most important thing, 858 - > like the way I'm looking at it as someone who's 859 - > excited about innovation, but also like I want to make 860 - > sure it's practical, gets adoption right, and part of adoption, 861 - > like to me, has four legs and one of it 862 - > that has to work right, not to just be on 863 - > the develop to Ben's point, right, like it has to 864 - > work otherwise you get the best people are really excited 865 - > and to get really frustrated, which I think is a 866 - > big one of the barriers to some extent with AI 867 - > is like people if they if the level of excitement 868 - > doesn't match that expectation, then they get burned out and 869 - > I don't know when the next time that the CIO 870 - > was going to sign off. 871 - >
Speaker 3: On something at the word eleven, right. 872 - >
Speaker 2: So. 873 - >
Speaker 4: If I'm looking at it from that perspective, I think 874 - > I think the key to success would be to focus 875 - > on what the intersection of what can be oblimated up 876 - > automated and what should be automated. Sometimes people try to 877 - > automate things shouldn't be automated, and or the things that 878 - > should be automated but cannot be automated with today's technology, 879 - > and because they're can accurate, they're inefficient, unreliable, all those reasons, right, 880 - > So if I'm looking at that intersection of what can 881 - > and should be automated. 882 - >
Speaker 3: One thing that's actually working on that thing is very exciting. 883 - > Is like with elms, we've seen massive success with actually 884 - > quite rewriting. Like as someone who's been like just. 885 - >
Speaker 4: Dealing with quities for the past twenty years, it can 886 - > actually rewrite quitties that we never thought possible. But it's 887 - > not just like, hey, chat GPT, can you please rewrite 888 - > this quity for me into a more efficient form, because 889 - > actually four out of five times, or I should say 890 - > eight out of ten times, it actually gets a quated 891 - > either doesn't even compile, or it actually compels, but it 892 - > gives an incorrect answer, or it compels gives the correct answer, 893 - > but that's actually slower than the one I started. So 894 - > like we've created this framework around and like the papers 895 - > out there for those who are interested in audience is 896 - > called the general light. 897 - >
Speaker 3: We actually we've created this really cool. 898 - >
Speaker 4: Cycle where we basically get that we're interacting with the 899 - > LM actually come up with what we call human readable 900 - > rewrite rules. So like when we ReLit it, we actually 901 - > turn it once we valuve it turned into a rule, 902 - > and then when the quit comes and we use those 903 - > rules as actually as hints. 904 - >
Speaker 3: To the l ELM. 905 - >
Speaker 4: So now we basically get pretty accurate, like ninety plus 906 - > percent accurate with in the sense that we can actually 907 - > whenever we relte the quity, we have pretty high confidence 908 - > that's actually correct, and it's actually more efficient than the 909 - > original equity. And the nice thing is that this database 910 - > of human I forgot what we call it in the paper. 911 - > I think it's human understandable or human rewrite rules something 912 - > like this HR two L something like that. That database 913 - > actually keeps growing, So it's like more along this enough 914 - > creating the database that keeps getting smarter over time, like 915 - > chat GBT, the more people are interacting with it, it's 916 - > also getting smarter and smarter. So like creating a system 917 - > that gets smarter over time the more we use it, 918 - > I think is also super exciting. But I think we 919 - > will get to a place like when that we will 920 - > be able to explain a lot of interesting things like hey, 921 - > why did my sales negotiate department of this particular Walmart 922 - > store you know go down last month compared to the 923 - > you know, other stores, comparable stores, right, And we will 924 - > never be able to fully automate experimentation and causality. But 925 - > at least we will be able to show them most 926 - > likely causes to the domain expert. We will then have 927 - > that domain expertise which should not be automated or cannot 928 - > be automated at least today, to tell us, hey, you 929 - > know what, these are the top series. And I think 930 - > it's because we had too many people, you know, out 931 - > of office, or there was this local event at this 932 - > other place. 933 - >
Speaker 3: That was not here. 934 - >
Speaker 4: So I think that's the that's the line that I'm 935 - > really excited about, just working on things that can and 936 - > shoot the automate. 937 - >
Speaker 2: Yeah, that example brought to mind an old example that 938 - > I used to use when when teaching new data scientists 939 - > the teams at past companies about the difference between correlation 940 - > and costality and intelligence systems, and like, here's this model, 941 - > and I had this data set that I would always 942 - > use that was it was basically like year round temperature 943 - > at a park in New York City, and then another 944 - > column was like amount of ice cream sold. And you 945 - > build a very simple model, a regression model, and then 946 - > use explainability tools and costality tools on that data. And 947 - > of course it's like, hey, I want to optimize sales, 948 - > and what does it come up with? It's like the 949 - > thing that you need to change is just increase the temperature. 950 - > And that's definitely trying to teach people like, hey, be 951 - > careful of how you interpret things that come out of, 952 - > you know, an algorithm exactly. 953 - >
Speaker 5: I think that. 954 - >
Speaker 2: The thing with like the explosion of jen Ai and 955 - > its popularity and it's democratization. The only thing that I 956 - > see as potentially disillusioning in that as these these capabilities 957 - > become greater and greater over time, I'm like, hey, I 958 - > can query all my data and I can ask whatever 959 - > question I want, and I can bolt onto this tool 960 - > that's going to do this causality analysis for me, and 961 - > somebody's like, inevitably a system is going to be built 962 - > that has those features that can do these sorts of 963 - > things and it can query the right data. And then 964 - > somebody's going to say, how do I make my sales 965 - > go up? And they're going to ask that to the system, 966 - > and the system's going to go and it's not going 967 - > to say increase the temperature of the planet Earth in January, 968 - > but it'll it could do something similar to that in 969 - > their business, and they might not know like, oh, maybe 970 - > if I, yeah, focus my efforts here turns out you're 971 - > cannibalizing from another part of your business, and you know, 972 - > creating chaos or whatever. I think with incredibly intelligent and 973 - > reliable systems, it could create trust issues with people with 974 - > those systems. Is that something that in academia people are 975 - > thinking about. 976 - >
Speaker 4: I think so, not as much and not as many 977 - > as but there are some actually. But you published a 978 - > paper called dB Sherlock a few years ago. It wasn't 979 - > using l lens, but the idea was like, how we 980 - > can actually incorporate CAUs on models into a system that 981 - > can show the most likely causes and then use it 982 - > the cause on model to actually help users so that 983 - > we don't tell people what caused. 984 - >
Speaker 3: The rein was that you know your wife took. 985 - >
Speaker 4: Down Berella, but we say, hey, these are most correlated 986 - > with each other, and then we can use causality models 987 - > so the system learns over time. 988 - >
Speaker 3: But I think there's some people we are looking into it, 989 - > but not as. 990 - >
Speaker 4: Many, to be honest with you, not as many as 991 - > I you know, wish these days. 992 - >
Speaker 2: So I've got a silly question for you. You've been 993 - > in this space for a while and been doing research 994 - > for a very long time, and we're likely exposed to 995 - > the things that everybody thinks is pure magic nowadays. 996 - >
Speaker 5: About like, oh my god, chat GBT is the best 997 - > thing ever. It's it's so smart. 998 - >
Speaker 2: Anybody who's been in a like dealing with advanced computer 999 - > science for decades is going to look at that and 1000 - > be like, Yeah, we had these like a while ago. 1001 - >
Speaker 5: They've been around a while. 1002 - >
Speaker 2: Maybe not transformers models, they're they're slightly more advanced, but 1003 - > they're growing off of the shoulders of giants that they 1004 - > came before. Were you doing like a table slap or 1005 - > knee slap with a bunch of other professors saying I 1006 - > called it? I knew it was going to happen this year. 1007 - > Where my grandma knows the name of something that is 1008 - > involved with artificial intelligence. 1009 - >
Speaker 3: That's a really good question. 1010 - >
Speaker 4: I think when you spend a lot of time in 1011 - > a space, you actually see like certain things that become trivial, 1012 - > like or look becomes certain to you and become clear 1013 - > to you, right, but like to outsiders because that's all 1014 - > you know, right, Like if all you've done all your 1015 - > career is like this very narrow area, which is you know, sadly, 1016 - > the situation with a lot of us in academia is 1017 - > like we know everything about a very little narrow topic, right, 1018 - > so it becomes pretty clear, but to outside it looks 1019 - > like magic. So yeah, I would say, like, you know, 1020 - > I mean I had students who work time like transformer 1021 - > models and whatnot, so like we were seeing the advances 1022 - > that are coming. But you know, I think what surprised 1023 - > all of us is how quickly the public kind of 1024 - > was impressed with it. Right, Like we go to a conference, 1025 - > we say, hey, we improve this accuracy that like you know, 1026 - > half a percent, and we clapped for each other. We 1027 - > get excited, right, but like eventually when it becomes good 1028 - > enough that everyone else also gets excited about it because 1029 - > they're not there in the journey where like it was 1030 - > just growing a little by little, little by little, where 1031 - > it's harder to see it like they saw like hey 1032 - > there was like sci fi movies and now this is 1033 - > actually here. 1034 - >
Speaker 2: So right, yeah, we even got to see that over 1035 - > the last you know, eight years or so at data 1036 - > bricks with even traditional IML, where you look and like 1037 - > the first couple of months or probably the first year 1038 - > that mflow is out and you're looking at the statistics 1039 - > of like how many people are saving what types of models. 1040 - > You're like, oh, yeah, we've got like one hundred users 1041 - > that saved sk learned models and deployed them and a 1042 - > bunch of people doing an extra boost Like this is exciting. 1043 - > And then you look now and you're like, how many 1044 - > millions of these were were saved in the last week alone. 1045 - > It has become so commonplace, like every business has these. 1046 - >
Speaker 5: Things, and not just one. 1047 - >
Speaker 2: But you look an account might they might be hitting 1048 - > that API for logging that thing five hundred thousand times 1049 - > a week. 1050 - >
Speaker 5: It's like, wow, that's crazy. It's so commonplace. 1051 - >
Speaker 2: But ten years ago that would have been like whoa, 1052 - > this is state of the art. And then people that 1053 - > have been doing that stuff for a long time, Like 1054 - > when I came in and would talk to to like 1055 - > new accounts that we got on, they're like, we want 1056 - > to learn more about this this new thing called data 1057 - > science and we want to like understand it, like new thing. 1058 - > This has been around for a long time, like but 1059 - > way before I was born. 1060 - >
Speaker 5: They're like, what, no, we just heard about this thing 1061 - > you can do. 1062 - >
Speaker 2: I'm like, yeah, the paper for that was written like 1063 - > before one. 1064 - >
Speaker 5: Before computing, so. 1065 - >
Speaker 2: That speed like what you talked about it surprised me 1066 - > a little bit. I didn't think it would hit psyitgeist 1067 - > level of like, hey, everybody knows this thing and everybody's 1068 - > got an account on this thing. It's exciting, but it's 1069 - > also very surprising. 1070 - >
Speaker 3: No spot on exactly. 1071 - >
Speaker 5: Cool. 1072 - >
Speaker 1: So I know we're coming up on time. I'll quickly 1073 - > summarize really interesting conversation. Some things that stood out to me. 1074 - > Our research skills are very valuable for innovation, and in 1075 - > academia you can typically learn the fundamentals of research, at 1076 - > least one would hope. And then also fast failure is essential. 1077 - > Sort of at a macro level, a lot of organizations 1078 - > are turning off knobs, so there's less configuration and customization 1079 - > you can do. But despite that, there will always be 1080 - > additional layers of infrastructure to optimize. People will start using 1081 - > those as discrete blocks in more complex systems. And then 1082 - > some future areas of innovation that we're excited about. Our 1083 - > agentic querying and then query rewriting, And if you guys 1084 - > are curious about the paper, it's called query rewriting via 1085 - > large language models. So Barzon, if you want to learn 1086 - > more about you or your work, where should they go? 1087 - >
Speaker 4: If you google my name, or go to Kibo dot ai. 1088 - > That's our company. You can get live demos of what 1089 - > we do. You're using you know, snowflake or in your 1090 - > cloud to table house in any capacity, and you're interested 1091 - > in auto optimizing it, you know, diverting some of that 1092 - > manual effort or infrastructure bill to some other areas. 1093 - >
Speaker 3: Of your business. It's sound a keyboard that ai. 1094 - >
Speaker 4: Or google my name and look at my academic homepage 1095 - > with a bunch of papers, or reach out to me violected. 1096 - >
Speaker 3: Cool. 1097 - >
Speaker 5: Thanks so much. 1098 - >
Speaker 1: All right, Well, until next time, it's been Michael Burke 1099 - > and my co host and Wilson, and have a good 1100 - > day everybody. 1101 - >
Speaker 3: Thank you so much. Nex
Other episodes covering the same guests and topics, from across The B2B Podcast Index.