The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Data Radicals
Data Radicals artwork

Why AI Builders Need a Metadata Goldmine with Chris Aberger, VP at Alation

Data Radicals · 2025-06-18 · 47 min

0:00--:--

Key moments - from our scoring

Substance score

48 / 100

Five dimensions, 20 points each

Insight Density10 / 20
Originality10 / 20
Guest Caliber13 / 20
Specificity & Evidence10 / 20
Conversational Craft5 / 20

Numbers Station started in 2021 with a focus on applying large language models to structured data, initially targeting data transformations before pivoting to business-user workflows. Chris Aberger describes the journey from chat-with-your-data interfaces to a multi-agent platform capable of executing complete workflows - not just surfacing insights but taking actions like generating reports, finding vendors, or triggering automated processes. The critical realization that drove the acquisition by Alation: 80% of Numbers Station's engineering effort wasn't building AI agents, but curating metadata. Aberger found that organizations lack clean, accessible metadata about their data models, column descriptions, and query patterns, making it nearly impossible for agents to operate reliably in production. Alation's existing metadata platform - what Alation calls its data product (combining knowledge graphs and semantic layers) - represents exactly what Numbers Station needed to accelerate deployment. This acquisition positions Alation to help data teams become builders of AI applications rather than just maintainers of infrastructure, addressing the budget pressure enterprises face while capitalizing on the broader trend of generative AI democratizing technical capabilities.

Key takeaways

  • →Metadata curation, not model training, consumed 80% of Numbers Station's effort - making Alation's existing metadata platform the strategic asset that justifies the acquisition.
  • →Multi-agent architectures that retrieve existing insights from BI tools (Tableau, Power BI, Looker) before querying databases directly avoid reinventing the wheel and reduce computational overhead.
  • →The feedback loop between production AI applications and metadata quality is essential: agents perform real work, users see results, and that closes the loop to continuously improve metadata and expand agent capabilities.
  • →Chat-with-your-data is a distribution mechanism and organizational entry point, but text-to-SQL alone is an unsolved, commoditized problem - real value emerges when agents move beyond insights to action-taking workflows.
  • →Data teams must evolve from supporting infrastructure to building AI-powered applications, with the right platform removing the production deployment bottleneck that stops prototypes from going live.

In this episode

  1. 1Introduction and Chris Aberger's Background
  2. 2Founding Numbers Station and Early Vision
  3. 3Product Evolution from Data Transformations to Chat with Data
  4. 4Moving Beyond Text-to-SQL: Multi-Agent Framework and Workflow Automation
  5. 5The Critical Role of Metadata as the Foundation for AI Agents
  6. 6The Elation Acquisition and Metadata Synergy
  7. 7Data Teams as Builders and the Production Challenges of GenAI

Mentioned

ElationNumbers Station AIChris AbergerSamanova SystemsGoogleAppleIBMStanfordVenki GuntyOpenAIJLLTableau

Guests

Chris Aberger

Topics in this episode

AI agentsSemantic LayerLarge language modelsMulti-agent systemsKnowledge graphsmetadata managementData transformationText-to-SQLNumbers StationAlation

Questions this episode answers

Why did Numbers Station spend 80% of its effort on metadata instead of building AI models?

Production AI agents require accurate, well-organized metadata about data models, column descriptions, and existing data assets to operate reliably. Without clean metadata, agents fail in real-world deployments despite working well in prototypes.

How is Numbers Station's multi-agent approach different from simple chat-to-SQL tools?

Numbers Station's agents first check existing BI tools (Tableau, Power BI, Looker) and documentation for answers before querying databases, and they're architected to take actions - like generating reports or searching for vendors - not just surface insights.

What was the strategic value of Alation acquiring Numbers Station?

Alation's metadata platform (combining knowledge graphs and semantic layers) eliminates the data curation bottleneck that slowed Numbers Station's deployment cycles, allowing the combined company to help enterprises build production-ready AI workflows faster.

Why is chat-with-your-data not a sustainable business on its own?

Chat interfaces are easy to prototype and commoditized (OpenAI can generate SQL directly), but customers really want complete workflows that get business outcomes done - chat is just the entry point, not the end state.

How do data teams transition from supporting infrastructure to building AI applications?

They need a platform that makes production deployment easy; prototyping with generative AI is fast, but hardening it for production requires both clean metadata and an agentic architecture that can integrate with existing systems and take actions.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

10 / 20

The episode contains a handful of genuinely useful claims for data/AI practitioners - the 80/20 split of metadata vs. agent-building effort, the 'unsolved problem that everyone thinks is solved' characterisation of text-to-SQL, and the AI engineer identity crisis framing - but these are heavily diluted by extended anecdotes, mutual congratulation, startup platitudes, and M&A small-talk that consumes roughly half the runtime.

almost like I would say 80% of our time at Number Station was not spent building those agents, it was spent getting that metadata correct
it's largely still an unsolved problem that everyone thinks is solved. So it's like you're kind of in this awful death trap, I would say if you're just solving that problem

Originality

10 / 20

A few genuinely contrarian observations - the seesaw shift from unstructured to structured as the hard AI problem, and the argument that only ~5 companies in the world should be training foundation models - give the episode some fresh texture, but the bulk of the conversation recycles well-worn startup mantras (iterate fast, easy to prototype hard to productionise) without pushing into first-principles territory.

the seesaw here between unstructured data and structured data has actually shifted recently
there's actually only like maybe four or five companies in the world. I would argue that should be training These models

Guest Caliber

13 / 20

Chris Aberger is a legitimate practitioner - Stanford database PhD, ML leadership at Sambanova, and founder-CEO of an acquired AI startup - with directly relevant technical depth on LLMs and structured data; however, Numbers Station was a small early-stage company and the episode's M&A framing limits the candour one would expect from an independent practitioner appearance.

worst case, optimal join processing. So really hardcore kind of database infrastructure pivoted more towards the AI side of the house
we had a concept that we called a knowledge layer. You can think about this as like a combination of like a knowledge graph and a, and a semantic layer

Specificity & Evidence

10 / 20

The episode offers scattered concrete detail - JLL as a named customer with a work-order use case, the 91 F1 score benchmark comparison, named acquisitions (ServiceNow/Data World, Salesforce/Informatica), and the 80% metadata time figure - but is entirely absent of revenue numbers, customer counts, growth metrics, or measurable outcomes, which are the specifics most useful to a B2B operator evaluating the claims.

one of our largest customers is in Commercial real estate jll they've been a fantastic customer to us. You can look at things like work orders over a property
if you talk to a data analyst and I tell them you're going to get a 91F1 score, they're like, what the hell did you just say?

Conversational Craft

5 / 20

The host is the CEO of the acquiring company interviewing his newest VP in what is functionally an M&A press release in audio form; there is no independent challenge, no pushback on any claim, and the host frequently delivers extended monologues or answers his own questions rather than drawing out the guest, making this a PR conversation rather than a substantive interview.

Yeah, I couldn't agree more.
Yeah, it, it, it absolutely does. I mean, even in the early days, you can just see that happening with the, the, with the pace of code that you guys are able to put out.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker D57%
  • Speaker C39%
  • Speaker A3%
  • Speaker B1%

Most-used words

data78metadata28structured25agents22build21number21first21problem21space21station19building19elation18level18back17different16problems16

Episode notes

The future of business intelligence is being rewritten. Have you ever wondered how AI will unlock the power of unstructured data? In this episode of Data Radicals , host Satyen Sangani is joined by Chris Aberger, newly-minted VP at Alation to discuss building AI-powered data workflows. A startup pioneer, Chris explores the importance of metadata in enhancing AI applications within organizations, the significance of quick iterations, and the evolving role of AI engineers. “ That two-step realization is what's causing a lot of this activity that we're seeing in the market, which is, I know I need to plug into databases.

Full transcript

47 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: Welcome back to Data Radicals. Today's episode is special. We're joined by Chris Aberger, the co founder and former CEO of Numbers Station AI and now Elation's newest vp. Following our recent acquisition of the company, this conversation reveals why we've joined forces and our shared vision for enterprise AI. Together, we're tackling a critical challenge, unlocking the potential of structured data with large language models. Chris shares the origin story of Numbers Station and the fast paced innovation that fueled its rise. He also reveals why metadata is key to AI and what that means for the powerful synergy that forms this partnership. If you're curious about where AI is headed and how it can make structured data more usable, intuitive and impactful, this is an episode for you.

Speaker B: This podcast is brought to you by Elation, a platform that delivers trusted data. AI creators know you can't have trusted AI without trusted data. Today, our customers use Elation to build game changing AI solutions that streamline productivity and improve the customer experience. Learn more about Elation at A L a t I o-n.com today on data

Speaker C: Radicals, I am thrilled to welcome Chris Aberger, newly minted VP at Elation. Though most recently Chris serves as the CEO and co founder at Number Station, a ah, startup pioneer for building AI agents for data workflows. He was also the Senior Director of Machine Learning at Samanova Systems where he built the AI team from the ground up. He has work experience at Google, Apple and IBM. And this is a super exciting time for Elation because we just acquired Number Station and we're joining forces to really make a new generation of structured data and AI and workflows and like literally every buzzword that we're going to actually deliver on and make real. Chris, welcome to Data Radicals.

Speaker D: Hey Satyan, thank you so much for having me on the show. Excited to talk to you today.

Speaker C: So maybe we'll just start with Number Station. You founded it along with a couple of other folks. Tell us about the story. Tell us about how it started and why you started it.

Speaker D: Yeah, so it was back in early 2021. I was at UH Sava at the time and I'd always kind of. So I met sun and Anas, my other co founders and Chris Ray back at Stanford when we were sun and Anas were doing their PhDs with me and Chris Ray is all of our advisors. He was a, uh, he's a professor at Stanford and so we all met back at Stanford and back in early 2021. This was pre ChatGPT and kind of all the AI and LLM hype that's come out since then. We saw this trend with foundation models and large language models coming and Senn had written and pioneered some early work on um, taking structured data and applying that to LLMs. And I read that paper, kind of circled back with those folks and was like, wow, there's really interesting opportunity here to build a layer that sits on top of the LLM. So we're not, you weren't training these models ourselves, but to build out this layer that sits on top of the LLMs that makes it easy for organizations and data organizations in particular to build out applications on top of their classic databases and structured data. So that was kind of the origin of. It was this paper that ah, Aness spearheaded back at Stanford when she was doing some, some research us kind of all sorts of circling back together after, after many years they prototyped some ideas around this and then eventually got enough conviction in the fall to say this is something that we have to do and took the leap of faith and decided to start the company at that point.

Speaker C: So from the point that you read the paper to the point of founding the company, how many months was that?

Speaker D: That was about six months, maybe eight months. Uh, I read it before it was released. I believe it was released in like May of 21. And we formally started the company in October, so it wasn't a time frame.

Speaker C: And then you went straight to basically raising your first seed round of capital from Chris and his, his, his crew.

Speaker D: Yeah, so we raised our, our first round of capital late in, in 2021 and we're kind of off to the races on, on building out the company since then. Yeah.

Speaker C: And then the first time we met, I can't remember, I think it was 2023. Right. Or was it, was it 22?

Speaker D: Um, I'm, um, I think it was 23. Yeah, it sounds right.

Speaker C: Yeah. So the, the story and the, and the fun connection there is that Venki Gunty, who was of course one of Elation's co founders, you know, journeyed through a couple of different places. His own startup, then Google and then most recently Number Station. And you know, now is, now is elsewhere. But Venki was your, what was your head of product. And he was like, sathya, I've got to show you this, I've got to show you this thing. And he, he, he did, he did at the time, um, which was, which was super cool and, and I guess then reintroduced us about. Oh My God, what, 90 days ago?

Speaker D: Yeah, not too long ago, I think a couple, couple Months ago. Yep.

Speaker C: So tell a be about the journey between sort of, I guess, you know, you're founding the company, getting your initial funding and what did you learn along the way of building Number Station and what was, what was the journey like? And you know, like all startup journeys and you're still on it on some level. What were the ups, what were the downs and what'd you learn?

Speaker D: Yeah, I mean like all startups, it's a roller coaster as the kind of first and foremost thing. So learned a ton of stuff uh, along the way. The vision for the company has always been the same, right? This foundational layer that sits on top of these foundation models and large language models and largely targeted for structured data. But a lot of the things that we did on top were like quick iterations and testing directions and market and where we think we can actually build a sustainable business. And so kind of the first approach that we took was actually looking at data transformations because it's like the most gnarly problem I would say, that's out there in terms of getting your data model correct is like step zero to do anything on top of your structured data. So we're like, you know, as academics, like oh, this is like the, the toughest technical problem that's out there. Like clearly we should just go solve this. And we really like dove in solve that. I think it was about half a year had our first product released and we were getting decent traction, but we were finding was it was really tough to turn that into actual dollars coming into the business. And it makes sense, right? As you kind of took a step back, it's like, oh, if you go out looked at and look at these organizations who actually controls the budget and what they want to get done, enabling their data engineers to be, you know, slightly better is a decent value proposition. But actually solving that executive purchasers and level problems is a much more attractive one. So we kind of up leveled after that first initial product iteration into actually solving that purchaser's end level problems. And where that started was actually with this kind of canonical chat with your data style application. So you know, think of like ChatGPT on top of a database and use that as a starting point application. But the idea is really to be extensible and broaden out into a variety of different workflows on top of of structured data and actually solving that end, um, business users problem or executive purchasers problems. And so that was, you know, it was a lot of like testing the market I think really iterating quickly. One of my My favorite quotes that I learned early on was like, when in doubt, iterate faster. And I'd say like the only regret I would say at Number Station or like the main regret I would say I had is like, I just wish we would have iterated faster as you go back and kind of look at the company, even though I do believe we iterated extremely quickly. So I'd say that's like the number one lesson lesson that I learned. And yeah, the data transformation stuff, we still use it, right. To get that data model in the, in the correct form for, for what we're doing on top of these, these databases. But there's a lot of evolutions of the product to get to the state that it's in now.

Speaker C: Yeah, and I mean this is one of the observations that Venki made about you as well. And he's like, this is. This team just moves so fast. Like, you know, and, and I remember the early elation founding team. One of my co founders was a guy who, you know, we all know his name is Feng Nu, who now, you know, runs his own company. And Fung was just like, you know, we, the three of us, I um, mean, we, we tell these funny stories where Aaron, who was one of the, who's the fourth co founder, and I would just debate things endlessly and you know, and I don't code so like, you know, what else could I do but like, you know, pontificate about crap? And so he, he would, he would basically, we would talk, we would argue and then FIG would be like, here, I've already hacked it up, it's done. And, and it felt like you guys had literally a team, uh, that just, just did that all day long. I mean it was not just, it was not just obviously you who, who had the capability to hack, but, but Inez and Sen and you know, all of the other folks on the team who, who had joined over time, which was just an incredible superpower. So you built this chat with your data interface, but your chat with your data was different. Like we've actually seen, you know, in our travels, gotten pitched by a lot of chat with your data companies. Everybody wants to do chat with SQL. You know, there are lots of papers out there talking about the efficacy of that. You guys took a slightly different tack though, because the chat was an interface, but it was an interface as a window to not just a sort of SQL authoring agent or a querying agent, but you had other agents tell us about that path. How'd you get to that discovery and what What'd you do to get there?

Speaker D: Yeah, so I think I've always had this like love hate relationship is probably the best way to put it with, with chat with your data. I think one of the, the things with chat with your data is there's a ton of noise out there in the market. It's super easy to rip up a text to SQL prototype. In fact you don't even need to rip anything up. You can just talk to OpenAI and it'll spit out SQL for you. So there's not a huge technical barrier in getting like a proof of concept running. But I think it's largely still an unsolved problem that everyone thinks is solved. So it's like you're kind of in this awful death trap, I would say if you're just solving that problem from, from my perspective. And the second thing is I don't think it's kind of, I don't think Chad is where things end in this space. Right. Especially you uh, know I think this has been validated right in the era of agents. But our idea was always to like get something done. So having a conversation is nice, finding an insight is nice, but that's really always been a stepping stone for us in terms of building out other agents that uh, enable people to capture full workflows and actually get stuff done on their data. So I think it's a very easy, digestible application kind of first step that you can take on top of the database. That's why we started there. Like everyone understands it. There's a need in most organizations to kind of lower the barrier to entry to get insights on top of your data sets. So all that's good, it's an entry point into these organizations. It's a way to establish trust. It's a way to also train our AI systems to become more acquainted with this organization. But every company that signed up to work with us did not sign up just for that chat capability. Even though that's where they started. They signed up for the longer term path to solve and level workflows. Right. Whether that's in commercial real estate, fitness industry, whatever it might be. They actually wanted to get things done on their data, not just serve up an insight or explore what's going on over their data. And so that's always been the hypothesis. I think it really got validated when all the agent hype kind of picked up because we were a little bit ahead in that we had an agentic platform from the start. And then I kind of validated that uh, okay, it's clear how all this stuff becomes immediately extensible and how you can build really interesting systems on top of structured data. So that was kind of the evolution of the business. I think a lot of people that look at us at a high level were like, oh, you're just doing text to SQL. Which often was our motion to go work with companies. But really what we were doing from a foundational level was actually much different than just that application.

Speaker C: And uh, talk about some of those differences. Talk about sort of what else you built besides sort of this, you know, text to SQL capability.

Speaker D: Yeah, I think even just so let's just like zone in on um, the text to SQL and how we're different there. So like the first realization, even if you're doing chat with your data that we had and I think what a lot of companies were taking in this space was that oh, you do everything kind of net new like you talk to your new, whatever chatbot that you're deploying in your organization. This will emit all the. NET new SQL and like you're good to go. Go throw away your bi tools that you were using in the past and use this new tool. That was not our stance. In fact like we looked at who is having a lot of success in the AI space. We saw glean and early on we were saying we're glean for structured data.

Speaker C: Right.

Speaker D: So we want to plug in and actually operate across a bunch of different systems. If the answer already exists in tableau, power bi looker, uh, documentation, wherever it might be across your organization, we want to actually pull from that, not do something net new in the database. Don't reinvent the wheel if you don't have to. But, but the first thing there is that kind of connectivity into a bunch of other systems. And the way that you do that connectivity is through a multi agent framework. So we always had that kind of architectural thing set up. That's still we're talking about kind of chatting and getting insights from your database still, even though we're talking about plugging into a bunch of systems. But what that set us up for right is the ability to eventually go and take actions in those systems too. Right. So you can think about plugging into the simplest example here would be like producing a PowerPoint slide. Right. So I have an insight, but now I need to serve up that insight to my boss. I'm m going to automatically produce a PowerPoint slide. Something more complicated is I'm going to go look for, you know, one of our largest customers is in Commercial real estate jll they've been a fantastic customer to us. You can look at things like work orders over a property and say like your air conditioner is breaking down or there's a heat wave. I want to proactively go find vendors to do maintenance on my air conditioning unit ahead of this heat wave. Go search in my area, find these vendors, send me an email with the list, maybe even go out and contact them directly. But that's actually going into the end state of taking actions, not just surfacing, you know, pretty bar chart that's giving me an insight over what's going on in my data. And that's where I believe the space is going to more and more tilt over time.

Speaker C: Yeah, I couldn't agree more. I mean I think we were late to have relative to you and maybe to others, but probably not the world to that realization. I mean the way you think about the last decade of building data tools, you get this proliferation of productivity based capabilities. Whether that is BI tools or the ability to build ETL better or the ability to be able to build ETL in different ways and different modalities or even the databases themselves. And then of course in our space you've got the catalogs and the governance capabilities, but all of it has been off to the side. Like it's been largely stuff that makes other people productive. The end is a report. The insight then feeds a decision and people are doing these things sort of in, in, in two different planes. You're making a decision and then you're doing something. And I think the, you know, the two things that, that seem to be happening today or that are happening today is that you know, data teams are under this assault. There's like, look, the secular investment in all of these tools is, has better proven some ROI or you're not going to get funded. So there's this really clear budget motivation. But on the flip side there's also this issue which is that uh, Gen is eating everything. And what I think to your point this means, and you said this so eloquently, is that these data teams have to become builders. These data teams have to actually go from you know, just intellectually solving the problem to actually solving the problem. And you can view that as a threat. But I think you and I both see this as an incredible opportunity.

Speaker D: Yeah, I mean like, I think the, the advent of like chatgpt and all these gen AI tools has been crazy. Like you know, I even go talk to my mom and she's doing like you know, things on the computer that she couldn't dream of before. Whether it's like editing images and like, it's very basic things to be clear. But my mom, you know, believes with these tools like ChatGPT, she's turning into more of a, you know, know, quote unquote builder. Sorry mom, don't mean to offend you. Using these tools. And we see that kind of bug being bitten in all the data teams that we're talking to, where they now have this belief that like, hey, I can actually go out and build a fair amount of this myself using these tools or on top of something like OpenAI. But what they find when they're doing this very quickly is that it's easy to get a prototype up, it's very hard to get it into production. So, uh, I like one of the quotes that I always like in the space is it's very easy to get started, it's tough to get right. And so oftentimes when we were talking to teams at Number Station, we wanted to go after these builders because again, we have this platform that we want to make it easy for people to build on top of and going after these teams that we're building. Often, many times they've tried it and then realized that it was really hard and that they needed to partner with someone to kind of get that last mile and to get it into production. And the interesting part about that kind of partnering and getting into production is we would build these agents out on top and we have a bunch of ML talents. We actually are able to rip that up fairly quickly, these different agents on top. But where the real kind of moat and difficulty came in the space, there's actually two parts of it was on the metadata side. So we had a concept that we called a knowledge layer. You can think about this as like a combination of like a knowledge graph and a, and a semantic layer. It's actually the same thing as a lation's data product, which is a very interesting mashup that we kind of saw early on when we were talking here. But almost like I would say 80% of our time at Number Station was not spent building those agents, it was spent getting that metadata correct, right? Working in these organizations, going in, trying to get access to query logs, trying to figure out how to mine descriptions of columns, trying to figure out how to, you know, everything in between, look at documentation, et cetera, how do we get this metadata into a reasonable state? Because that was actually the key thing to making these agents work well in, in production. And so I think that Was like, you know, one of the things that clicked for me, you know, I think even in our first meeting in 23 or 24 when, when Venki was at our organization, I was, I remember having these conversations with him. I'm like, dude, Elation is sitting on a gold mine, right, for all the metadata that I have in their organization. Because we are like, you know, crying tears and sweating a ton of bullets to go into these organizations and curate this metadata. You guys already have it in a lot of cases. And so that was always something that was really interesting to me about partnering with Elation was the fact that we could almost like jumpstart or take a shortcut, right, in working in these organizations to access all this gold mine of metadata and which is actually the key thing you need to make these agents work well on top. So long winded answer there. But that's kind of been our tour of building out agents from chat all the way to workflows and getting things done for all the applications. What it came down to was metadata in the end.

Speaker C: Yeah, I think that's absolutely right. And I would also though say that metadata in and of itself has been, you know, the challenge, um, of on some level, bi and databases, you, uh, know, probably for the last 30 or 40 years. I mean, I remember when I first started working at Oracle, my mentor at the time was building these like metadata driven apps to do, you know, these deterministic, these deterministic calculations on data. And in his case he was like, it's all about the metadata. And this was like back in 2006. And you know, people have been talking about that since the days of like business Object and everybody's been trying to get the metadata. Right. I think what's exciting about what you guys do is that the cycle time, so it goes back to this entire premise of speed, which is that the cycle time of improving the metadata, getting to an outcome, being able to do something, seeing the impact of that thing, going back and correcting the metadata, so I can even be more powerful to do. The next thing that I want to do is I think the loop that you have to create. And to me that's what I think is so exciting about this combination, which is that we got this metadata, but it's off to the side and you've got these applications and these agents that actually give me the power to do something with it. I can go build a presentation, I can go write a query, I can go do these things that like uh, give me so much more power and capability. Than I otherwise had. And I can do them right. And, and I think that dual benefit is just, I mean I, I remember when I first met you, uh, the second time around or a second, you know, like 90 days ago when we first talked.

Speaker D: Second. First time.

Speaker C: The second first time.

Speaker A: Yeah.

Speaker C: It felt like. I think you were. It's funny because as I reflect back to the first meeting I, ah, this is really interesting. I think there's something here, but like I clearly was like not totally getting it as quickly as I ought to have. But it did feel like the second time around we were just completing each other's sentences and I think largely because we got to the same place through different paths.

Speaker D: Yeah, ah, I mean I probably didn't get it completely the first time we talked to you, so it's probably both of us kind of converging on similar conclusions here. But yeah, to your point, like you know, solving that end level use case and then having that feedback loop back in to correct the metadata in terms of how these agents are working and being organized underneath the scenes. Absolutely crucial, right? And I would say the kind of weird thing that we saw, just to go on another tangent for what's happened at Number Station was we actually had some, we started the company, I Forget, with like 5 or 6ml or AI PhDs, ton of firepower in the AI space. And what we had spent our PhDs doing was actually training models, right? So you go collect the training set, you go train a bunch of floating point numbers effectively underneath the scenes, watch some loss curves and make sure that everything's performing well. But the space completely shifted, right? Where for AI talent and what you're doing for the most part in 2021 when these large language models came out and that. Well, nowadays there's actually only like maybe four or five companies in the world. I would argue that should be training These models, open AIs of the world, et cetera. But the AI or ML talent at the remaining companies actually should be spending their time working on these feedback loops. It's almost like a form of reinforcement learning on this metadata, but it's how do you work on these feedback loops? So all this is a long winded way of saying there's actually been a transition for AI and ML talent outside of kind of those, those model providers in terms of what they should be doing in organizations and that it's largely a lot of prompt engineering and then building these feedback and evaluation loops into these agents themselves. Which is a tilt right from kind of what we were doing in our in our PhDs 10 years ago. So it's like, I almost call this like, uh, the identity crisis of AI engineers is like they feel like because they spent their PhDs training these models, they still need to be training these models. And it's like, no, no, no. To go solve these end level applications. Actually in most cases you shouldn't be touching the weights of these models. What you got to be getting right is this metadata and evaluation feedback loops. And if you can nail that, that's how you have success and solve N level applications in the AI space. So super interesting problem, ton of interesting things to solve there. And it really has been a huge focus for the majority of number station in terms of what we've been looking at.

Speaker C: Yeah. And this is like one of the things that I think we've, you know, I think seen in greater resolution with you guys. Like there's so many variables to any given problem. So like, you know, there's the model and which model you choose to use and there's some that are obviously more powerful but more expensive and cheaper and smaller. And you know, there's, there's open and closed and whatever. There's that entire thing and then there's this entire question of like, well, okay, great, but what's the data that I'm working with? And you know, what, what is the interaction between that model and this data and how much metadata do I need and actually to make these things actually allow these things to talk to each other. And, and, but then there's also the use case which is like, what am I actually trying to do? Because if I'm just trying to answer a very simple question, it's one thing, if I'm trying to do something very complicated, I might need a whole different level of metadata. And to your point, the metadata is the thing that actually sort of does all the translation. But it's a, it's not a fixed outcome that you can deterministically know. It's literally always this thing that you have to find and evolve and test. And so it's a lot more like software testing and evolution than it is like, you know, sort of just like data modeling where, you know, in the old world of data modeling we'd basically go say, oh, you have a star schema. Great. It's answers, uh, some set of questions. Now you're going to like live that, have that thing live for a couple of years.

Speaker D: It's a living organism.

Speaker C: It's a living organism. And so the real question is, how do you set up that? How do you set up that. That feedback loop that you're pointing to? We have a guy here at Elation who literally every paragraph he's going to say the word fly, Flywheel. But how do you set up that flywheel? Shout out to Jonathan Bruce. But like, yeah, no, it's. It's a funny. It's a funny. It's a funny thing. And I think. I think building these apps is what I'm super, uh, I think so excited about, because I feel like it'll. The power is going to be really incredible. Although it's going to be hard. I mean, I think there's going to be lots of hard technical challenges and lots of hard user interface challenges too. Now, you know, I guess everybody in the world seems to see this at the same time because right before we announced, or actually before or after, before we announced the acquisition of number station, ServiceNow was gonna announce the acquisition of Data World, or Data World. And then I guess a week after Salesforce announced its acquisition of Informatica. Talk a little bit about that. Like, so, you know, we're in this moment where, where metadata is super interesting. You know, we're going up the stack, though. Those guys are coming down the stack. I don't know that I even say that we'd compete against each other, but certainly it's a lot in the same domain and space. How do you think about that?

Speaker D: Yeah, I think it's interesting. I think a lot of people are coming to interesting realizations here. Uh, one thing just to take a step back on how I view this space. So actually if you look like four or five years ago, again, um, pre these LLMs, the difficult and sexy problems m where you saw a lot of big acquisitions getting done, I think even Fung had one in this space was around unstructured data, right? So it's like, how do I take this, like, web of dark unstructured data and turn this into something usable? And it was like a huge problem, right, in the AI and ML space and the enterprise data problem. And it still is, to be clear, not claiming this is a solved problem, but I do think the, like, seesaw here between unstructured data and structured data has actually shifted recently. And why I say that is if you look at how these LLMs are trained, they're actually, you go and take like all the unstructured data from the web and train over that. And so these LLMs are actually well suited for unstructured data. Do they work perfectly out of the box for unstructured data? No, but like they are much better suited for unstructured data than they are for structured data. And I think over time more and more of the unstructured data like PDF parsing and all the unstructured data that you might have in an organization is going to be commoditized by these LLM or model providers. But what I do think people are realizing lately, or what's kind of become the cool problem to look at, or people at least realize it's hard, is making the structured data useful. These LLMs are not trained on structured data. So like again, I can rip up a prototype on OpenAI very quickly to show you talking to my database or doing something on top of my database. But when you go into a Fortune 100 company with really complex enterprise structured data, like, good luck trying to get that to work well by just using a model plugging into that. And so again, uh, I think there's like a couple things that have gone on here. First is people have realized that, okay, like structured data is actually like the hard problem to get right and all these organizations really valuable data is inside their databases in the structured formats. We have to figure out how to make this ready for the AI era. And then the kind of second level problem that people are discovering is how do I make this structured data actually work? Oh, it's metadata.

Speaker C: Right.

Speaker D: And I think that realization, that kind of two step realization is what's causing a lot of this activity that we're seeing in the market, which is, I know I need to plug into databases. I'm um, now coming to terms with the fact that this is actually a really tough problem to get right. In order to get it right, I need to effectively go build a data catalog or metadata provider. And therefore you're seeing a lot of activity in this space, at least from, from my biased perspective.

Speaker C: Yeah, I mean Inez and sort of our, our, our. And David Chow, who runs marketing for us, you know, called it this sort of precision agentic workflows, which, I mean, you know, two of those words are used by literally every vendor to talk about every single thing, agentic and workflows. Because, you know, why not? They do think this precision thing is the entire fundamental point. Like you've got these like lossy, stochastic, you know, LLMs that are prone to hallucinations. Then you've got the structured data that has to be perfect and precise and you can't be like, oh, it might be a, ah, debit or a credit or it might be like, you know, positive or negative or it might Be like, you know, a customer or a vendor like you. You have to know exactly which one it is. And, and that's, I think, the opportunity that we have in front of us, which is to make these models talk to this structured data, both both on reading and on writing. And, and it's what I think is really fun is to watch the progress that, like the speed of the progress that we're making together. I mean, I feel like we were moving at a pace, you know, for the last six to seven months that were, was fast. You guys were obviously doing so as well. And I feel like we've almost weirdly accelerated each other, which you almost never see in these acquisitions. And I don't know if that's the ethos of the team, the tools of the moment or the people, but it's really cool.

Speaker D: Yeah, I think it's an obvious like one plus one equals three situation. So I think it's awesome on both sides, but just to double click into what you're talking about on the precision side. So like, you know, I kind of set up this unstructured versus structured data. There's this, you hit this very well and eloquently. So I should have said this, but like, there's a user expectation difference between these two kind of worlds as well. So like on the unstructured side, you're typically starting from nothing. Or like a typical search experience, like getting a 91F1 score is acceptable. If I go talk to a data analyst and I tell them you're going to get a 91F1 score, they're like, what the hell did you just say? Right? Because there have been systems on top of these databases since Cobb in the 80s. And the answer is the answer, right? So that is the expectation on the structured side of the world. So there's this user expectation side where like, you have to get it right on the structured data side. And that's why the bar is higher and why it's uh, a harder problem to solve going along that line of kind of that precision agentic thing and why it's so important. For the problem that we're looking at is also just kind of the user expectation and experience on this side of the house.

Speaker B: Yeah.

Speaker C: And it's funny because I talk to a lot of people who are like, look, this is, you know, man, this game's going to be over. Everybody's going to do this work and like in three to six months it's going to be the, you know, the unity of like, all problems being solved because Somebody's going to win this race. And you know, I, I sort of, look, I obviously think that scale has its own beauty and a lot of these players have a lot of, you know, interesting things to contribute. I also think there's lots of hard problems and complicated problems to go solve. And I think, I think the, so we were, you know, at, at Columbia there's a professor, ah, who I met, guy named Eugene Wu. And he's basically starting what he, he terms as sort of an agent lab which he's trying to compare to the AMP lab at Berkeley. And you know, he talks about, the way he talks about this lab is he's like, look, you know, in the early days of databases in the 80s, people were talking and you mentioned Cobb, but like there's all of these problems to get to true asset compliance and it took years, decades to get to that, to the moment where databases could be truly relied upon. And he's like, we're going through the same stuff. There are same fundamental research problems in agents that need to be worked through or with agents that need to be worked through and with models that need to be worked through to get to a level of reliability. And I'm on the side of like, yes, the future will be really exciting, but it's going to be jagged and it's going to be unexpected and we, you know, I think what's fun for us as builders is they're the cool problems to solve. Like, they're just great stuff to go build and figure out.

Speaker D: Yeah, I think there's like infinitely large number of cool problems to go solve. There's also the same amount of hype too. So I think having like realistic expectations in terms of kind of taking stepping stones in terms of how you're going to get to fully automated agentic solutions again. I think what we saw even throughout the duration of of Number Station was like buyers getting super burnt out on just like hype and demos that like anyone can rip up and like what's real here, what's not. But being really pragmatic in terms of how you can go into organizations and incrementally provide value kind of along this trajectory towards this end level vision that we all believe in and realize right now, I think is essential and crucial when working with companies.

Speaker C: Yeah, I, I do feel like that too. And I know there's like, you know, there's a lot of pressure from the world of customers and employees and investors and all these people who are like, let's get out and make these sensational claims and Just say, you know, AI is going to, I don't know, like, what is it? Like, Dario, everybody is like, oh, like, you know, 25% of all white collar jobs are going to go away in two years. Like, okay, I guess, maybe, but like, maybe. And, um, like, who knows? Like, the world is like, the future's hard to predict. But I do feel like, I mean, the other thing that I loved about you guys is that you just stood for a very pragmatic and very authentic and very clear and purposeful brand of, like, look right in front of us. There are problems to go solve and we're going to go solve those problems and we're going to help customers right now and we're going to make claims based upon what we can do today as opposed to these sensational things. And I, and I do think that was maybe one of the hidden but really cool values that I think both companies share. And so that's, uh, one of the things that I really admired about you. And as. And Sen as well. As you forecast forward, what are you excited about? Where, where do you, where do you see the world going? What things do you. I mean, we've obviously talked about some of the problems, but like now, you know, forecasting forward, what do you think is going to be things that people are not paying attention to that they ought to be things that you're excited about, people that people should be thinking about?

Speaker D: Yeah. So I think there's kind of two angles of things that I, you know, I'm a technologist at heart, so I always focus on these technical problems that I think are really interesting, but I think they tie back to huge enterprise value and downstream value for customers. So one is this whole curation and feedback loop around the metadata itself. And there's a ton of innovation to be had in this space in terms of just automatically going in and curating and producing this information, as well as having these really tight feedback loops for a bunch of different agents on top to go in and kind of maintain this living organism over time. As I said, it's a problem that we focused on throughout the duration of Number Station. But I would not claim as a fully solved problem, there's a ton of technical innovation to still go on in this space. And I think the companies that really nail that loop are the ones that are going to have the most success and the end level applications that they're doing on top. So I think there's a ton of technical innovation to be done and had on that side that I'm really interested about and then the second part is actually those applications on top. So, you know, I think we're just scratching the surface of what you can do with these agents on top. And you know, one of the things here is having this web of interconnectivity across a bunch of different tools. And then the second thing is actually having organizations not be bottlenecked by the vendor. And so what I mean by that is like early on in Number Station we were going, just working, building out these first versions of agents. It was a lot of us building it out, right. To go solve these end level use cases. What I've been excited about in the past six months is I'm seeing our customers start to come up to speed and build their own agents on top of our platform. And that's when you're really starting to cook with gas from my perspective. Because you're not the domain expert should be the one going to build these applications, not the vendor. Right. And so I think giving that power and kind of unlocking this power with organizations where you can go build your own agentic applications really quickly and having that platform that powers that sitting on top of the metadata and structured data underneath is something I'm super excited about and excited to see all the different types of things that customers will build up on top of our combined platform.

Speaker C: Yeah, I mean that's one of the things that has struck me about Number Station and about just generally how we're seeing AI use evolve, which is that the entrepreneurial, curious, high power, high agency, high motor individuals are people who now have sort of superpowers. I mean they just have this ability to take that knowledge and do, you know, Everybody talks about 10x but the uh, ability to do, you know, at least an order of magnitude more than they, what they otherwise would have done and it lowers the barrier to being that person. It sort of allows you to, if you, if you have an idea and you have initiative, you can go do all this stuff. It is a little, you know what I think is also a, you know, the counterbalance to that is it's a little scary. Like if you're just kind of somebody who's like, yeah, uh, you know, I just want to show up and like, you know, ask random questions and not really do a lot of work and not think really, really hard. I think those are probably the areas where I think, you know, you're going to be put under pressure because these tools will force you to the limit of what you can, what you can do. Um, and you know, recently it has Been I think the case that software actually really catered to the lowest common denominator. And interestingly, these tools are now creating to like, almost the best in class, which is, is really fun. So you are, you're based in Seattle. It tells us a little bit about you. Tell us a little bit about the culture of the team that you'd like to build and the cultures that you think are winning in the customers that you're seeing.

Speaker D: Yeah. So about myself, like a little bit about my background. So went to stanford for my PhD. As I mentioned before. I actually originally was like super into the hardware side of the house. So it's like computer architecture as an undergrad. That's what I wanted to do for my PhD. So initially started working with Khutin as the father of multicore. And you know, I remember sitting outside, I basically like stalked Kunle to work with him, but I was like sitting outside of his office waiting for him and I was like, I want to work with you. I'll do anything to work with you, basically. And he was like, great. But I don't want people building, I don't want people working on hardware. You need to work on software. I was like, okay, I guess I'm working on software. So I started working on software throughout my PhD. Eventually met Chris Ray through a class he was teaching. He came to Stanford as well and really focused on databases for my PhD. So I don't know why you would go look up my PhD thesis, but if you were to go look that up, is on, um, worst case, optimal joint processing. So really hardcore kind of database infrastructure pivoted more towards the AI side of the house. After I'd finished my thesis. Eventually went to Samova Systems where they were building hardware for AI. Like, great combination of some of the early things I was interested in. And then Number Station, of course, kind of combining databases with AI has, has been my, my trajectory. So in terms of kind of culture and building out companies, I think the, the number one thing that I look for in people is the ability to learn quickly. Right. So, like it almost to a certain extent. It sounds weird because I actually hired a bunch of really credentialed people. Both companies I've been at, you know, PhDs, really smart people from, from top labs. But I think one of the qualities of, uh, PhD students that's well suited for, for startups is you're a little bit fearless in terms of you're really, you're used to these quick iterations. You're used to failing right when you're doing, you're doing research and you're interested, you're continually interested in learning. Right. You're not set in your ways in terms of how you're adapting to the market landscape, et cetera. So that kind of ability to learn and as I talked about earlier, iterate fast is kind of the number one thing that I look for from a cultural perspective as well as, you know, of course, just being a good person to work with. No one wants to work with someone that. That's extremely difficult. And so just finding it could be an AI PhD, it could be someone straight out of undergrad, it could be someone later in their career. It doesn't really matter. But that kind of hunger to learn and adapt and fail and iterate fast is the number one thing that I've always looked for when I'm going out and hiring. And I think our culture hopefully embodies that to, to a large extent.

Speaker C: Yeah, it, it, it absolutely does. I mean, even in the early days, you can just see that happening with the, the, with the pace of code that you guys are able to put out. And I mean, you know, you know, candidly, it's funny. I mean, you know, even revealing a little bit of our own security and insecurity and some of the, some of the things that we were thinking about and we were talking about, like, hey, these guys are going to come on board and wow, are we going to be able to move as fast? And I think it's been both fun to actually have the team move as fast, but also really inspiring to watch how quickly you guys move. And I think it just, you know, it, it shows you that there's always another level to go reach.

Speaker D: Yeah, I think everyone's moving fast across all angles from, from what I've seen. And it's what's needed in this market. Right? Like, I mean, the space is moving so fast and we're going to get things wrong right. It just matters that you get them wrong quickly. I think one of the, like, another, just another quote that like, popped in my head as we were talking about this. I remember early on at, uh, Sam Benova, and it's just like, Chris Ray has been my mentor for quite a while. I was like, struggling between a big decision that I had to make and, um, like, so concerned about getting it right. I was like, I have to get this right. Like, life or death, turns out, didn't matter. But, you know, I was like, coming to him and I was like, so distraught about it. And he's like, dude, you spent A week thinking about this. The worst thing you can do is not make a decision. It's like just make a decision. Who cares if it's wrong? Like just go change it after the fact. It's not like to use Bezos's analogy, like a, a one way versus two way door. And like that is also something that I've tried to embody. It's like it's fine to make wrong decisions, just like admit it and quickly course correct. Right. As you're kind of moving through this and especially in this space, like the world is shifting every month, we're going to get things wrong. Just got to move fast.

Speaker C: Yeah. And I think that's an interesting modality for customers because I think customers obviously want a little bit of a roadmap and they want predictability. And so I think one of the things that it also means is that you got to, with both team and customers, be really transparent about like what have you figured out and what haven't you figured out and what's open and what's not. And you know, really have to be pretty high of your expectations for the level of sort of attention that people are paying, but the maturity that they come at the problem space with because everybody's just running super fast. And that means you're going to run into dead ends as much as you're going to run into, to some, some success.

Speaker D: Yeah, I mean, I think you should isolate your customers from that process as much as possible. Right. Like for the things that you've solved, that's what you go like, you know, roll out to customers and then the things you're iterating on hopefully is shielded right. From the customers.

Speaker C: I think that's right. Although I think I think that's right. But like, you know, one VC early on told me, he's like, look, there's things in any endeavor that you've proven, there's the things you know, you can get to and then there's like, you know, and then there's the dream. Uh, and, and I do think that everybody wants the dream. And I think you have to be pretty authentic about the first thing and then pretty clear about how you get to the second thing. And I, and I, but I think there's a lot of people that are trying to do great work that just want to know and, and, and so yeah, it's pretty cool. So, so I, I, I, I hear that you're interested in moving from engineering to, to becoming a lawyer here at Elation. Tell us a little bit about that transition Because I know you love the, you love the law discipline so, so much.

Speaker D: Actually, funny enough. So if you do want to know, I did apply to law school out of undergrad and did get in actually, with like, scholarships to some, some pretty good universities. But I. So Satyan's making a joke about the acquisition process, which I think any CEO or founder who's been through an M and A knows. It's, uh, a lot of fun times with legal teams. But actually tying into my background, that actually was originally why I did engineering. I thought I would be a patent lawyer. And I did an internship, I think it was like, at IBM or Apple. I'd have to remember the particular summer. And I was like, whoa, this is way better than being a lawyer, but I'm going to go be an engineer here. I decided to stick to engineering, so I think I'll stick to engineering for now, but maybe we'll. We'll keep that option open for a later date.

Speaker C: Yeah, it was. I mean, especially inside baseball there.

Speaker D: The.

Speaker C: The, like, I think we both got you. You get the. The lawyers put you in this place, and not because that's what they're paid to do, but they put you in this place where you have to think about every worst case, like, corner scenario. And then you're just like. You m. Put yourself in these circles where you're like, God, this is like the most horrible thing that's going to ever happen. Even though what you're doing is something really great. And I think both of us, it's probably not. It's not the space that we both want to operate in.

Speaker D: Yeah, it's, you know, fun. Fun times in the M and A process, but worried about what happens if we get hit by an asteroid tomorrow, which we, you know, all have bigger problems in that case anyways, so.

Speaker C: For sure. Chris, welcome to Elation. Thank you for joining Data Radicals. It's been awesome to talk to you here and obviously to work with you and, uh, yeah. So excited for what we're going to do. It's going to be amazing.

Speaker D: Yeah, I'm super excited about it as well. And thanks for having me on the show.

Speaker A: That was a fantastic conversation with Chris. What stood out to me most is how clearly Chris grasps the real challenge in enterprise AI. It's not just building the agents, it's building the foundation those agents rely on. And that foundation is metadata. That's what makes AI work in the real world. And that's why this partnership makes so much sense. Elation has the goldmine of metadata numbers. Station has the agentic apps and platform to act on it. As Chris put it, we're just scratching the surface of what teams can build when they're empowered to create their own AI agents on top of structured data with feedback loops that constantly get better. That's the a faster, more intelligent, more actionable future for enterprise data. And we're building it together.

Speaker C: Together. I'm Satyan Sanghani, CEO of Elation.

Speaker A: Thanks for tuning in to Data Radicals. Stay curious, stay bold. See you next time.

Speaker B: This podcast is brought to you by Elation. Your boss may be AI ready, but is your data. Learn how to prepare your data for a range of AI use cases. This whitepaper will show you how to build an AI success strategy and avoid common pitfalls. Visit Elation.com AI Ready that's Elation.com AI Ready.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Radiology Can't Keep Up. Here's Where AI Actually Helps | Dr. Nina KottlerRethink Imaging · on Large language models96 / 100
  • Why AI Pilots Fail: How to Escape AI Pilot Purgatory and Scale Enterprise AI with Ronnie Kwesi ColemanUsing AI at Work · on AI agents92 / 100
  • More Agents Than Employees: How Zapier Disrupted Itself Before AI CouldTalking AI · on AI agents90 / 100
  • PMs Are Fintech: How Column is Rethinking Property Management BankingThe Profitable Property Management Podcast · on AI agents89 / 100
  • Inside an AI-powered personal dashboard | Matthew TiemannPossible · on AI agents83 / 100
  • 90% of AI prototypes never reach production (w/ Temporal's Samar Abbas) | AI BasicsThis Week in Startups · on AI agents81 / 100

More from Data Radicals

All episodes →
  • Perfume, Power, Prediction: Inside a Luxury Giant's Data and AI Strategy with Julie De Moyer, Chief Data Officer of LVMH Beauty
  • The Rise of AI: Voices from the Frontlines
  • Delegate to Innovate: How Letting Go Makes You a Better Leader with Todd James, Founder & CEO of Aurora Insights
  • Data Products for Dummies with Sanjeev Mohan, Principal at SanjMo
  • GTM is a Data Management Problem - How AI (& Better Data) Can Fix It with Copy.ai’s CEO, Paul Yacoubian
Explore the best B2B AI & Data podcasts →
All Data Radicals episodes →