
Data Unchained · 2026-05-29 · 40 min
Key moments - from our scoring
Substance score
42 / 100
Five dimensions, 20 points each
Pradnesh Patil founded Ultima AI to tackle a persistent problem he observed across his roles at Symantec, VMware, Cisco, and Palo Alto Networks: data teams are drowning in routine, complex work while struggling to hire talented engineers. Ultima AI builds an independent AI layer that automates three critical areas: data pipeline building and maintenance (using integration with GitHub Copilot, Cloud Code, and CLI/UI), data infrastructure management and optimization (through ambient agents that continuously tune warehouse configurations and query performance across Databricks, Snowflake, and BigQuery), and data migration automation. Rather than being vendor-locked like Snowflake's Code or Databricks' Genie, Ultima AI operates across hybrid ecosystems, lets customers bring their own LLMs, and maintains governance through context graphs and agentic harness architecture. Patil emphasizes how the company addresses enterprise AI pain points: cost containment via model gateways and context compaction, data sovereignty through on-premise and open-source options, and the ability to deliver world-class results by feeding LLMs rich internal context rather than raw data.
Data teams lack sufficient engineers to handle the immense workload of building hundreds or thousands of data pipelines, managing complex infrastructure across multiple platforms, optimizing queries, and executing multi-year data migrations - all while maintaining governance and meeting business deadlines.
Ultima AI operates vendor-agnostic across hybrid data ecosystems, allows customers to bring their own LLMs (on-premise or open-source), and avoids cost markups and biased recommendations inherent to walled-garden solutions locked into a single vendor's tools.
The platform uses a model gateway to route tasks to the right LLM (eliminating overuse of expensive models), employs context compaction to reduce token consumption, and automates infrastructure tuning so data warehouses run at 90-100% utilization rather than being manually optimized.
Yes - the AI agents handle 90-95% of migration work by understanding architecture differences, SQL dialect translation, schema mapping, and data validation across source and target systems, eliminating the years-long, manual validation-debugging cycles that plague traditional migrations.
It's a specialized architecture combining context graphs (knowledge from internal environments), governance layers, tools and skills (MCP servers), infrastructure components, and agents to deliver high-quality data work at lower AI cost while maintaining compliance and data sovereignty.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains a handful of genuinely useful concepts - context compaction, ambient infrastructure agents, and the five-component agentic harness - but the majority of airtime is product pitch and high-level AI generalities. Non-obvious claims are sparse relative to the 40-minute runtime.
some of these agents are so advanced they're completely ambient, can take hundreds of actions within an hour
context compaction. A lot of time when you are trying to do a data work for a few minutes, the compaction happens. You must have noticed this in cloud, et cetera, tools. A lot of times if the LLM itself, uh, isn't aware of what data I need to preserve in that compaction
The core arguments - vendor lock-in is bad, open-source LLMs reduce cost, AI can automate boring data work - are well-worn positions in the data/AI space. The 'agentic harness' framing adds modest structure but is not a genuinely fresh framework; no contrarian or first-principles claims appear.
That AI layer itself is bias. That's why I think using some independent AI, uh, layer to get the data work done, which talks with all of your tools
locking your sort of bets in. Sort of um, one vendor which is completely locked um, is not a good idea at this time
Pradnesh Patil is a legitimate operator-founder with senior product and engineering roles at Symantec, VMware, Cisco, and Palo Alto Networks, which grounds his domain credibility. However, Ultima AI is an early-stage startup with limited verifiable scale, and his claims rely largely on self-reported metrics rather than third-party validation.
worked for a bunch of big B2B enterprise companies. So I did engineering initially for Symantec, long back the makers of Norton Antivirus, and then evolved through my career, started doing product management, product marketing, eventually leadership roles for companies like VMware, Cisco, Palo Alto Networks
Our product gets used in 200 plus countries. We raised a couple of rounds of funding
The episode name-drops real platforms (Databricks, Snowflake, DBT, Airflow, GitHub Copilot) and cites two named benchmarks (ADE by dbt Labs, DAB by Apache Spark community) with a number-one ranking claim, plus 10 million free tokens and the 90-95% migration automation figure. However, no customer names, no actual dollar-savings figures from real deployments, and no timeline data are offered, leaving major claims unsubstantiated.
right now we're trending number one on that benchmark
we are giving away 10 million free tokens which should be good enough at least depending on at what level you use. Hopefully at least for a month
The host structures the conversation reasonably across pipeline management, cost, migration, and deployment, but consistently validates rather than challenges. The 90-95% automation claim, the number-one benchmark assertion, and vague customer scale figures all pass without any probe for evidence or counter-scenario.
I have my suspicions, but probably better just to hear um, directly from you
So that seems like kind of a utopia
Computed from the transcript - who did the talking, and the words that came up most.
What is really slowing down enterprise AI? According to Pradnesh Patil, Founder and CEO of Altimate AI, the answer is not always the model. It is the data work behind it. In this episode of Data Unchained, host Molly Presley talks with Pradnesh Patil about the pressure facing modern data teams as AI adoption accelerates. They discuss why data engineers are buried in pipeline backlogs, how token costs can spiral out of control, and why enterprises need more open, vendor-agnostic ways to manage AI-driven data work. Pradnesh explains how AI agents can help with data pipeline creation, infrastructure optimization, governance, context management, and the painful migration projects most teams try to avoid. He also breaks down why walled-garden AI tools can limit flexibility, why context graphs matter, and how enterprises can reduce cost without sacrificing control. Cyberpunk by jiglr | Music promoted by Creative Commons Attribution 3.0 Unported License
Transcribed and scored by The B2B Podcast Index.
Speaker A: Hello. Welcome to the latest episode of Data Unchained. I'm your host, Molly Presley. Before I introduce today's guest, let me tell you a little bit about this show. If you're new to it. We started this conversation around the idea of how do we take advantage of decentralized data? And, and when I'm talking about decentralized data, you can think about how data creation has evolved over the last couple of decades. From originally, everything was kind of built into a building, a mainframe. The people, the applications, and everything were sitting in a building together. And then, of course, in today's world, there's data being generated everywhere off of sensors and phones, across data centers, clouds. Data is being generated increasingly in space, and those data sources have a lot of value, but there's also complexity in how do you gain access to them, how do you know what data is there, how do you move it around, and how do you do it in, uh, a responsible and governed sort of way. Data Unchained delves into some of the challenges around accessing this data, the solutions and best practices and outcomes that you can achieve as a business or as a data user, uh, as you implement some of the technologies that we talk about in the show. Without further ado, let me introduce today's guest. Pradnish Patil is the founder and CEO of Ultima AI. Pradnish, welcome to the show.
Speaker B: Hey, Molly, thanks for having me. Um, very excited to be here and chat with you about different things related with data.
Speaker A: So, um, as we jump into talking about data, we'll take a step back and talk a little bit about you first, which may be a data conversation too. But tell us a little bit about your background. I always find it interesting to learn about how founders end out on the path of founding companies. That takes a unique background and unique personality. So tell us just a little bit about yourself.
Speaker B: Yeah, I was always attracted towards the technology and, um, that's why I was trained as an engineer many, many years ago. And as I started working for different companies in the San Francisco Bay area, or sometimes people call it Silicon Valley, which is the cradle of the technology, um, worked for a bunch of big B2B enterprise companies. So I did engineering initially for Symantec, long back the makers of Norton Antivirus, and then evolved through my career, started doing product management, product marketing, eventually leadership roles for companies like VMware, Cisco, Palo Alto Networks, and throughout my journey as a part of team, I was responsible for building these amazing B2B infrastructure data products which provided tremendous value for the Largest companies in the world and then also brought in a lot of value to the companies we worked for. At the same time when I was working for these companies, I started noticing one problem again and again. Because data work is so complex, right? You have been talking about there is so many desperate sources of data right now. Finding people who can do that data work is always challenging. There is a lot of complex work and even beyond complex work, there is a lot of routine run of the mill work which usually is left behind. If you ask any data team, there is a backlog of couple of years that they need to get on with and it's always delaying business objectives, et cetera. As the AI waves started coming on few years ago, we realized this is the sort of a perfect marriage made in heaven. We have bunch of data work that's sitting there. We can take the AI, combine it together and start delivering the work faster and make the data teams accelerated. That was basically the inception of ultimate AI. So started company with my close friend, we founded it together. Initially we didn't have anything, just started talking with users, got a bunch of feedback and based on that started building initial prototype. The piece we built really worked out, it really started resonating. We built an open source, free to download, sort of a project. Initially a lot of people started using it. We started building enterprise offering around it and now tons and tons of people use it. Our product gets used in 200 plus countries. We raised a couple of rounds of funding, could hire amazing people to the team and basically working on this mission and vision to expand it more and more and identify more different types of data or different ecosystem, different sources where we can automate this work more and more.
Speaker A: So can you tell us a little bit? Maybe just imagine an enterprise customer in your head, um, how would they put your platform into place? Do you have humans and developers, the data engineers that are helping augment and train their team? Or is it more about the processes that they put in place and where to run the data engineering? Can you talk just a little bit more about what is the steps putting it into place and who's using it within the company?
Speaker B: So if you look at typical data team, there are different roles within a data team. You have your data analysts, you have data scientists, you have data engineers, sometimes data governors, data stewards also. Now fundamentally the first and foremost thing is bringing right data together from different sources and putting in some sort of data lake, then massaging and transforming it in the right way so it's usable for different use cases, whether it's analytics or machine learning, et cetera. Now that first and foremost use case of building data pipelines, maintaining those data pipelines, that's tons of work. Usually people have hundreds, sometimes larger companies have thousands of data pipelines. What ultimate AI does is we have built an AI layer that works across different tools and ecosystems and can help people build these data pipelines natively by using their existing coding assistants like GitHub, Copilot or Cloud code. Our system integrates with that. We also give a CLI and UI layer that they can use and in that way they can automate bunch of work when it comes to building and managing that data pipelines. Not just that, when you build and manage these data pipelines, they run on data infrastructure. And over the time data infrastructure has gotten more and more complex and then the scale has just gone crazy because the amount of data is exponentially increasing. Maintaining that data infrastructure is also very, very challenging. You turn your infrastructure configuration knobs all the time as data workloads change, or for existing data workloads, you need to optimize those workloads, optimize queries, optimize the larger workloads so that they perform better. And you run these workloads most efficiently. Also on different platforms like databricks, Snowflake or Bigquery, that requires a lot of human work as well. So we have built agents which can automate large chunk of that infrastructure management work as well. And some of these agents are so advanced they're completely ambient, can take hundreds of actions within an hour. So humanly it's not even possible to make that level of infrastructure changes very, very rapidly. So that's the second big area where we have built a bunch of functionality. And third area is in the data space. There is always this modernization going on, whether modernization of infrastructure, modernization of software. In the one word that scares people a lot, including myself, is migrations. Those are messy, complex years. Sometimes people spend years and still those projects are not successful. You get sometimes service companies throw millions of it. I have seen that movie so many times, um, with the AI coming in into that migration layer. Also what we have seen, you wouldn't believe me, almost sometimes 90, 95% of that work. Can AI come in and do it? Also a lot of times it's a messy work nobody wants to do. Even service companies don't want to do it. Um, that's where we provide tremendous value to customers also. So in the three areas, by and large, one is data pipeline building and maintenance, data infrastructure management and optimization, and modernization or migration of your data workloads and infrastructure. Also that's where we offer tons of functionality. And since the pain is so big, people love it coming and doing this work for them, maintaining pipelines or even documenting pipeline or optimizing their infrastructure and saving them millions, um, it's like amazing godsend for them.
Speaker A: So to help provide a little more context, pradnish, can you tell us who are your key competitors or what are the types of technologies users are typically comparing against when they assess you?
Speaker B: Most of the times we get compared against these AI solutions built by individual vendors. Like for example Snowflake has built code, text code or DVD has their own AI, or um, Databricks has Genie code, etc. Genie AI, I think they don't have the code version of it. Now people compare us against the data vendor individual AI tools. But usually the large limitation of these tools is they are within their wall garden, maybe work with one or two tools and they don't work with really hybrid data ecosystem with bunch of different tools we work with. And at the same time they force you to use their own LLMs which are not that token efficient, usually higher markups, m on the cost, et cetera. It takes you usually in the direction of using the features and functionality from that walled garden. It wouldn't recommend something which is outside of their wall garden. So that AI layer itself is bias. That's why I think using some independent AI, uh, layer to get the data work done, which talks with all of your tools, which is token efficient, gives you that flexibility of choosing LLM and at the same time very transparent and vendor agnostic, um, with open source is very, very important. And that's what we are seeing with our user adoption as well and with the components we have built in from the functionality perspective on the benchmarks as well, like ad benchmark or DEB benchmarks that are out there.
Speaker A: One of the areas we see a lot of customers struggling with is they don't know what next year holds for their AI initiatives, that they're just trying to get this pilot going or this project successful. There'll be a lot of innovation in the next 612 months. The business will have different needs and that opens. And flexibility gives them the ability to shift and be dynamic where if they lock themselves into a tool it has limitations that you've already spoken about and also reduces their flexibility for the future. Does that come up much for you as well?
Speaker B: Absolutely. Because the space itself is changing so fast, locking your sort of bets in. Sort of um, one vendor which is completely locked um, is not a good idea at this time. You want something open that work across your tools and at the same time gives you flexibility to bring your own LLM and use LLMs on your own terms as well.
Speaker A: That makes sense. So did you design the company for this era of AI and model building and agent training, or did you guys get started earlier on with kind of a focus in things like BI and other areas of data analytics?
Speaker B: No, not really. So when we started initially, we always heard this pain point that hey, we don't have enough people, there's just so much work, how can you automate it? And when we initially started company, there was nobody else. It was just me and my co founder, right, Trying to figure out things. Um, so initially we started very small. We started with use cases like hey, we got this data table, can you just document it? Or I have got this data pipeline. I don't like going in and testing data manually. Can you help me write some data tests? Because usually that was the work that was boring. Engineers like to write new pipelines, do more data modeling, more interesting where nobody wants to do documentation. Um, so we started with those use cases initially. Data ecosystem is so vast. There are so many tools. We initially started with just one ecosystem, which was DBT Ecosystem. And later we expanded to others like Snowflake Databricks, Airflow now and a lot of other ecosystems Spark ecosystem. But initially it was just those two use cases. We did it for only one ecosystem. And once we realized there is a lot of value in it, we started expanding use case by use case as well as ecosystem by ecosystem as well. But our mission was always how we can make this work faster. So ultimately the business users can get their dashboard or their ML inside or whatever they're looking for from data as soon as possible. Because that was always the big pain for heads of data and VP's of data. How I can satisfy my business counterparts and I have lived through that pain also. You ask for a simple insight or dashboard, it takes three months to deliver. Then people think, hey, my data team, why they are not delivering me? Are they ignoring me? But lot of people from the business side don't understand the complexity under that data layer. How many sources are there, the layers of infrastructure, the layers of governance that you have to do as well.
Speaker A: I'm curious, as AI got rolling, um, and the company I work for was very involved in building models and worrying about keeping GPUs saturated. That really big GPU is valuable, um, and having it not used is a massive waste of money. And resource. It feels to me as though that conversation, especially in the enterprise world, not the big foundational model builders, but the rest of the world who's using AI, that the streaming, keeping the GPU busy isn't really the challenge. I mean it's important, they're expensive. But it's more this idea of data readiness, data preparation, getting data to the right places, the right governance. Um, what is your experience when you're thinking about enterprise data teams? What are the pressures and big pain points as they're getting their AI projects launched, whatever those may be.
Speaker B: I think one you already touched, which is also the cost of the AI that is coming through. Because even though people are using foundational models like Anthropics of the World or there are different models, uh, available, they cost a whole lot and it's also based on how people use it. If you give wrong instructions to LLM models, it's going to pull a lot of information, going to consume a lot of tokens and that's going to cost you a lot of money. Um, you must have seen all these famous stories coming out of for example Ubers of the world, et cetera. Hey, I blew my year worth of LLM budget in just a couple of months. So that's what we are seeing happening very strongly as well. There is one side of it is the cost side of it. Second is what we see more in different uh, industries like finance industries, healthcare, etc. More around data sovereignty and data governance. How I can keep all this information in my environment and how I can do these things on premise in my LLMs as well. So the trend we are seeing is because of these two strong reasons a lot of people want to deploy the existing models in their environment and just use open source LLMs, etc. Because another thing of this is there are all these cutting edge models like uh, for example OPUS is out there from anthropic of latest GPT models for a lot of work we do, not just for the data space, even for the other segments. You don't really need latest and greatest of the models. Even a lot of open source models now have really big context windows. They are capable enough deep reasoning, you can use those models for your specific work. So that's where we are seeing when it comes to AI. Those two things from cost and governance perspective, how I can keep the entire thing in house, that's where a lot of people are moving towards also now and that the third thing is um, how I can actually get the results from the AI. I'm Using what I mean by that is if you push some tasks, some information to AI is going to generate some results for you. Are those world class results are ah, those equivalent to if I put my amazing engineer to do this work, are those on par with how the engineer would deliver it? And most of the times we see those results are not at that mark. And the reason behind that is LLMs are amazing but they are only as good as the information you feed into them. So there is a lot of this discussions around um, context layer and context graphs which will pull all the knowledge, all the information from your internal environment and how we can feed that information to LLMs and get the amazing results at the level of we want the level of output you want from these AI systems, et cetera. So there is that context layer part of it where a lot of innovation needs to happen. That's where we are innovating also and some other components. I'm not sure if you're hearing a lot about the whole agentic harness concept which is applicable for the data space as well as anything else which is a specialized work. That's um, where the context graphs are important, governance layers are important. Tools and skills, MCP servers, those components are important as part of the harness as well as the infrastructure part that is needed by the agents. All these five components, um, we have put a very nice diagram on our website like how the agentic harness should look like and what are the key components of that. That's our vision, that's what we are building. So you can deliver data work with really amazing results and at the same time they are lower, the cost that is coming from these AI models is lower and you maintain your governance layers also.
Speaker A: Those are some I think really complementary trends things. Other guests on this show I'm seeing at my work as well are seen in the space. Um, can you go another layer deeper into how you help? So let's talk about the cost side. Money is going everywhere. It seems like it seeps out of every hole possible in AI strategies. How do you help with cost containment?
Speaker B: There are two things there. So one is what we have built is a model gateway that is part of our product. So as you send a data task it goes through our model gateway and it gets, you can use the right LLM through that gateway also whether it's on premise LLM or you want to some um, level two model for certain tasks, et cetera, that model gateway will help you tremendously to reduce the cost by utilizing the right model. The second part we have done on the AI related data work is we have this functionality we call context compaction. A lot of time when you are trying to do a data work for a few minutes, the compaction happens. You must have noticed this in cloud, et cetera, tools. A lot of times if the LLM itself, uh, isn't aware of what data I need to preserve in that compaction, what is going to happen as the session resumes is going to go and pull the same context again. So in that context compaction layer, it's intelligent enough for data work that we have built, so it will compact the context without losing most of the useful information. What that translates into is less amount of token consumption as well. So in that way when it comes to LLM layers, we help save costs. And then there is a second part also we talked about data infrastructure a little bit. Infrastructure always costs a bunch of money. Uh, all these different data warehouse vendors without funding one of them. Um, usually if you are big enterprise companies, you are paying millions and millions in your compute and storage costs. But a lot of times what happens is your data workload changes throughout the day, throughout the week. Your infrastructure doesn't change automatically according to that because previously we tried to manage the problem a little bit where admin will change some configurations or sizes of the machines and configuration. But it's an efficient approach as a human how many times you're going to change it? We have built these ambient agents which will directly interface with your data infrastructure. Every few minutes as your workload changes, they adjust size of your warehouse or they adjust the um, clustering configuration, they adjust your idle time configuration and hundreds of other knobs. The result of that is your infrastructure is at Ah, almost 90 to 100% utilization. Now if you use your infrastructure that efficiently automatically translates into dollar savings as well. And that's one of the jobs data engineers did. Usually we save a bunch of cost on that. And third is I think probably it's obvious if you're going to agents for doing lot of your data engineering work or data analyst work etc. You're going to save on people cost also. So those engineering hours are actually much more valuable. I always tell people, hey, don't try to use your engineers too much for data infrastructure optimization because it's a losing game. Your engineers are actually much more expensive. So, so try to get your most um, ambitious project done then rather than just spending time optimization, let agents take care of it.
Speaker A: Yep, absolutely. So saving on the people cost, the infrastructure cost, optimizing, moving faster, that all fits right in With I think massive pain points that organizations are dealing with. And then you also mentioned the data migration and how you kind of overcome that massive pain point. And I definitely have. I worked in companies that we were service providers who helped with that and it was the thing no one wanted to do and gave them hives and all the horrible things, nightmares and everything. Um, tell me a little bit more. How do you avoid data migrations or make them more efficient? How are you helping there?
Speaker B: So we make it, I think we don't make it easier, we just make it. Automated data migrations are inherently complex and the reason for that is between two different systems. Usually there are differences, the architecture differences. That's why you need to do data modeling differently or there is a SQL dialect differences, so certain functions don't exist in your target system from source, etc. Because of that, if you try to do the translation manually, first you need a person who is expert on both systems and understand the nuances. After they do the manual translation, usually the human workflow is first you do the translation, second step is you go in and validate the data. I bet most of the cases the data doesn't match in the first time, Then you figure out, okay, what's wrong in data, then go back again, you change some code, then go back, validate the data again. This process goes on and on. I'm sure you must have seen tables with hundreds of columns. I have seen those. It's a grueling work, like going through that data, identifying exactly which row is not matching, then figuring out in the code what's not happening, et cetera, now and then later trying to explain it to other people. Also why I did that architectural change or in this system, this support doesn't exist, so I need to write more code, et cetera. And then you connect BI system or ML system to it and making sure those work okay. Also this is um, A lot of hard work is involved in this and usually people put bunch of engineers to get this done. Now imagine you have a migration AI agent which is going to look at everything. It is going to look at rest of your information in your environments also. Then what it will do is it will figure out exactly how the translation should be done. LLMs have all this information about different systems, etc. Plus we juice it up with our knowledge graphs about different systems as well. It will perfectly know how to do the translation. It will also give the explanation to you how the translation is done based on architecture, decision, etc. You need to model it a little Bit differently. It will do that for you. But the hard part of data validation, it will actually go row by row or create a data profile first, do the validation. If that matches, then it will go row by row and do the validation. And if it doesn't match, it will do the whole loop again. And it's a machine, it will do it 15 times, 20 times that usually us ah, as humans, I don't want to match rows.
Speaker A: I'll do it once, maybe twice.
Speaker B: That's not what I signed up for when I took this job. Right.
Speaker A: Um, yeah, exactly.
Speaker B: They do it so well that only usually 5 to 10% of cases where agent stumbles and then for a variety of reasons and then human has to step in and then only fix the part which agent couldn't complete. So that's why I was saying earlier I have seen cases where Almost we did 95% automation and just 5%. You have few engineers looking at it and then people don't mind. 5% is not that hard to get them.
Speaker A: So that seems like kind of a utopia. And you're focusing it sounds like largely around structured kind of database data. Does this reach into semi structured unstructured data as well?
Speaker B: Right now we are more focused on structured data pipelines, et cetera. We haven't started tapping into unstructured.
Speaker A: But I'm not just another hairball to unwind.
Speaker B: Yeah, there is just so much work to do and then even around unstructured data that is a whole different ball game once you start going in there that way we have unlocks and puzzles there.
Speaker A: So let's talk about agents a little bit more and kind of you've, you've touched a lot on them, but you referenced the agent harness. Um, I think we'll put into the show notes the link to your website where you, you mentioned there you have kind of a, a visual of what this looks like. But can you talk about what that is and kind of how it helps um, to both optimize, simplify the deployment.
Speaker B: So in, in simple terms, agentic hardness for any space are kind of components that particular work needs to make LLMs work really well for those particular use cases and stuff. Now when we look at the data space that work is your context graphs where you are pulling the information from your data stack, which can be your data lineage, how different tables are connected, you are pulling in query profiles to understand how well your workloads are running. You are pulling in data pipeline logs on whether there are failures or whether they are running most efficiently, et cetera or from governance perspective, you're pulling in access history and just making sure from data governance perspective, things are right right people are accessing right sensitive data, or even the right agents are accessing your right sensitive data. That part is the context part of it. You're just pulling the metadata and then the information that is sitting there in the data stack and you can pull it. The second part is the information that is not sitting in your data stack, for example, your business glossaries, your business requirement information. This might be sitting in your Google Drive or you have confidence documents or heck, it might be just sitting in a meeting like this where there is just a transcript that's available that you can utilize. Right? So second part, we feed in that information as well. And then the third part is the knowledge part of it. You know, the, the data functionality and features are evolving so fast because everybody in the data vendor companies also has access to cloud code and cursor of the world. So the feature velocity has doubled and tripled already. I think, um, in a year or so it's going to be very hard to keep track of all the new features and how to best utilize those features as well. What we have done is we have built knowledge graphs to capture all these best practices information. Like for example, we're talking about migration. What are the best practices around migration? Or what are the best practices around infrastructure optimization or building pipelines in certain ecosystem. So we have created this knowledge graph and all this together makes that context graph that is one of the big components of the harness. Second component is governance layer. So as we were talking about, it's not just humans. Agents are going to consume data also and they are going to run queries and everything. How are you going to govern those, especially across different ecosystems? And there are going to be hundreds of them, probably more than even humans. We have built governance layers around rules, permissions, access, which cut across the tools and give a common governance layer for these agents also, which is very, very important. Now the rest of the pieces I believe most of the audience probably already knows, like the MCP servers, right tools and right skills, those are part of the harness also. So you don't have to fumble with which MCP server I should install or which skill I should install. How should I configure this tool? It comes as a part of that harness. You click install it, you are set to sort of, at least for most of the major tasks we have seen. And of course you can extend it further by installing additional MCP server. Last but not the least, where we are Going to do some work. Also I was talking about that model gateway or context compaction or another piece. I think there is um, some discussion that's been happening in the industry is sandboxes. So as humans we always tested our code in development environments before pushing to production and we needed sandbox environments for that. Now imagine agents are coming in and doing some of that work also they need to test it somewhere. You wouldn't trust the agent to take their code and just run it in production. So how we can have data specific sandboxes Also it's not just about running code but validating that data and making sure right access levels are done. So that fifth component is the uh, infrastructure component. I highly recommend if this piece is interesting for our audience, um, definitely go check in our website in the platform section there is a nice diagram and much more information than I could give you in just last couple of minutes.
Speaker A: Yeah, a picture speaks a thousand words, right?
Speaker B: Exactly.
Speaker A: No, absolutely. And we'll point to that with the show notes as well to make it easy to find. So you've talked a lot about some kind of far reaching, um, different types of people in the data teams, different users of some of the reporting and the information that you provide. How, who typically do you work with to deploy this? And I'm thinking where does the budget come from? Um, who's the buyer, who's managing this? Is this going to AI teams? Is the kind of traditional IT team managing it? How do you see deployment happen within an enterprise specific?
Speaker B: Yeah, first and foremost we have a simple open source project as this agentic harness that anybody can go and just install even on their local machine and start using it. Um, it's delivered as a part of NPM package. So it's just a single command, NPM install and it gives them this agentic harness. Even with our uh, cli, which is ultimate Core, which is basically a fork of open code. But what's important is what's underneath that layer, uh, that we talked about as agentic harness platform. So that's at individual level you can just get it installed even on local laptop, it will install DuckDB for you, it will install some other open source tools and you can start building the pipelines. Now the layer above that, when you're trying to do this in enterprise environment, et cetera and you have your own systems, you can create the connections to this system also. So in that case we usually work with data engineering teams and one of the data platform admins who will configure those connections and get it Going and individual users can get it connected also. So it's not like one platform that everybody has to use. You create an instance, SAS instance and individual, your CLI code, uh, will get directly connected to that SAS instance. You put your key and it will be connected ready made for your data warehouses, your data, pipelining solutions, orchestration solutions, even BI solution. And it will start pulling that context and enforcing governance for you for that agent work as well. Now sometimes um, people try to integrate this in CI CD pipelines, et cetera. Depending on how the team is structured, you might need to get a DevOps person in the heart of a CI CD integration.
Speaker A: So usually you get started with a curious data engineer or someone of that type going out, getting started with open source download and then proving value, proving how it works within their environment. It expands from there. Is that fair to say?
Speaker B: Yes. Yes. And that's why we started even giving with our model gateway. We are giving away 10 million free tokens which should be good enough at least depending on at what level you use. Hopefully at least for a month to use even the LLMs through our model gateway to try out these different things as well.
Speaker A: Okay.
Speaker B: And this system works with a lot of times people have their own coding assistants like GitHub, Copilot or Cloud code. It works amazingly well with those coding assistants or their internal IDEs like cursor or VS code etc. Nicer.
Speaker A: So I think we touched on some of the key pain points that are very real I think for pretty much any data team today as we look at AI continuing to evolve and kind of looking into the future a little bit. Um, are there any tips you have for data teams on how to avoid mistakes? Maybe there's some secret report outs that you have that will tell them you're falling into this pit. Watch out. Whatever it may be something in your software, something you're seeing kind of as directionally where you're taking your company right now that might help them over time in planning for the future.
Speaker B: So I think that's a very interesting question because the space is moving so fast, especially AI, uh, every week, if not every day there's something new coming up. First and foremost, it's very hard to keep track of, hey, what are the things happening in the AI ah space and how many of those are relevant to the data work that I do? Also we actually heard this from multiple of our users. That's why what we have done on our website is there is a section where we have an agentic data engineering training area. It's completely free. So we take an effort to keep it up to date as the newest things happen. It starts from actually level one that hey, if you're new to this, how to use AI, uh, for data work, start here. After that you go level two, level three, etc. We keep that entire thing up to date as well as with the latest news, etc. So anybody who's trying to figure out what exactly is happening, at least they get our perspective on this. Like, hey, this is probably still very flaky, don't touch it. Or these areas are pretty mature, you should get on it as soon as possible.
Speaker A: Like you say, with the rapid evolution, new companies, new technologies are coming up, you know, probably by the minute. We probably have two new companies that just happen to come out in this space while we were talking today. How do you compare to maybe built in tools that are available in Snowflake or databricks or aws, um, and kind of why would someone select your technology versus something that's built into one of the um, analytics or warehouse companies? And I have my suspicions, but probably better just to hear um, directly from you.
Speaker B: No, you are right. I think the space of innovation is crazy in this space and what industry has done for it is there are a couple of benchmarks that are out there that people have built. These are open source benchmarks. For example, one is agentic data engineering benchmark. Ad benchmark is built by DBT Labs. Um, it's open source benchmark and right now we're trending number one on that benchmark. And second is what uh, is that measuring?
Speaker A: I'm not familiar with that one.
Speaker B: So that benchmark actually has a list of the tasks that usually data engineers and analytics engineers do and basically use your solution to complete those tasks. There is a whole benchmark evaluation process and you publish the results on hey, how many of these tasks, how much accuracy we got, et cetera. And based on that iterating, there is another benchmark called DAB also which is by people who uh, build Apache, Spark, etc. Which is also uh, very reputable. And even on that benchmark we are number one. And the reason for that is these are the components I was talking about context layer, governance layer, some other tools and skills we have given ready made in the harness. These components for us actually go across the data stack. If you look at any typical company, most of them don't have a single vendor covering their entire data stack. You have a different tool for pipelining, you have different tools for data warehousing. You have different tools for bi, so you need a layer that cuts across these different tools and talks with every tool to create that common AI layer for you. That's what we have done really well. You don't want an AI layer which is pro to one of the vendors as well because it's always going to recommend the features and then the components from that particular wall, garden or ecosystem. So because we have solved that cross sort of a tool problem so well across the ecosystem, um, that's why we are I think talk most on the benchmark as well as in the user adoption as well.
Speaker A: Makes sense. I was expecting the cross platform was probably the answer, but I find the, the benchmarking of the process itself really interesting. Something I'm going to go take a look at that. I've not. I've heard of all kinds of benchmarks in this space, but not that one. And that's a really valuable one.
Speaker B: And then as you can see, uh, it's a recent benchmark I think that came out a few months ago as agent data engineering started picking up. So definitely check out. It's very interesting.
Speaker A: Okay. And then as we tie up because it's hard not to hear the word tokenomics and cost of tokens and token efficiency, you know, without, you know, some kind of context to what where people are talking about. It seems like it's everywhere right now. Um, how do you play into that conversation? Do customers come in and say I'm worried about token efficiency and uh, can you. Is there the benchmark that you're tying into there as well or do you kind of tie into the people efficiency, the pipeline efficiency and all of this leads to token efficiency.
Speaker B: Benchmark I think right now is much more geared towards how successfully you complete the task. Token efficiency, the measure of that is basically the amount of dollars or euros you're spending you're going to get monthly bill and you know how you're doing your tokens?
Speaker A: Yes.
Speaker B: What we do in that area is um, as I was sharing earlier, we have a model gateway which people can utilize and they can use bunch of open source LLMs. They can have even LLM deployed in their cloud or on premise and use that. And in that case you're just paying for infrastructure and you can keep the costs really, really low or you can use some of the other providers like open router, etc. We also give our LLMs if people want to do as a service. We don't train LLMs or anything like that and that people can use which is much more cost effective compared to directly using LLMs via anthropic or OpenAI's of the world or directly using it from some of the, um, data providers also, because usually there is a markup.
Speaker A: Awesome. Hey Pradnish, thank you so much for joining the show. M. You're smack in the middle of a really exciting, um, time and exciting place of innovation. I wish you all the best of luck and definitely hope that some of our users are reaching out and thinking about, um, downloading some of the tools, the open source tools, and getting rolling and checking out your technology.
Speaker B: Yeah. Ah, it's been a pleasure. Amazing conversation. Thanks for hosting me, Molly. I appreciate it.
Speaker A: You bet. Thanks for listening to Data unchained, powered by Hammerspace. To learn more, visit Hammerspace.com if you have a guest you would like to hear on the show, email me@mollyammerspace.com.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.