
Data Engineering Weekly · 2025-04-25 · 37 min
Computed from the transcript - who did the talking, and the words that came up most.
In our latest episode of Data Engineering Weekly, co-hosted by Aswin, we explored the practical realities of AI deployment and data readiness with our distinguished guest, Avinash Narasimha, AI Solutions Leader at Koch Industries. This discussion shed significant light on the maturity, challenges, and potential that generative AI and data preparedness present in contemporary enterprises. Introducing Our Guest: Avinash Narasimha Avinash Narasimha is a seasoned professional with over two decades of experience in data analytics, machine learning, and artificial intelligence. His focus at Koch Industries involves deploying and scaling various AI solutions, with particular emphasis on operational AI and generative AI. His insights stem from firsthand experience in developing robust AI frameworks that are actively deployed in real-world applications. Generative AI in Production: Reality vs. Hype One key question often encountered in the industry revolves around the maturity of generative AI in actual business scenarios.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Hello everyone. Welcome to another episode of Data Engineering Weekly. We took a little break this week. We are hitting the um, important aspect of AI and data that we are going to talk to industry leaders and trying to understand what is in reality, what is the ground truth of AI and data is happening.
Speaker B: Everyone, thank you for tuning in. Like Anand said, long time, but we have some fantastic lineup of speakers coming in. We haven't asked today over to your winners for a quick introduction on what you do and what's your area of specialty. And we dive deep into the topics that we have planned for you.
Speaker C: Hey everyone. Solutions for Coke industries. I've been working in the data and analytics for a couple of decades now. In the last couple of years, been focusing more on the AI side. When I say AI, traditional AI, bit of operations, generative AI. So we build modules, we deploy and we scale.
Speaker A: That's what I do.
Speaker C: Excited to be here.
Speaker B: Thanks Avinash. Let me start with some quick one before I turn it on to Anand. I work with a lot of AI conversations. Every time I meet somebody like you on my day job. Everybody wants an AI perspective. There are a lot of pilots that I see. Let me start with a, uh, more obvious one. Are you starting to deploy your models, gen AI models or gen agents into production yet? I wanted to hear your perspective in terms of what are the things that data leaders are doing today. Has it reached that production maturity yet? Are you seeing some success?
Speaker C: Yes, absolutely. So we were one of the early birds to get started, right? It's been more than two plus years we got started in this whole generative AI journey. I won't be able to share specifics, but we've got quite a few things in production already. So it could be in the various domains that Coke operates. So we've got quite a lot of things deployed, being consumed. There's a continuous feedback loop set up. So our journey is going pretty strong as well.
Speaker A: So Anush, you mentioned that you already are getting a lot of agents building into this production ecosystem. One of the things that in everyone industries and even though, even if you see this recent development from cloud about MCP servers and the tooling capabilities they are building, the underlying theme for that is like the data readiness part of it. So when you are experiencing in the real world how good bad is that, the data readiness that either preventing or accelerating a agent's getting into the production systems.
Speaker C: So what we do is we try to focus initially focus is really to define the problem statement well. So once we define the problem statement really well, we get to the value part of it. And to get to the value we work with different the business teams, finance teams, and we collectively agree on the value. And then we do something called a discovery session where we it could range anywhere between four to six weeks trying to explore the feasibility of doing something. And in that the quality of data accessibility, do we have the right kind of data availability? All those things comes into play. So we do a very holistic assessment of whether problem to solve. And mostly at the end of four to six weeks we have go no good. That's how we go about deciding on the project.
Speaker A: Okay, is there any standard framework that you use to assess the use cases? Because I'm pretty sure every in every company has this problem. Okay, I hear this AI very shiny things, but what is in it for us? Like how do you do that? Do you have any structured thoughts on that? Like how a company has to go ahead and then um, evaluating the use cases?
Speaker C: Yeah. We follow a couple of concepts called spontaneous order and Republic of science, which is very embedded to our culture of our day to day work. Right. We want people to come up with ideas, but we also at the same time don't want to discourage anybody. So ideas come up from conversations, excuse me, your lunches, coffee, wherever. And then we want to encourage people to come up with ideas and um, provide some kind of a playground for people to really play around and just ask the right questions. That happens. Pick of science is really the concept of many multiple groups experimenting by providing the right set of tools. So the first thing is really the biggest thing really is what are we trying to solve? Right. What problem are we trying to solve? That's where a lot of our energy goes to really define is this the right problem to solve? Do we have the right value? What would be the adoption rate? Uh, and that's what it spend some time on really on the demand side of it. Ananth if I want to put it. And then the supply side, I think there's enough technical experts to figure out how to do all the magic there. So this also involves talking to business folks trying to understand, hey, if you build a solution, are you going to use it? How is it going to create value? How many people are going to use it? And then you're open legal compliance, finance partners to come in and help us through it. That's a uh, little bit on the problem statement, the value side. Then there is this whole aspect of does this align with our enterprise? Is this part of does this sync up with our data strategy? Is there An AI strategy in place. What is the enterprise vision for managing some of these AI agent workflows or whatever solutions.
Speaker B: Right.
Speaker C: Does it. So usually what we've seen is that some of these use cases stir up slightly bigger conversations and then there is a whole list of when it comes to the textile that we love to explore. So just collectively, collectively defining the problem with articulation of value and doing some kind of a discovery session is a good way to start.
Speaker A: Yeah, I think there's a two key takeaway from what you are selling. Like I really like that idea of like developer sandboxing so that people can freely not only express the ideas but they can actually try to implement. I'm trying to see this one before introducing any technologies then create the same playground for the users and then trying to for the developers to come up with an ideas. And the Republic of science is some of a new concept here. So very interesting to look. Learn more about that.
Speaker B: That's a valid answer in terms of how you go about bidding, but I'm more curious in understanding what, what vertical team or what vertical use cases do you see as a common trend amongst your network of CDOs and enterprise AI teams? What are those early use cases that you see go live into production? Which domains do they represent?
Speaker C: Yeah, so before I answer that I want to share a little bit of uh, of how I believe this would develop, uh, everything we do. So this was shared by a professor from University of Mekong who had been to Bangalore and I got a chance to interact with him. Certainly he's been working on the AI side the last 40 years. So he believed it is going to turn out in three different ways. One is really just the basic stuff using generative AI which is you just get a lot of documents, you chat over data that's you know, first one or based on your pyramid if you can think of it, the second one is really automation of your workflows. Right now I won't go into the definition of an agent but anything that mimics human action is your AI workflows. The way I believe it is you can use different tools and to make that happen. And then thirdly the last one that the biggest one is really how do we leverage uh, AI change our business model. Right. Those are the three things. The way he explained, I think it resonates with me very well. So we also did a lot of channel over data and then for exploring use cases on what are the workflows that we can automate. But by keeping human in the loop which is extremely critical because we don't want to let the AI just go do stuff uh, without any governance and guard. So we are very careful. But from a domain perspective, Ashwini, just the common things that you're seeing in the industry. Stairs related then there is a little bit of finance operations, HR operations, ah, manufacturing related use cases. So that's what we've been trying to explore. Technology has just evolved quite a bit and they're also rapidly happening on the go. On the go. Because where we started two years back and where we are now, the massive change and we also keep seeing so much of content around mce. So it's just been, I would say it's not been an easy journey. It's been quite challenging in terms of understanding how the technology changes. Ground on which you are standing is moving Right. All the time and you're trying to fix a tire which is flat while the car is moving. So that's exactly what's up.
Speaker B: Yeah, I'll come to the MCB part in a bit. But then I think going with your own sequence right now, let's say sales, marketing, those consuming interfaces where there's a lot of manual loops going on today you're saying you'll apply AI or Genai or Darpus A to solve those use cases. And now how do you go about building the blocks here? So let's say you have decided to build an agent. Let's go into the tech a little bit in terms of unwrapping the onion peel and explaining people. I'm, I'm happy that you've defined what an agent is very simply. Let's not go into a very complex definition but if you want to double click on the stack once you started to build an agent, what does it look like? Where does the model layer fit in? If you can explain and double click on that based on what you're built or what you're seeing the industry getting built.
Speaker C: Yeah. So I can even take an example which also is being talked about in the uh, AWS re invent sessions. Ashwin, is there is this knowledge agents just to use a very generic name when you have experienced folks leave our uh, manufacturing clients there's also a bit of erosion of knowledge that goes with the person moving up. So then you have new people coming. How do you train this new people? Such complex supply chains and manufacturing processes, information about systems. How do you do that? We thought we started to experiment it with a couple of plans and a couple of production lines where we got a lot of time to it and we started creating the simple chat over data and we partnered very closely with Amazon. They also strategic partner in this whole journey. And we tried to bring a simple prototype of having people ask questions in a particular interface. And that was just taken really well because some of the manufacturing plants can be really large, right. And you may be confined to one section of the plant and you don't know what's happening any anywhere else. So just having some knowledge about the production lines, the machinery, uh, what's the history of that particular equipment. So that information was easily available for
Speaker B: people at the plants.
Speaker C: So that's a good example of knowledge agents and you can just extrapolate this to really any kind of domain. So then that's when the Amazon partnership, Amazon always has been a partner for us, right. And then we started to jointly explore this thing out. At any point in time I would never say we have figured it out especially in this realme.
Speaker B: Right.
Speaker C: So there's everything new that comes out every day. So that's how the whole partnership went to. We've uh, deployed a lot uh, people in Europe using this whole thing. So it's uh, I would say decent success.
Speaker A: Like you mentioned about this for knowledge agent. Right? Searching for knowledge is not a new problem, it is therefore existing and there is one company baited and it's, it's now it's a trillion dollar market cap right now. Right. But yes Google obviously so the search technologies or the vectors such as BM25, it's still, it's not, it's not something new that is already available there now how do you see or uh, what do you think that has happened different that a foundation model comes in that really help us to build this, the knowledge agent here.
Speaker C: I think when you saw today's announcement Also on the llama 4 everything is getting cheaper, better, faster and when you bring in the uh, SDKs of MCP in the mix, your integration to all the different data sources is getting easier which means that it is moving towards democratized way of doing things. It is moving towards, I wouldn't say commoditized right now, but probably it is heading in that direction. So it becomes easier and easier for people to get the data together, index it and then start uh, chatting over data. So that's where I am seeing that it's going to, the proliferation of these kind of knowledge agents is going to grow much bigger when the technology continues to evolve and I don't know where it really would be, hey, we are done. This is the new normal and let's go use it. I don't know when that would, when that would happen.
Speaker A: Expanding the question a little more. Traditionally if I wanted to work forecasting like for example I want to a search engine I think somewhere in elasticsearch or some kind of a search technologies and I wanted to predict my sales and I do some, whatever the prediction model that we do, building a logistic regression or a random forest, uh, the autoimm are kind of playing on everything, all those things. Right. How do you see the traditional AI getting disrupted by gen AI? How do people should think about that?
Speaker C: I think I've seen a slide from Gartner which very clearly articulated when you should use for what a problem statement right now some of what you mentioned, your mind forecasting. I don't think generative AI can do that. It's not meant for that because generative is really generating content, art, images, video, audio, text.
Speaker B: Right.
Speaker C: I don't know if it will ever get to a place where it can forecast things, forecast demand or products or anything or do any kind of traditional. There's also these machine learning algorithms which generative AI is not meant to do. For example, some kind of a sentiment analysis, a simple example. What are people talking about? My restaurant, for example. Now you still need the security traditional AI to come in and give you the sentiment analysis. I don't think generative AI can solve for it. Likewise when we talk about optimizing things, you still need mathematics based operations and research to come and solve. Generative AI, at least for now will not be able to solve for it. Uh, simulations for example, especially in engineering groups and product teams engineering products, they do a lot of simulation to test products on before it's launched. So I don't think that can be started either. So there's very clear um, cases where traditional AI should be used and then there's very clear areas where generative AI should be used. Now will generative AI disrupt traditional AI? Maybe not at this point I feel these two as a combination is a very, is very powerful. So I expect generative AI to be complimentary to the traditional AI and we could see many use cases where both are being used in a single to
Speaker B: solve a single problem.
Speaker C: That's how I see it. Yeah, totally few exclusively.
Speaker B: Raj, just double clicking on what you just said on that's not the traditional one. Come from a background that you spend decades in the world of data and analytics. One of the things that we all
Speaker C: have as a dumpyard, for lack of
Speaker B: a better word, is a number of dashboards and Visualization that we have created, BI visualization layer, whatever you want to call it, or analytical layer. How are you addressing those use cases and expectations from business users that this BI layer, um, this swarm becomes certainly intelligent and starts answering the world's problem and automating the world's problem and start answering every question that the business stakeholder has. Is that like a unreal expectation from business? And how are you addressing this whole notion of BI over AI in your stack? But also if you can give me a general overview in terms of how should we look at AI over BI or BI over AI?
Speaker C: Yeah, I think that's a topic that I think the whole industry is really talking about. If you see the latest announcement on the SageMaker, uh, the new framework they came up with, so they actually elevated that and they put BI and SageMaker, the new model, whatever it was called,
Speaker B: the queue index and everything can.
Speaker C: So they're trying to just build this unified platform across and this is probably a little bit of a moonshot where I believe it is going that you will get to a place where just chatting over data or generating insights over data, uh, will become way more easier than it is right now. Uh, now of course there's a lot of components like your data model, your ETLS transformations that will come into play that will get taken care of in some shape or form. So bi, in all honesty, BI is needed. The whole world still relies on BI because without understanding how it historically, what has historically happened, we're just shooting in the dark for the future. So BI is needed. But I think the question will be asked how do we solve the problem? Is it through a bi, uh, based solution or is it generative AI? Is it traditional AI or is it something else? Or it's just a core tech solution. So I think those questions will start to be become more prominent. Why historically, hey, I want some insights. Okay, let's go build a dashboard was the historical response. So I think people are being more literate, I would say illiterate. So I'm asking the right question, right kind of questions.
Speaker B: See one of the discussions that I typically have in my day job also Anaj is because work for AI, uh, company as well. Coming from Qlik where people build tons of dashboard and want to get your viewpoint on this particular uh, scenario that I am observing also in the industry people have invested significant amount of time building these dashboards and money, time, money efforts building strategy and building these BI dashboards. One very good thing that happens with these BI dashboards, architected, well defined modern dashboards is the whole modeling layer. It's model structured to a point that it is addressing a domain, a uh, particular segment within that particular company. It brings in the data. So there's a lot of data problem solved to feed in the BI layer. It's modeled, it's to a degree trustable. Obviously no data is 100% trustable. We come into the green and speed limit. But isn't adding just intelligence on top of it and augmenting that with unstructured information, information that's scattered across your organization a low hanging fruit in your perspective or where should people and how should people look at this particular problem? I am sure every leader is going to face this question where is my starting point for a lot of my gen AI deployments? The AI and ML I think people have invested over the last one or two years. Where do I see that low hanging success and where should I start?
Speaker C: No, I think the natural evolution of the existing BI is to get to evolve phase power, BI tableau, um, any visualization where they are able to integrate some kind of generative AI where the context of the question is taken and you get a dashboard right there in front of you. So that is going to happen at some point in time but that would probably be going from phase one to phase two. Then what does phase five six look like? Is where I was trying to articulate that BI may not be the initial choice to get insights. The challenge right now is just your data sprawl across sources in an enterprise. That's what a lot of folks are trying to solve for and it is very structure that we. But if those things tend to get simplified and I was reading a little bit about the PA engines but if those things get simplified then the way you consume information also could change. That's where I feel the next frontier of BI could be strength. Yeah.
Speaker B: In fact I was reading an article today morning or uh, Anand who was sharing. I don't remember that is we were using Google to search. Now you're asking Google to ask questions. I think we have now started seeing a questioning culture which is what prompting is. Right. So you go ask a specific question rather than finding more generic stuff.
Speaker C: Right.
Speaker B: We have specifically asking questions. I think the whole notion of how BI may want is also it becomes a questioning first approach. Dashboards will be a byproduct in which you may want an answer. Uh, or a chart could be an answer but not necessarily the only interface that you'll start from.
Speaker C: Because what do you do with dashboards? Right.
Speaker B: Download Excels and Do process.
Speaker C: Keep in mind there's a lot of questions when I hear from my industry friends that model outputs are shown in a dashboard also correct? Yes. They may not have the capability to build SaaS based solutions. So then the alternate approach to show the output of models in a dashboard so it will stay for sure. I don't going away. But then there's another extreme of dashboards where people read 2,000, 3,000 dashboards that just overkill and somebody has very good phrase called dashboard. So we should not go to that reason.
Speaker A: I think between this question and answering there is a middle ground. I think this is where my challenge or like many of the data leaders going to face the challenge. We talked about deadlock dashboard like without now LLMs are good at asking question
Speaker C: and um, giving you an answer.
Speaker A: But the question is if the knowledge source where it is deriving the answer is trustable or not trustable. And you have two C dashboard which dashboard that I should trust and which should know my trust. That also goes back to multiple layers that you are also mentioning about building a knowledge agent, codifying some sort of domain expertise and their knowledge and is that the right expect to codify and is it a trustable source? So there is a set of like you creating a fact and there is a verification process whether it is a correct fact or not. How do you see this layer uh going to evolve? Is it like AI can play a role strengthening this particular layer or it's more of a human problem?
Speaker C: I think it is maybe evolving to become data product which is governed. There is a clear ownership of that data product and that data product can feed your uh, machine learning models, operations research, generative M AI it can go to your website, it can go to your ERPs, it can go to your chatbots. So I think that's where it's going to evolve is my sense of again data products are also. People have many definitions to brag. We don't have debate about it. But anything that is governed secure meets the regulatory legal compliance requirements is probably where I see it.
Speaker B: That could be the middle ground.
Speaker C: And this data product based on a
Speaker B: particular uh, domain solving, catering to certain
Speaker C: requirements or solving problems in a particular domain, whichever domain.
Speaker A: Okay. So yeah, it's all perfect in building a customer source. Right. Do you follow any kind of a framework to establish in terms of a cultural aspect of it, in terms of the uh, process aspect of it and the tooling side of it? Is there any process that you see that work very well?
Speaker C: The first thing we did is we just Put out a notification to everyone saying don't put any company data on ChatGPT. That's the first thing. And we just kept as leaders kept reiterating this every now and then. So that's one thing uh, that we did. And then building a governance framework responsibly uh ethical. So there is multiple groups that we are talking to are uh legal and compliance regulatory trying to and the technology HR trying to formulate what it should look like because you don't want these controls to slow down innovation also. Right. It should expedite the innovation process while keeping your company's data and IP safe. So it's a trade off you both know very well. So we're trying to find the right balance out there. Of course from a security standpoint there are certain that's already in place and we've taken care of that, whatever we could control. But these are some of the things like all companies in the industry we're also trying to explore that.
Speaker A: I think you made a very good point. If an organization is not ready to use an AI, if you're not AI, don't put your data tool in a large language model because these systems immobile and you don't know uh, you are feeding a false information or the right information. So when you really wanted to do increment into a system you create a separate instance and you're completely do you need to nuclearize that memory? It's going to harmon hurt you a lot.
Speaker B: Yep.
Speaker A: Yeah.
Speaker B: A follow up on that, the data readiness piece again coming back. I come from a world where uh, until a year back I was largely dealing with structured data. Even if it was implementing lakes or lake houses 95,6% of the time it was largely structured data or if even if it's um, unstructured that it was like moving OCRs, moving video files, moving, not even taking front, just moving, just storing it for some magic to happen. Now this whole unstructured thing coming in. How are organizations expecting to solve the life cycle problem? And let me double click on what I'm asking as a question on life cyclists. How are people automating the movement of data uh into the models for memory or into the vectors? And then how do you ensure quality controls are there, trustability, fingerprinting and how does that entire lifecycle should be built? And again how uh, do you see that evolution happen?
Speaker C: So I think that's an area could see a mix of additional AI and generative AI coming together. You've got these, all these extractless gen AI solutions Coming into play Azure as a document implications that you put out. You've not tried it yet, but I've heard the cost is high. So we'll have to figure out what is a cost effective solution when it comes to unstructured data. I'm talking largely PDFs and documents where you extract data from. So I think that will evolve to something that can be easily customized, easily managed and is cost efficient. At the end of the day you should at least generate some ROI by solving some of these problems. That's where the focus is and that's where the whole industry is trying to build something that can go across domains and solve problems. So they also are looking at is it a SaaS solution, is it uh, an enterprise plot integrity trying to think of. So I think when I talk to my friend that's where I see things moving. The solution again can just be very different for the kind of problem we are facing. What kind of partnership you have with your hyperscaler system integrators. So it just a whole host of things to decide. And when we talk about technology there are at least 10, 15 things we'll have to keep in mind before we decide, before we pivot back to agents a little bit.
Speaker B: And just to follow up on my previous question, what does AI governance mean in this entire life cycle, particularly in
Speaker C: the context of unstructured? That's a good question. So generally if you think of a governance, uh, it should not just be the AI market, but it should also start with the original origination of data, the consumption of the data. So it's the whole cycle is what we should be governing really. So when you define a framework then you can probably double click on are we really looking at it unstructured data or structured data and go deeper into some of that. Mostly what I have seen in the industry is it's just limited to documents right now I have not seen a whole host of solutions deployed when it comes to audio or video that I've heard some of my friends experimenting where they've taken some audio videos and then converted that to images, process that or text videos where text is extracted, subscripts extracted and trying to use them as a base to form your knowledge as a knowledge base. I have not seen a governance framework for unstructured data, but that is definitely there being discussed right now.
Speaker B: Wonderful.
Speaker A: All those things right? If an organization is starting right now and what is the basic skill set they should look for? Let's say if I wanted to build my team on you build a toast team from um, the ground up is there pick one skill set to rule them all. Like you will see the mix of signals to that one. And um, how do you do that?
Speaker C: I think we probably had an advantage, a comparative advantage being a part of data analytics in ait, we had a bit of a comparative advantage compared to other groups, technology groups, infrastructure groups. When it comes to generative AI, there are multiple components here. One is just your basic understanding of uh, general statistics. How does OCR work? NLP techniques, which the team already knew, which helped, they have done. They had deployed traditional AI models, they have explored OCR techniques, NLP techniques. So they understood the whole ML engineering, data engineering concepts during CICD deployment and the infrastructure piece. Maybe not as experts, but they knew how to do it. So that gave us a bit of a head start compared to other teams. But that didn't mean that the other teams did not explore. People who really understood the whole tech engineering infrastructure also were well equipped to go do that. But I think when it comes to just blending traditional AI and generative AI together, I think that's where AI team stands to have an advantage compared to other tech teams. But if it's a solution, it's just pure tech with absolutely no nlp, ocr, uh, anything required, then they're also able m to do that. But we had a net because we understood how data flows from your generation to consumption, how people adopted. So we had that kind of knowledge which helped. And when we talk about just how did people really cope up with this change? I think a lot of it just happened on the go learning on the go. If you were to start from scratch, I think it's a fairly steep learning curve for anybody to go do it. I mean if, for example, if an infrastructure, uh, engineer wants to do it, that's pretty much head start because a lot of it is in the, in the generative AI side. We call it AI engineers and not AI scientists. So there's a reason for that. So we've covered a lot of ground, but you still have to write, learn to write a lot of basic code, SQL, Python, PySpark, those kind of like continuous upskilling, continuous learning and learning on the go. In fact, I read almost every day to keep myself up to date. Otherwise I think even if I'm just missed for a couple of weeks, I feel unbelievable.
Speaker B: I think that's becoming a problem in the world of a too much to consume and in fact I saw the announcement from Lama, like what naming convention these guys come up with, Maverick or whatever, the latest One, um, Llama, uh, Scout. Scout and yeah, Scout, Behemoth and Maverick. So that's the one. I'm like, can you not come up with like there's so much overload in our brains already, right? Too many names to remember. And I'm like, can you not just keep it? 4.14, 1.1, 4.2 max. I'll do the AWS model. Right? Max, whatever. Is the uh, easier way to do this?
Speaker C: I think.
Speaker B: Yeah, probably a feedback for the metasoaks here. A quick segue into the world of agents, right? What frameworks are enterprises adopting more commonly and how do we go about designing Avinas in your point of view? A quick touch point on that.
Speaker C: And we go into how your organizations
Speaker B: and how you're preparing your resources and preparing for the agentic world. Then we go from there.
Speaker C: See, Frameworks as of now is really technical framework. What kind of cloud are you using? What models are you going to use? Is it open source, closed source? Working out what kind of orchestrator has vector data? Really, it's just the whole infrastructure that we are trying to figure out. When it comes to framework, of course we involve our partners, uh, from compliance, from domain experts. All get together on single card. What is the best way to solve for it. Usually the business and the compliance folks, they have a point of view, but they don't dictate the technology that needs to be used. That's when some of the CTO CIO firms come in and help us guide through it. But frameworks per se is really on the technical side, of course. Security, governance, all those comes into play. So there's multiple groups that come into play and try and figure out the way forward. I don't think there are clear frameworks out there that I've seen. It's just collectively leaders coming together and figuring this out. Eventually this will start to become templatized and frameworks start to evolve. That's how I see it evolving. Wonderful.
Speaker B: Just a complete pivot from there. Let's leave the framework spot aside. We all would inherit or most data leaders listening to this podcast. We'll inherit a team that's coming from ML data, uh, integration, data analytics to some extent.
Speaker A: How are you?
Speaker B: How should we all prepare from a resourcing perspective to meet the modern demands of an ashram? Where are you putting the bets on? How are you preparing your teams for this? And we go, uh, if you can also address the skill set and the knowledge that people need to pick up to become ready for the next wave,
Speaker C: that will be Helpful. Yeah. I was reading about the llama for today and they talked about this whole, it was moe as a mixture of experts. I think that was a phrase that was used if you have, if you're having a team it doesn't mean everybody does everything right. There are uh, experts who are good at a certain domain in a certain capability and we help them get better at it. So it could be a person who's somebody who's really good in supply chain solving supply chain problems using AI then how do we enable that group or
Speaker B: a few people to go deep?
Speaker C: Because having business understanding becomes extremely important in today's world. How do we equipment him about how our business operates, how does it make money, what's happening right now? And then there is the other algorithm side of it. So um, identifying pockets in the team who does what kind of work becomes the first step and then how do you help them go deep and build expertise becomes a second step. I suspect your question was more around in the world of gen, how do we do it? So I think there's multiple resources available, number one you and this just spread and people can start to read up there's playgrounds that people can start to leverage for free and start to use some of it. Plenty of data sets available to explore and then just keep abreast with some of the biggest developments that's happening. The best thing I've seen is just if you get an opportunity to work on a problem, a ah statement solve it using generative AI. I think the learning is the fastest
Speaker A: if that happens on the same front of keeping up things we as a developers are trying to build automation to out job or something and one of the bigger things that people are talking about is AGI. So should developers should worry about it. What is your take on AGI and how should an engineer or a graduate coming out of now should be prepared for that? Um, you know the industry data analytics
Speaker C: AI is all if it's, if it's just generative AI then first learn about the foundational blocks and I believe the foundational blocks is really SQL, your Python and the visualization tool right to stack with that, those are your foundational blocks because end of the day every person who's working in AI is trying to take a data set and generate some insights out of it or in an uh, agent thing they're trying to mimic human actions. But then even there there's a lot of. So start with these three frameworks then start to get to technicalities of generative AI, understand how deployments Happen let's say if somebody has a problem statement, they build a prototype, it works every that works. Then let up on your cloud platforms, your deployment infrastructure, learn about DevOps CI CD and learn about how what happens after people start to use the information or the insights. Given that whole understanding of the cycle and the technologies associated with that would be the way to look at these.
Speaker A: Yeah, I think one of the things that you mentioned about that like even as an AGI stick to the fundamentals trying to learn the basics because everything is based on the data and data continuously going to be missing and somebody needs to do the genitive job anyway to that quality.
Speaker C: Most important, data quality is not data quality is not a team anymore. Data quality is becoming very cultural in nature like finops, right People any person who's. Any person who is accessing some data and feeding it to a model or a dashboard, it is their responsibility to make sure data quality is taken care of. Of course there's a bigger data, raw bad data being ingested into the ERP where things happen, don't tell that's a whole different problem people are trying to solve. But whatever uh, data you need to solve a particular problem and ensuring that the quality is intact is really everybody's job. While there are data quality teams set up to do it, I actually feel the whole model has to change and becomes everybody's responsibility. Just finops right? Cost, cloud costs and the way you spin up instances, auto shutdown those things companies have teams which do that. But if every person is accountable for ensuring to select the right instance for example then naturally you are keeping the
Speaker B: cost on should be everybody's affair. And I think we've had these problems and the tool stack for what, 10, 20 years now. I don't think we have moved a needle to the right. It's probably become more worse than I
Speaker C: think with AI coming in.
Speaker B: Uh, there are two school of thoughts here like AI will be intelligent enough to ignore DQ problems or improve on what it already has. The other one is you give more
Speaker C: context and rift data set so data
Speaker B: quality becomes a fundamental denominator and shifts. But I think both the school of thoughts I think would mature on its own maturity curve. I think the future is looking a lot more promising than where it stands today.
Speaker C: And just one comment, Ashwin, you also mentioned about these whole data lake Delta lake lay houses. So there's a whole host of massive repositories we built. We're bringing data from 10, 20, 30, 100 different sources getting into a single place and then you Realize that AI never needed so much data.
Speaker B: Right.
Speaker C: And then imagine doing data quality gigabytes and terabytes of data versus identifying. Hey, this is my problem. I just need data to solve this problem only and doing data quality over that. So these are very two different ways of handling it. So I think should go down the data lake. Data lake houses path only if it makes absolute sense. Otherwise just identify the problem, take the data related to that.
Speaker A: Yeah, I think that's like one of the interesting theme throughout organization that is emerging. We talked about knowledge graph and knowledge agent and the importance of what is the right signal versus a wrong signal and then continued impact of data quality issues and all those things. Like a couple of months back I was blocked about the rise of AI data engineer. Essentially I was imagining the similar thing. Some kind of an AI data engineer role will emerge. They're purely focusing on building those or they might be called as knowledge engineer. I don't know what knowledge engineering means. It's purely focusing on building that layer to do that. Do you see either a data engineering role emerging of this or um, like a new breed of engineering role will companies adopt? Do you see any analyst science in your companies?
Speaker C: Yes, I think I really like the one post you put out done on that data engineer's role shift from writing code from scratch to debugging to integration to data quality and other things. So I think that is true because now AI is able to write a lot of code. So as long as we good prompt which covers a lot of context, you are getting a lot of code. So the cycle time, development time is reduced greatly. Now you have a base code which you have to improvise on and then even to do that there are I think I heard a couple of startups are building these data engineer Personas which will actually do everything for you. So I think there is just a big shift happening in terms of data engineering roles. Also um, now for what is the end destination? I don't know because while of course I've used the whole wild coding as a phrase but people will still generate a uh, code and is that code, it's never copy paste. Right. We all know that. So there's a lot of work that needs to be done and then how do you take it and scale it across the different volume altogether? So we've at least solved the building blocks I would say. But again the whole uh, foundational elements of data models and all still have to be done. So I think there are pockets where AI is aiding expedited delivery in terms of code base and all. And that will continue to evolve.
Speaker A: Okay, awesome. I think overall I would say like this is a golden age of data engineering. Right. I mean there's the we accelerate to the manufacturing one as it, but we have not accelerated the standardization, trustability and uh, of the system. Uh, that is a mix week of challenges that we've been severed. Avinash, before we are closing up, do you have any final thoughts to share to the new engineers or the companies that looking to make their AI? Ah, data practices. From your experience, what advice that you will give?
Speaker C: Yeah, I think I feel most of the folks are talking about supply side of the equation and hey, what technology is needed? There's a lead days, great days, graph ramp, there's agent, then there is mcp, then two weeks later there'll be something new. So people are really a lot of focus on the supply side of the equation. But then I have not seen many people talking about the demand side. What is our AI strategy? What is our gen AI strategy? Can we have five different uh, applications for generative AI solutions or do we need to be integrated into a single platform? If you go down that path, what's on, how do we do that? Who's accountable for defining a framework for identifying a problem? What kind of value? ROI is still a big question in adoptive industry leaders, right? ROI of generative AI is not clear. So how do we address those kind of problems and then identify problems which really businesses care about is not loved enough. So I think that should be the focus area. And then there's other group of people who are trying to sell generative AI because they want to do some work. Now that will not work either because if your leadership is not AI literate, you're not going anywhere. So you need TIEC at the top, sponsorship at the top, OS at the top. Um, think of this as a journey and not go product project by project. And you need to have a three to five year vision when you make these investments. If you don't have it, then you are just doing project by project and you very quickly get to a point where you somebody says, hey, what's the ROI of all investments happening? Do you not have good answers there?
Speaker A: Wonderful.
Speaker B: Anas, I think you nailed it in terms of the valuable advice here.
Speaker A: Awesome.
Speaker C: Great.
Speaker A: Thank you so much spending your valuable time.
Speaker C: Very appreciate it.
Speaker A: Thank you.