
The Data Engineering Show · 2026-06-16 · 20 min
Key moments - from our scoring
Substance score
35 / 100
Five dimensions, 20 points each
The conversation explores the fundamental transformation of data engineering in the era of generative AI, with Pranav Motarwar positioning the field as bifurcating into two critical domains. Traditionally, data engineers focused on structured ETL pipelines, dbt transformations, and BI consumption - a pattern that remains relevant but is accelerating through AI-assisted tools like Databricks Genicode, Snowflake Cortex, and LangChain. The emerging second domain involves building data pipelines for AI agents consuming unstructured, multimodal data (audio, video, documents) through chunking, embedding, and vector storage workflows. Motarwar references an MIT Technology Review report showing AI use cases growing from 19% in 2023 to a projected 60% by 2027, signaling a structural shift. The discussion covers how feature stores enable millisecond-latency requirements for real-time recommendations and gaming ad systems, how individual contributors now own end-to-end pipeline ownership previously handled by seven-person teams, and why companies like Vespa and LangChain will compete with traditional data warehouses as multimodal data volumes explode. Rather than eliminating data engineering roles, the explosion of data generation (more data created daily than from humanity through 2008) ensures sustained demand - even as individual productivity increases through AI tooling.
AI for data (using AI tools like Snowflake Cortex and Databricks Genicode to accelerate traditional pipelines, modeling, and validation) and data for AI (building pipelines for agents to consume unstructured multimodal data through chunking, embedding, and vector storage).
Modern tools have reduced project timelines to approximately 30% of what they took three to four years ago - for example, creating entire dbt flows now takes days instead of months.
Traditional pipelines process structured, text-based data in defined formats, while AI agents consume unstructured multimodal data (video, audio, documents) requiring new workflows: chunking, embedding, vector storage optimization, and latency planning.
According to an MIT Technology Review report cited by Motarwar, AI-related use cases are projected to grow from 19% in 2023 to 60% by 2027.
Yes, traditional tools remain relevant because the core steps of data pipelines, ETL, and data modeling persist; however, they are accelerated by AI-assisted tools, allowing engineers to focus upstream on product requirements and governance rather than manual coding.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode has a few useful data points (the MIT/Snowflake stat progression) and the 'AI for data vs. data for AI' framing, but the majority of airtime is filled with vague career advice, repetition, and obvious observations about the industry trending toward AI. The ratio of novel insight to filler is poor for a 20-minute runtime.
19% of the use case in 2023 was related to AI. Uh, like providing the data to AI. Now from 2023 to 2025 it has been like 37% and the projection is by next year 2027 it will be 60%
you need to understand the market dynamics are completely changing in the sense like you need to be aware about the process of chunking, embedding and how you are planning the vector store
The 'AI for data / data for AI' framing is the only structuring idea offered, and it is not particularly contrarian or first-principles. Most of the episode repeats widely circulated discourse about AI changing engineering roles, PM/engineer blur, and data volumes exploding, capped by an explicit recycling of the tired 'data is the new gold' cliché.
Data is the new goal. Like, trust me, this line is very much important. Data is the new goal.
there are Clickbaits on the YouTube like Hey, data engineering is going away. There is no work for data engineers. How it is transforming into AI engineers domain. That is like clickbait. That is not true.
Pranav is an early-career practitioner who self-describes as 'pretty young in this particular space' and lists general domain exposure across a few companies without demonstrating leadership, scale, or a specific hard problem solved. The observations feel like those of a thoughtful junior engineer rather than a senior operator who has built something at meaningful scale.
I'm pretty young in this particular space
I've worked across different product based companies in different domains like risk and product, uh, as well as privacy and the core data engineering teams
The MIT/Snowflake report statistics and the rough dbt time-reduction figure give some concrete grounding, and specific tools (Vespa, LangChain, Databricks Genicode, Cortex) are named. However, most claims lack company-level detail, dollar figures, or personal case studies, and the Apple 'Tiro' job application reference is unclear and unverifiable from context.
There is this MIT technology review, there is this entire report that they have released along with Snowflake...19% of the use case in 2023...37%...projection is by next year 2027 it will be 60%
for dbt we were spending like maybe one month to create a uh, entire flow or something like that. Right now it has been reduced to almost close to 30% time
The host adds genuine value by injecting Firebolt-informed perspective on the BI vs. embedded analytics split and pushes on the multimodal infrastructure question with a reasonable follow-up about compute-intensive pipelines vs. serving. However, several questions are generic prompts ('What else is top of mind for you?') and no weak or vague claims are ever challenged.
I think personally beyond that by the way and this is something we see a lot of Firebolt. There's also a uh, split in like how analytical databases are used
You're mostly now talking about the serving side, right? So something like Vespa as a retrieval engine for like fast vector search and so on. I think the more compute intensive part is actually the whole embedding pipeline
Computed from the transcript - who did the talking, and the words that came up most.
In this episode of The Data Engineering Show, host Benjamin Wagne r sits down with Pranav Motarwar , a data engineer who worked across major tech companies, and the intersection of AI and data infrastructure, to explore how artificial intelligence is fundamentally reshaping the data engineering landscape not by eliminating roles, but by bifurcating the field into two distinct, equally critical domains. What You'll Learn: - Why the "data engineering is dying" narrative is clickbait: Data engineers remain essential because 60% of use cases by 2027 will involve providing data to AI agents, while simultaneously human-facing analytics demands continue growing, meaning more work, not less. - How to future-proof your career by mastering "AI for Data" AND "Data for AI": Modern AI Data Engineer roles now require both using AI agents to accelerate traditional ETL/DBT workflows AND building entirely new data pipelines (chunking, embedding, vector storage) designed specifically for agent consumption.
Transcribed and scored by The B2B Podcast Index.
Speaker A: There are two different aspects to data engineering right now. AI for data and data for AI.
Speaker B: The Data Engineering show is brought to you by firebolt, the cloud data warehouse for AI apps and low latency analytics. Get your free credits and start your trial@Firebolt IO.
Speaker C: All right. Hello everyone and welcome back to the Data Engineering Show. Today I'm super happy to have Pranav Motawar on. He's a data engineer, worked across multiple big tech companies, was a research assistant at New York University and is generally thinking a lot about how AI is changing data engineering. So great to have you on the show. Pranav. Welcome. Do you quickly want to introduce yourself?
Speaker A: Yeah. Hey. Hi everyone. So I'm Pranav. I've worked across different product based companies in different domains like risk and product, uh, as well as privacy and the core data engineering teams as well. So, so quickly to drive this conversation in the direction of like how data engineers have evolved in the last five years from starting with like very core data engineering job descriptions to like how it is evolving today with AI coming into the place starting like 2022. So that's where like I'm in the cross junction of like applying both AI and data skills right now in my current domain as well.
Speaker C: Okay, very nice. That's cool. Like I guess you kind of started your career around the early chatgpt moments roughly. Right. So you've gone through like the full transformation of like, okay, pre LLM, you actually write your data pipelines manually to like whatever's going on. Now take us through that transformation. How you think about like a modern data engineering workflow in like 2026 onwards?
Speaker A: Yeah, absolutely. So when I started my career, as you mentioned, it was pre chatgpt. So the work that we usually did was like creating the entire data modeling stuff and planning the data design, creating the ETL pipelines and writing the scripts and completing the last mile delivery where the entire product that we were creating were consumed by someone at the end. Like humans at the end. So that's what the case was pre like Covid or like post Covid as Well as pre chatgpt suddenly like from COVID even during the chatgpt era till like 2023, we were in the transition phase of applying data engineering beyond on premise infrastructure to cloud based environment. So that was the first transition that data engineers went through because that's how industry transformed and evolved. And speaking about the latest transitions like in 2025 and 2026, if you see the data engineering is completely se two different categories. One is where the end consumer is basically human or product which is used by humans eventually. And that is first category and then second category is where you are building the entire data engineering flow, pipelines, design and everything for the agents to consume, models to consume. And like that's how the field is like evolving in the last couple of years.
Speaker C: Right. I think personally beyond that by the way and this is something we see a lot of Firebolt. There's also a uh, split in like how analytical databases are used and like this is something we're seeing a lot I think right now. There's like this split between your traditional like BI and reporting stack at kind of enterprises that's powered by a snowflake, by databricks, by a bigquery, maybe kind of these types of systems. And I think the other part of the market nowadays is increasingly like building software on top of analytical databases which is relatively new. Right. Like historically there's always like look at any Fortune 500 companies like roughly the same stack for analytical databases, right? Like okay, fivetran to get data in, dbt, to run transformations, looker, tableau, Omni, um, hex, kind of whatever to run analytics. Nowadays also with just so much software being built there is more and more embedded use cases where similar systems actually power some specific product. I don't know, let's say like a cybersecurity, uh, threat detection product or like uh, real time gaming ad recommendation system and so on. And like for us that's been. Another thing is actually interesting to see is that there's more and more software being built by very technical teams on actual engines that are quite similar to like Snowflake, uh, or Databricks. But well these teams then build on different systems right? Like they build on a Clickhouse, they build on a DuckDB, they build on a Fireball kind of in our case. And yeah, uh, I think that's also like an interesting split that I've been seeing which I think is becoming more and more pronounced with the Gentex coding as well.
Speaker A: Yeah, absolutely. I mean if you see traditionally the analytical field is also evolving where right now the requirement of Feature Store is more and more relevant into the market. Like the use case that you mentioned, real time recommendations or maybe some data uh, infrastructure being developed for companies like Uber to map the drivers, fetch the estimated times and everything where the milliseconds, you know, requirement of the data latency is quite important. So that's where we have transformed from the traditional offline as well as online. We had like real time data infrastructure earlier as well. But right now even that is getting transformed into the case where the seconds are reduced to milliseconds with feature stores coming into the picture.
Speaker C: Right, Nice. So as you think through then like AI forward data engineering, what do you think is going to change about like the ecosystem landscape? Right. Like do you think people will still run tools like uh, dbt? Do you think people will still run their traditional BI tools? Like take us through what you think is emerging as like the next gen data, uh, stack basically?
Speaker A: Absolutely, yeah. That's where I mentioned when I started this conversation. Like as I mentioned, the entire traditional data engineering infrastructure or the steps that we were following are going to be there forever. Like it's not going to be completely building the data pipelines for agents somewhere. People are going to use the data, people are going to use the applications where the data layer is pretty much like sending the output for all these applications. So traditionally all this DBT stuff, ETL pipelines and data modeling will be there still relevant in the market. But the fact that products like databricks, uh, Genicode or even like Cortex, all these products are pretty much changing the dynamics in how we were traditionally writing the things. Like for dbt we were spending like maybe one month to create a uh, entire flow or something like that. Right now it has been reduced to almost close to 30% time that we usually spent like almost close to three to four years ago. So the steps are pretty much there. It's just that these tools are like helping us to fast pace the process so that we are moving into the upstream layer where upstream in the sense, like we are more focused on defining what the product requirements are, how we can drive the product, uh, revenues and all the stuff, governance and these are like some of the examples. But we are moving m one step ahead. Like all the developers, even at the individual levels, they have this like individual contributors. First it was like manager, then there is a team who used to like create these different individual tools. But right now that individuals are also like utilizing the agents as a individual assisted process. So that's where I feel like we are moving one step upstream and driving the business rather than writing this boring traditional DBT stuff. Uh, ETL stuff.
Speaker C: Yeah, that makes a lot of sense. And I think like they are basically, I think the also role of basically what used to be like an engineer versus a product manager are getting increasingly blurry. Right? Because the reality is, and you're seeing it on LinkedIn more and more, PM's actually coding more and more engineers kind of like taking on product responsibilities in some way. I think another big part of that dynamic by the way, is that historically to really leverage enchanted coding as much as possible, you need as fast of an iteration loop as humanly possible. Which basically means if you split kind of any decision into an engineer plus a kind of separate person on the product side driving that, well, you're not going to be limited by the speed at which you can do agent coding. Right? Like you're going to be limited kind of by your conversation. The fifth, uh, spec review, the going over a design review again, trying to like run some internal, find time for some internal stakeholder alignment, like all of the, I think also past shapes, especially in big software companies, around how things were built. Don't map cleanly onto how the next generation of software companies will work. Where it's going to be much more about kind of like someone owning something end to end, building it, shipping it, because that's the only way how to basically move and accelerate as fast as you possibly could with the technology.
Speaker A: Yeah, absolutely. And that's a great point. I mean we are in the process of owning the entire pipeline flow which was like maintained by a team of like seven to eight engineers previously. Right now that is being transformed into a place where as I mentioned, the data pipelines which are being used by humans, the data pipelines which are being used by AI. Ah, but the process is pretty much there. We have sources, we have the processing pipeline and we have the consumers. Consumers are like two different things, like one, humans and AI in the transformation phase. Now data engineers, first they were aware about ETLs, DBDs and all the stuff right now apart from that, you need to understand the market dynamics are completely changing in the sense like you need to be aware about the process of chunking, embedding and how you are planning the vector store and how you're optimizing the entire process. So the transformation layer is split into two different categories and you need to work on both the categories right now to be relevant in the market. I mean you can pretty much take one field and go deeper into it, but to grasp more opportunities to have that ownership of end to end process, a person should focus more on like even evolving and taking the knowledge over like all this rag based pipelines and all the works that data uh, side from the agentic view.
Speaker C: Completely agree with that. Yes. Cool. Nice. Yeah, that's exciting. If you think about data pipelines like for agents versus humans, like what's actually the difference? Why don't the past pipelines we had for humans don't just work for agents.
Speaker A: Yeah, because the fact that we traditionally were consuming and processing the data in a very structured format like text, as well as structured right now that is getting transformed for E agents. It will be pretty much unstructured files, audios, videos, it can be pretty much anything in the market right now. So if you go with the traditional flow, it won't work because it is uh, definitely defined for structured flow for unstructured and all these things where chunking. First we start with getting these unstructured files into a, uh, place. Then we start planning the chunking phase, then we embed, then we find out which is the particular vector db we need to store this in how the latency for the product needs to be planned out. This was not the use case two years earlier. And to be honest, like I'll quote one fact here. There is this MIT technology review, there is this entire report that they have released along with Snowflake. It's an excellent review for any data engineer to go through because they are speaking about how it is like the field is evolving. Like 19% of the use case in 2023 was related to AI. Uh, like providing the data to AI. Now from 2023 to 2025 it has been like 37% and the projection is by next year 2027 it will be 60%. So you understand like from 2023 to 2026 it is going at a pace from 19% to 60%. So for this to stay relevant and to help the companies own the entire structure by one person or maybe couple of people, rather than like having a 10 people's team earlier, you need to learn all the skills like from AI agent perspective.
Speaker C: Cool, interesting. Take me through in terms of unstructured and multimodal data. Right. Because another thing you're going to see happening is that data volume will just explode over the next couple of years. Especially as there's more robots and kind of physical AI coming online. Right. Which is going to be so deeply wired into the real world and physical world. Basically. What do you envision is going to happen there basically is. Do you think that mainstream data warehouses will. Because at the moment you can't run a video decoding pipeline plus embedding pipeline like in Snowflake, right. Or databricks. It's like, do you think there's going to be separate systems for that? Do you think it's going to actually end up being the same systems just like eating everything? Take me through your take on like this multimodal data infrastructure world Absolutely.
Speaker A: I mean that's a great point. So if you see all these companies right now, their focus has not been shifted completely to both these domains. They are still figuring out the AI stuff, AI pipelines and stuff. But there are companies like take example of Vespa, they are building a uh, data warehouse which can pretty much handle any amount of multimodal data. Right now these companies like these startups will be relevant in the market in the near future. And then once the impact is observed by all these companies, they will start planning more deeper into like how their products should be evolved for multimodal stuff. And right now which is focused on structured and unstructured but like text level data.
Speaker C: Right. Okay, interesting. You're mostly now talking about the serving side, right? So something like Vespa as a retrieval engine for like fast vector search and so on. I think the more compute intensive part is actually the whole embedding pipeline, right? Like chunking it kind of like all of these things. How do you think that part will evolve basically because in the end, okay, you put it in like a retrieval engine, like fine, that's only like 20% of the problem actually.
Speaker A: Yeah. So for chunking, embedding and all the pre planning stuff for the data, there are companies like LangChain and all. So they are working on the layer of improvising the entire memory layer as well as data processing latency and all the stuff. So if I am from the companies that you mentioned, like maybe databricks or like snowflakes, so they will eventually evolve and be as a competitor to LangChain eventually. Right now they are not, but eventually they will plan out BAS basically based on the report that hey, 60% of the use case now will be providing the data to the agents.
Speaker C: Makes sense. Cool, Interesting. Nice. Well what else is top of mind for you?
Speaker A: I feel like honestly the more important aspect here for data engineers is like there are Clickbaits on the YouTube like Hey, data engineering is going away. There is no work for data engineers. How it is transforming into AI engineers domain. That is like clickbait. That is not true. Honestly, there are two different aspects to data engineering right now. AI for data and data for AI. Both the things are essential for a uh, engineer to you know, plan their future. I'll explain the AI for data and data for AI part. So there are companies even like Apple Tiro, they are raising this relevant job applications in the market known as AI data engineer. If you see the job requirements, they have mentioned two things. Are you aware about the process of creating the Data pipeline for agents. Secondly, do you know how to use AI agents into your data engineering flow, like the process that you are following right now. So these are the two main requirements. So to stay relevant from the AI for data side, you need to get hands on with snowflakes and databricks, the Jenny, the cortex on all these products that they have built to plan everything like the data layer, the data modeling, the validation layer and the automatic test case, uh, features on your entire data pipeline. This is for AI for data. But secondly, as I mentioned, as we mentioned for the last 20 minutes, that's where you need to plan for data for AI. So these two aspects are equally important right now. You can't say that, hey, let me focus on AI for data rather than data for AI because both are going to be very much important for the next couple of years. Then the transition will smoothly go for the next five years.
Speaker C: Makes sense. Yeah. And by the way, I agree, like I think the field will be transformed in very deep ways. Right? Like I think the kind of like traditional job of the engineer is going to change. But the good news is like the amount of data volume is also exploding, right? Like there's so much new data infrastructure coming online all of the time and like new shapes of data. Right. We talked about multimodal data management now for a long time. So is a single data engineer more productive now than they were a couple of years ago? Yes. Does that mean we'll need fewer data engineers? Well, I don't think so, because at the same time the data volumes are going to explode so much and we'll just be able to do more.
Speaker A: Yeah, sure. Just to give an example of how the data level is exploding right now, the data which was generated by humans from the humanity Till the year 2008 was currently generated in a day. That's how the volume is like the data requirement. Data is the new goal. Like, trust me, this line is very much important. Data is the new goal. And that's where data engineers again, even if the team which was structured around seven, eight people, is now transformed to one or two engineers, they are going to stay relevant. Because the use case was only 10% requirement back in like five to six years ago. Right now that use case is exploding in two different fields as I mentioned. So even if you are doing the work for five, six people, there are more use cases coming. There are more data requirements coming into the market. The data is exploding like anything, so exponentially. You don't have to think about the fact that, hey, will I stay relevant in the market. That's not the case. Data is going to be exploding for the next decade or so for sure. It's not going to stop anytime soon. So that's not the case right now.
Speaker C: I agree. Cool. Well, very cool. It was great having you on the show, Pranav. Uh, thank you so much for taking the time. Any closing words from your end?
Speaker A: Yeah. To close this session, I would say that for the past couple of years I have been from the generation of data engineers which saw the transformation of traditional data engineering to big data engineering to cloud engineering to AI engineering in the span of five to six years. So I have pretty much seen the industry evolve, even though I'm pretty young in this particular space. So I feel like if you want to cope up with the market dynamics, you need to understand the requirements in the market and gauge your use cases, gauge your skills according to the market dynamics. So let's say I'm very good in the traditional data engineering flow. Hey, that's the good part. But now I need to evolve and understand the requirements from the next perspective. Like, hey, the data field is like evolving and how I need to plan my career so that my output is relevant for the company as well as myself.
Speaker C: Perfect.
Speaker A: Cool.
Speaker C: I think those are great closing words. Thank you so much for being on the show today. Uh, and yeah, looking forward to all your work. Bye.
Speaker A: Thanks for, uh, having me on the show.
Speaker B: The Data Engineering show is brought to you by Firebolt, the cloud data warehouse for AI apps and low latency analytics. Get your free credits and start your trial at Firebolt IO.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.