
What's New In Data · 2025-11-17 · 1h 5m
Key moments - from our scoring
Substance score
49 / 100
Five dimensions, 20 points each
This episode brings together two database veterans with complementary expertise: Sharish Prabhu, corporate vice president for Azure databases at Microsoft, and Aloka Gupta, co-founder and product lead at Striim. Prabhu oversees Microsoft's entire operational database portfolio spanning SQL Server, Azure SQL, Cosmos DB, and open-source offerings like Postgres and MySQL. Gupta leads Striim's unified platform for real-time integration, replication, and streaming analytics. The conversation explores how enterprises should approach database selection in an era of AI-driven requirements and workload complexity. Rather than chasing every new database option (there are ~400 listed on Stack Overflow), the discussion emphasizes strategic choices: Postgres as the modern default for most applications, SQL Server for mission-critical compliance-heavy workloads, and Cosmos DB for internet-scale, multi-region applications. A key insight centers on database logs as the foundation for replication, analytics, and AI use cases - a capability that traces back decades but now powers heterogeneous data integration and real-time streaming. Vector indexing belongs in operational databases like Postgres, not separate specialized engines, reducing unnecessary data movement and operational complexity.
Postgres is the default answer for a huge range of applications in 2025 - it's the choice for most SaaS apps, rapid prototyping, and workloads needing a reliable transactional store with vibrant ecosystem support. However, the right database depends on specific workload requirements around consistency, scale, compliance, and latency.
Cosmos DB is the right choice for internet-scale, user-facing applications requiring multi-region writes, single-digit millisecond latency guarantees, RTO-zero characteristics, and tunable consistency models. ChatGPT uses Cosmos DB for this reason, whereas Postgres is not designed for those requirements.
No - vector capabilities should evolve within operational databases like Postgres rather than requiring separate specialized engines. Postgres has strong vector extensions that cover most RAG and graph RAG scenarios without moving data from your transactional system, reducing operational complexity for most enterprises.
SQL Server excels for mission-critical, compliance-heavy workloads needing advanced security integration, in-memory and column-store optimization, encrypted columns, high availability features, and world-class analytics capabilities - scenarios where specialized relational design provides advantages.
Database logs have evolved from redo and recovery mechanisms to infrastructure supporting heterogeneous replication across different database types, physical replication, high availability, and now real-time streaming for analytics and AI model training purposes.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains a handful of genuinely specific insights - the Cosmos DB/ChatGPT architecture, MCP replica-routing as a production protection pattern, and specific Postgres failover-slot progress - but these are diluted by lengthy bio sections, generic AI commentary, and repeated platitudes about 'the right tool for the job.' Roughly a third of the 65 minutes yields substantive signal.
ChatGPT uses Cosmos DB, all the user-facing applications of ChatGPT. Every time you interact, any message operationally it goes through Cosmos DB
there are vitamins and there are painkillers in this space
Most frameworks deployed here - 'right tool for right job,' 'don't do lift and shift,' 'vitamins vs. painkillers' - are recycled industry staples. The most original idea is the concrete MCP-via-read-replica isolation pattern and the observation that cosine-similarity search is effectively what a 'semantic LIKE clause' would look like if SQL were designed today, but neither is developed deeply enough to count as first-principles thinking.
if we were to design like clause in the modern era, it wouldn't be like a regex-based like clause, right? We would have done something like semantic like clause, but we ended up with cosine similarities
expose kind of a database engine you know with some replicas, and then now having one of the replicas being exposed over MCP to an agent
Both guests are genuine senior practitioners: Sharish is CVP for Azure Databases, a founding member of Cosmos DB, and a SQL Server and SingleStore veteran - real depth at scale. Alok co-founded Striim after owning Oracle's LogWriter and spending time at GoldenGate, giving him unusual CDC-layer credibility. The episode is slightly promotional (Microsoft-Striim partnership context) which constrains candor, preventing a higher score.
I was one of the founding members of Cosmos DB
I used to own LogWriter for a long time
There are several concrete data points - Cosmos DB's four-nines single-region and five-nines multi-region SLA, single-digit millisecond latency guarantees, Striim customers running 10,000 - 15,000 Postgres databases, and specific Postgres version references - but the episode never names a customer, cites a migration cost, or gives a timeline with business outcome. Claims about 'largest banks' and 'manufacturing areas' remain unnamed and unverifiable.
single-digit millisecond latency guarantees with a four nines HLA guarantees in a single region, five nines with multi-region
I've seen single customers adopt you know 10,000 to 15,000 Postgres databases
The host sets up reasonable framing for each topic and occasionally surfaces useful connective tissue (e.g., linking MCP to the read-replica pattern), but questions are predominantly open-ended invitations rather than probes, there is near-zero pushback on vendor claims, and affirmations like 'absolutely' and 'great question' dominate transitions. No genuine disagreement or tension is surfaced across the 65-minute runtime.
absolutely, and you know, having the broad kind of reactive view of what's happening right now
Yeah, absolutely. It's changing very fast
Computed from the transcript - who did the talking, and the words that came up most.
Modernizing databases in practice involves more than just moving data - it requires rethinking how systems, developers, and AI interact. In this episode, Shireesh Thota, Corporate Vice President of Azure Databases at Microsoft, joins Alok Pareek, co-founder and Executive Vice President of Product Development at Striim, to discuss the evolution of operational databases, the rise of real-time data movement, and what it really takes to modernize at scale. Together, they discuss how Microsoft’s Unlimited Database Migration Program, powered by Striim, enables organizations to migrate heterogeneous sources - from SQL Server and Oracle to Postgres and beyond - into Azure with speed and precision, creating a modern data foundation ready for the next generation of intelligent applications. What's New In Data is a data thought leadership series hosted by John Kutay who leads data and products at Striim. What's New In Data hosts industry practitioners to discuss latest trends, common patterns for real world data patterns, and analytics success stories.
Transcribed and scored by The B2B Podcast Index.
1 - > SPEAKER_02: So first, uh Sharish, I'd love for you to uh 2 - > tell the listeners about yourself. 3 - > SPEAKER_01: Yeah, thank you. 4 - > Sure. 5 - > So um, you know, I'm corporate vice president for Azure 6 - > databases.
7 - > So what it means is that I have the remit to uh to manage all 8 - > operational all TP databases for Microsoft. 9 - > This includes on-premises instances like SQL Server, 10 - > obviously. 11 - > Um we have IaaS offerings for SQL Server and you know a couple 12 - > other databases, but primarily the pass offerings in Azure SQL, 13 - > Azure Customers DB, which is a very differentiated 14 - > non-relational database, and then the OS databases, both 15 - > Azure Postgres as well as Azure MySQL.
16 - > And we've recently shipped it, uh shipped the database into 17 - > Fabric as well. 18 - > So SaaS, so all the way from on-premises uh to IaaS to pass 19 - > to SaaS. 20 - > All anything related to operational databases is kind of 21 - > under my remit. 22 - > And you know, if you want me to talk a little bit about the 23 - > carrier, then I could just go in on a just a few bits here.
24 - > Um it's it's been a long journey through the evolution of 25 - > databases, I would say. 26 - > Uh I started out working on SQL Server primarily. 27 - > Uh I, you know, that's kind of really where I learned the 28 - > fundamentals of OTP systems. 29 - > Uh I worked on various aspects of SQL Server, shipped a bunch 30 - > of releases back in the day.
31 - > Uh, you know, it also helped me like learn the craft of building 32 - > software. 33 - > It was sort of my formation years. 34 - > Uh and then I moved on to being one of the founding members of 35 - > Cosmos DB. 36 - > Uh I that was a complete shift in mindset, you know, going from 37 - > a single node, high reliability systems, to kind of thinking 38 - > about globally distributed, elastic kind of databases.
39 - > And it's also non-relational, so there's no schema there. 40 - > Um and and it kind of taught the uh distributed systems uh 41 - > consistency challenges, uh, scalability, internet scale, 42 - > user-facing applications, and it's kind of a different muscle. 43 - > Uh along the way, I also took a single uh small detour uh to 44 - > join Single Store. 45 - > It's a great company, uh, which basically was focusing on 46 - > real-time analytics, you know, sort of really challenging the 47 - > transactional analytical boundaries with a very novel 48 - > architecture.
49 - > Um and that experience was valuable because it kind of 50 - > forced me to think about like speed agility, leading closer to 51 - > customers and deeply thinking about performance at scale and 52 - > such things. 53 - > And now I have the you know the privilege of uh effectively 54 - > leading uh the database portfolio that I was mentioning 55 - > a minute ago. 56 - > Um and I I think in a way, you know, I find myself privileged 57 - > and and grateful that I got a chance to uh to experience 58 - > different kinds of these experiences across the board 59 - > from SQL to Cosmos to you know single store and then back to 60 - > all.
61 - > And then now I'm focusing quite a bit on all these databases as 62 - > well as open source and Postgres in particular. 63 - > And you know, it kind of really puts all of these things into 64 - > the context of Azure's broader mission. 65 - > So so yeah, that's kind of like my uh the the the through line 66 - > in my journey is that you know I've kind of really been drawn 67 - > to the challenge of like developers, enterprises, 68 - > systems, etc. 69 - > So that's kind of really who I am.
70 - > And uh yeah, I I enjoy what I do. 71 - > SPEAKER_02: Yeah. 72 - > Incredible here, just the the uh both depth in what you've gone 73 - > into database systems as an engineer and also the breadth of 74 - > you know the different types of systems you've worked on, uh, 75 - > which is I think incredibly valuable, especially as 76 - > especially as these workloads are evolving so much and the 77 - > requirements are are now being completely transformed by AI. 78 - > We'll get into that.
79 - > Aloka, you know, this is your second time on the show, and you 80 - > know, you also have a very deep uh database background. 81 - > Uh, would love to hear about that as well. 82 - > SPEAKER_00: Yeah, thanks, John. 83 - > And uh Sharish, yeah, very impressed.
84 - > You know, uh I think the journey is kind of remarkably uh 85 - > parallel, I would say, right? 86 - > The formation years and the foundation years, I got to be 87 - > part of the Oracle database and uh kind of grew up there, um, 88 - > got to got to work on a lot of interesting features. 89 - > This is my second time, John, on your show. 90 - > Uh so uh hopefully you'll invite me again and we'll make this one 91 - > of those Saturday Night Live.
92 - > I'm hoping to get a robe from you as well. 93 - > Uh you know, SNL5 or maybe JK5 or something like that. 94 - > And hopefully Sharish will be back as well. 95 - > Um yeah, so you know, um like like you know, I said, uh I 96 - > started off in the Oracle database.
97 - > Um uh and and I've taken an interesting journey all the way 98 - > till till uh Stream where I am right now. 99 - > Uh so I'm uh one of the co-founders in Stream. 100 - > I run all of the product areas here, uh right from engineering 101 - > to uh product management to product support, uh 102 - > documentation, all of the product aspects come under me. 103 - > So I really enjoy doing that.
104 - > Um and then um also it's a it's a it's a it's a great place to 105 - > be. 106 - > Um, you know, on one hand, we are working with some of the 107 - > leading global large uh enterprise customers. 108 - > Um so one day, you know, I find myself in a conversation, you 109 - > know, with some of the largest banks in the world, uh, try to 110 - > look at some AI use cases. 111 - > And the next day I'm uh arguing with one of my engineers that um 112 - > the amount of you know, just debugging stuff in the in the 113 - > stream server log is excessive and I don't really like it, and 114 - > we need to do something about that.
115 - > So um, but just going back to the database, spend a lot of 116 - > time in in recovery, John. 117 - > Um uh I think at the time that I joined, I found that to be one 118 - > of the hardest areas. 119 - > You know, uh oftentimes uh in support, I would look at the 120 - > most challenging things, and uh, you know, people would take a 121 - > backup and they would restore it and the database won't open. 122 - > And um, you know, oftentimes these support engineers would be 123 - > like patching file headers and fixing you know checksums and 124 - > stuff like that.
125 - > There's like all kinds of strange utilities they had to do 126 - > block edits, which are at the in this day and age, like highly uh 127 - > not secure. 128 - > But at that time they used to do that. 129 - > So that's what kind of fascinated me towards uh towards 130 - > recovery. 131 - > Had a good fortune of uh uh getting involved in some of the 132 - > redo generation layers.
133 - > Uh so I used to own LogWriter for a long time. 134 - > And it's been fascinating, you know, what uh sort of that one 135 - > piece of code has uh how it has shaped kind of my own thinking, 136 - > because we went all the way from kind of logging uh for backup 137 - > recovery purposes to physical replication purposes, to high 138 - > availability purposes, to heterogeneous integration 139 - > purposes, and now to kind of you know streaming intelligence uh 140 - > purposes.
141 - > So it's been fantastic. 142 - > Um and I got to spend some time at a company called Golden Gate, 143 - > uh um which was acquired by Oracle. 144 - > Uh so uh got my hands dirty in not just sort of the replication 145 - > technology, but also in data integration, data quality, those 146 - > those pieces. 147 - > Um, and then um have been busy with uh with Stream uh for the 148 - > last decade, uh, where we're trying to build a unified 149 - > platform that brings together uh you know real-time integration, 150 - > real-time replication, and real-time streaming uh for 151 - > analytics and AI reasons.
152 - > So it's been it's been fascinating and excited to be on 153 - > the show. 154 - > SPEAKER_02: Also, just great to hear how we've evolved from 155 - > using logs for redo and recovery purposes on a on a single 156 - > database. 157 - > And if you if you take that further, you can use it for 158 - > heterogeneous replication where you know you can take those 159 - > database logs and and use it to replicate data to multiple 160 - > different types of databases. 161 - > Obviously, lots of challenges there uh that that you work 162 - > through.
163 - > And then can also be applied for uh you know distributed uh 164 - > compute uh use cases and a lot of the the address a lot of the 165 - > requirements for the modern uh cloud analytics and now AI use 166 - > cases, which we'll we'll get into why uh replication is 167 - > critical for for uh model context protocol specifically. 168 - > Uh so excited to get into that. 169 - > Uh but both of you, you know, you you you know, you uh the the 170 - > great part about you know speaking with both of you is 171 - > that the amount of experience in databases and different types of 172 - > databases.
173 - > And I wanted to ask this first question. 174 - > Uh Sharish, if you want to take the first one. 175 - > Um, you know, when you think about the right database for the 176 - > job in 2025, knowing all the challenges that are sort of 177 - > ahead of us now with the known unknowns and the unknown 178 - > unknowns that are sort of creeping in the background, I 179 - > think you see a lot of consensus towards like using popular open 180 - > source databases like Postgres, but there's also other great 181 - > offerings.
182 - > So I want to ask you, you know, how do you advise teams to pick 183 - > the right database for the job? 184 - > SPEAKER_01: This is a great question. 185 - > And I, you know, you can think of it in two different ways. 186 - > On one end, you definitely want to have unification.
187 - > You don't want to have too much of fragmentation, you want to 188 - > have fewer choices and just invest in it to get it uh right. 189 - > And then on the other hand, you want to think about best to 190 - > breed and what exactly is the right tool for the right job, 191 - > um, that approach. 192 - > We may have taken it a little too far uh in the DBMS market, 193 - > where you, you know, if you think about Stack Overflow, 194 - > there are around 400 databases or something like that.
195 - > Um, I don't know if the world needs 400 databases, uh, to be 196 - > frank. 197 - > Now, and you you mentioned Postgres. 198 - > Postgres is definitely having a real moment. 199 - > Um, you know, in 2025, it's the honestly, it is the default 200 - > answer for a huge range of apps.
201 - > Developers love it, the ecosystem is thriving. 202 - > Um, and cloud providers, Azure included, of course, you know, 203 - > we've made it easier to spin up Postgres uh sooner and easier to 204 - > manage, easier to upgrade, and all that. 205 - > Um, so if you're building a SaaS app prototyping quickly and need 206 - > a uh reliable transactional store with a vibrant sort of 207 - > ecosystem in Postgres, like in the form of extensions, um, just 208 - > use Postgres does have some uh meaning to it.
209 - > Um, you get good stuff, SQL standards, like different kinds 210 - > of data types, you know, all the latest and the greatest 211 - > indexing, including vector indexing these days. 212 - > Um great community that keeps it moving forward. 213 - > All the hyperscalers believe in it. 214 - > So um so I buy that argument.
215 - > Having said that, though, I will have a lot of caveats. 216 - > Um we are, you know, obviously uh definitely want to be very 217 - > clear that Microsoft believes in it. 218 - > We are investing deeply. 219 - > In fact, you know, amongst the hyperscalers, uh, we are uh 220 - > we're one of the, in fact, the top committer into the Postgres 221 - > source code.
222 - > Um we've been committing quite a bit. 223 - > We've been adding a lot of features in Postgres 17, we 224 - > added a lot, 18, you know, we've done quite a bit. 225 - > Um, and we'll continue to do that. 226 - > Um, but back to your point about the you know, the right uh 227 - > database for the right job, um, there are limits in terms of any 228 - > engine design.
229 - > Like, you know, there's nobody, no database engine can be really 230 - > designed in a way that it can capture all the workloads. 231 - > And it's always been the case, it will be the case for a long 232 - > time. 233 - > Database design is a non-trivial exercise, so you do have to make 234 - > some choices. 235 - > So whenever you make those choices, you'd have uh some 236 - > gaps.
237 - > Um, you know, there are databases like SQL Server for 238 - > instance. 239 - > Uh SQL Server is designed to be a general purpose, amazing 240 - > relational database. 241 - > So you could do uh through and through relational transactional 242 - > systems, but you could also do SMP kind of workloads in SQL 243 - > very well. 244 - > Um, on the other hand, you know, there have been lots of other 245 - > scenarios where you have things such as um mission critical, 246 - > like you know, there's been an in-memory for some time, column 247 - > store for some time.
248 - > There are lots of advanced security integration challenges, 249 - > uh, analytics integrations, et cetera. 250 - > And these are some other places where SQL really shines very 251 - > well. 252 - > It's optimized for compliance, high availability, uh advanced 253 - > security features, like we have you know, all this encrypted for 254 - > a while, uh I'm just as an example. 255 - > There are lots of other Q related improvements in SQL, 256 - > which are really world-class.
257 - > Um, so you know, it's not necessarily just that Postgres 258 - > is the only answer. 259 - > Uh, looking at the workload and deciding is important. 260 - > Uh, and there are also limits in terms of how far a relational 261 - > database can go. 262 - > You know, for that matter, if you are looking for something 263 - > that has uh a world-class RPO RTO characteristics, if you want 264 - > something internet scale, user-facing, then Cosmos DB is a 265 - > great choice.
266 - > You know, it handles multi-region rights for sort of 267 - > like RTO zero characteristics. 268 - > It has single-digit millisecond latency guarantees with a four 269 - > nines HLA guarantees in a single region, five nines with 270 - > multi-region uh lets you tune your consistency. 271 - > And there are many apps. 272 - > Um, of course, you know, one of the most prolific apps of the 273 - > day, uh, ChatGPT, uses Cosmos DB for that very reason.
274 - > Um you know, Postgres wouldn't be able to do that kind of work 275 - > to it. 276 - > It's not just designed for that, right? 277 - > So there's that. 278 - > Then the final piece that I would say is that um, you know, 279 - > this is a is an important category that was quite the rage 280 - > a few years ago, vector databases, right?
281 - > Uh and my view, uh my team's view has always been that we 282 - > don't need a separate specialized engines to do uh 283 - > vector databases. 284 - > You should be able to sort of evolve it in conjunction with 285 - > your operational systems. 286 - > And Postgres is a great example. 287 - > It has strong vector extensions, um, and you know, it covers most 288 - > of all the rag scenarios, graph rag scenarios, you without 289 - > having to move the data from a transactional system to another 290 - > kind of a system.
291 - > So um, yeah, dedicated vector databases will have their place 292 - > in terms of some extreme scale or extreme design point, some 293 - > exotic sort of scenario. 294 - > So I'm pretty sure there is always a scenario. 295 - > Um, I don't want to say that there isn't any, but for the 296 - > vast majority of the enterprises, um, the complexity 297 - > outweighs the benefits. 298 - > And so, you know, keeping them all together in one place and 299 - > having fewer choices uh is is is a is a good good way to go 300 - > about, but but I also don't want to subscribe to the um to the 301 - > dogma of like you know it has to be only one uh or none.
302 - > Um so both points are extreme, uh, but I don't think you know 303 - > there's a place for 400 different databases either. 304 - > SPEAKER_02: Uh and and fast forward to that question, you 305 - > know, how is AI changing your your business, your business and 306 - > in terms of you know, uh looking at the entire database landscape 307 - > and and and offerings, uh including examples of, you know, 308 - > you mentioned Chat GPT running on Cosmos DB, but you know, 309 - > other examples of how you're applying like the roadmap of AI 310 - > innovation to databases.
311 - > SPEAKER_01: In a big way, you know, a huge, huge way. 312 - > So obviously, um a few patterns that have emerged, and there's a 313 - > lot of frothiness in this space, clearly. 314 - > So like dust has to settle down a little bit, but there's 315 - > clarity. 316 - > Uh there's clarity on a certain aspects, and there are a few 317 - > things that are still evolving.
318 - > Um, the way I think about it is that one of the best the best 319 - > ways, simplest ways that I explain it to my teams is that 320 - > there are, you know, there are vitamins and there are 321 - > painkillers in this space. 322 - > Uh you absolutely need rag pattern because a lot of the 323 - > enterprise system of record information will be stored in 324 - > operational databases. 325 - > And that's never going to change, uh, right. 326 - > And LLMs are not going to be trained on those kinds of data 327 - > points because those are by design confidential and private 328 - > and secured, which essentially kind of comes back to the whole 329 - > challenge of how do you really get the entire value of LLMs, uh 330 - > semantic searches.
331 - > Well, you have to marry the foundational knowledge of the 332 - > training with that of some of the system of record uh sort of 333 - > engagements that you typically go through operational systems, 334 - > uh, and even analytic systems or whatever databases effectively. 335 - > So rag pattern is real, it's here to stay, and we'd have to 336 - > do everything we can to really support it. 337 - > The way people search data in databases is changing. 338 - > Um, so natural d vector indexing and trying to uh if we were to 339 - > design like clause uh in the modern era, it wouldn't be like 340 - > a regex-based like clause, right?
341 - > We would have done something like semantic like clause, but 342 - > we ended up with cosine similarities and you know a very 343 - > interesting looking uh syntax. 344 - > Nonetheless, uh the the search by meaning, not search by regex 345 - > or exact predicates, et cetera, that's going to take off and 346 - > it's already happening. 347 - > We are seeing that. 348 - > And it has some really interesting implications in 349 - > terms of how people think about searching their data, attaching 350 - > it to their AI apps and all that stuff.
351 - > Uh, and in particular, if you think about some of the major 352 - > problems that databases always had, um, and this is kind of 353 - > really where, so far, what I've said, I kind of think of them as 354 - > vitamins, the painkillers. 355 - > And you know, it's very important to think about the 356 - > painkillers as well. 357 - > The core cohort of databases, uh, either DBAs, develop data 358 - > developers, people who spend all their time day to day managing 359 - > these enterprise, really big applications, they really need 360 - > to understand the database schema, need to understand the 361 - > performance, they need to keep on tuning it.
362 - > Uh, they cannot really uh, you know, they you can't design a 363 - > database and forget about it, right? 364 - > Um, so how can AI really help them in their day-to-day jobs? 365 - > A very classic example that's something that you know we are 366 - > very focused on at Microsoft, is uh help them chat with their 367 - > query plans, right? 368 - > Not just do something simple on the periphery, but go deeper and 369 - > help them like solve their pain.
370 - > It's not just a vitamin, it's a painkiller, right? 371 - > Um, you may have query plans that are very deep, very 372 - > complex. 373 - > Uh, how do you really enable them to go chat with them 374 - > without having to go through all these graphic trees? 375 - > You need to keep, you know, there are many queries which the 376 - > tree itself is really hard to navigate.
377 - > Um, could you make it easy for them to in a natural language 378 - > chat with these uh and go deeper depending on their expertise, 379 - > depending on their problem? 380 - > We are not far away from a point where the data the these 381 - > developers, DBAs can come in and ask the question, hey, why is my 382 - > database slow today? 383 - > And it can give a simple answer, like, hey, you're missing an 384 - > index or some, you know, your queries have changed and this 385 - > this needs to happen.
386 - > Or it can go all the way, like, hey, I see these wait stats and 387 - > like you know, these deadlocks are happening here, or maybe 388 - > something's really happening. 389 - > Um, your write patterns have changed. 390 - > You go really, really deep, or it could say that, you know, 391 - > maybe it's time for us to look at your query plans for these 392 - > queries. 393 - > Uh, some things have changed here.
394 - > Let's go look into it. 395 - > You could go deeper and deeper, and it can be like a 396 - > conversation, doesn't have to really require some really deep 397 - > expertise. 398 - > So, those are a few examples, and I haven't gone touched on 399 - > some many other things that you could do. 400 - > You could rewrite queries to be more efficient, you could do a 401 - > lot of things.
402 - > Um, you could certainly, you know, the simple example is 403 - > natural language to query language, where people want to 404 - > just give you a prompt and then get a uh get get get just a full 405 - > SQL query. 406 - > Those things are real. 407 - > And and it's not just about the engine query, et cetera. 408 - > It could be uh cost governance, it could be sort of resiliency 409 - > governance, it could be security management, et cetera.
410 - > So yeah, it's truly pervasive. 411 - > It definitely captures all these scenarios. 412 - > Um, and you know, on on the other hand, I also want to point 413 - > out again, as as you were referring to, um, some of the 414 - > biggest workloads of the of the day. 415 - > Uh ChatGPT, of course, it's it's massive, it keeps growing 416 - > significantly.
417 - > Uh they all rely on our databases. 418 - > In particular, ChatGPT uses Cosmos DB, all the user-facing 419 - > applications of ChatGPT. 420 - > Every time you you interact, any message uh operationally it goes 421 - > through Cosmos DB, all the user-facing apps. 422 - > We also use Postgres SQL in that space.
423 - > Um it it that's sort of like the user-facing thing. 424 - > I haven't even touched on how we do our development and how we 425 - > think about using AI. 426 - > Um, but you know, suffice it to say that it has dramatically 427 - > changed in the past uh several months. 428 - > SPEAKER_02: Yeah, absolutely.
429 - > It's it's it's changing very fast. 430 - > Like that's and then that's usually the main takeaway from 431 - > every uh episode that you know we have folks talk about AI is 432 - > just the the rate at of of change is just remarkable and 433 - > something we didn't see in the the 2000s or 2010s. 434 - > Uh Alok, I wanted to ask you a similar question. 435 - > Obviously, you know, we we talked about uh vector databases 436 - > a few years ago, you know, stream, we you know, we're 437 - > focusing on database replication and data streaming.
438 - > Uh, and obviously vector databases came up when there was 439 - > like a hype cycle for it. 440 - > And you had a similar view, which is sort of like this is 441 - > more like an index type rather than you know a full class of 442 - > databases that we needed to support. 443 - > So, and and now a few years later, we've seen that play out 444 - > where you know vector extensions for the mainstream databases 445 - > seem to be getting the the lion share of a of adoption there.
446 - > But I also want to ask you because you know, you by working 447 - > in change data capture, getting into the logs of the databases 448 - > when you truly know how the database works, so you have a 449 - > unique perspective on you know how to choose the right database 450 - > or the right job in 2025 and beyond, you know, given all the 451 - > AI transformation we're seeing. 452 - > SPEAKER_00: Yeah, um, I mean it's a great question. 453 - > Uh, and I think Shirish kind of covered it uh, you know, fairly 454 - > in detail, right?
455 - > I I do think that, you know, one size kind of doesn't fit all. 456 - > Um, you know, and and that's something in the database 457 - > community, you know, we've published multiple papers on 458 - > that and debated it uh far and wide. 459 - > I I personally uh still believe that to be very true. 460 - > You know, there are some cases, uh, John, where it makes sense 461 - > to kind of invest the energy and the effort and the resources to 462 - > address a very specific workload, right?
463 - > But then, you know, the other aspect of it is treating that as 464 - > more of a general workload for which there is an engine. 465 - > So I think databases, you know, have come a long way. 466 - > And I do think at this stage, take a vast look at, you know, 467 - > maybe largely from the customers that we interact with, how do 468 - > they think about it? 469 - > I think that's kind of an interesting question.
470 - > On one hand, they have like sort of as their legacy workloads, 471 - > right? 472 - > And they're sort of invested in that. 473 - > Um, and so on one hand, I think what happens is because of that 474 - > legacy investment and some of the engine choices and the 475 - > feature choices that have been made maybe 20, 30 years ago, 476 - > they're not able to change and evolve that very fast, right? 477 - > So that's kind of like one thing that I see.
478 - > So that begs the question that what do I do if I really want 479 - > something fast? 480 - > I want, you know, some new uh type of uh MFA, or I want to 481 - > introduce some new security uh level uh feature, which 482 - > otherwise wasn't thought through in the original design and so 483 - > forth. 484 - > So um I do see them sort of then saying, hey, how can I actually 485 - > introduce these newer services and newer workloads, newer 486 - > applications, but I don't have to be sort of limited to the 487 - > choice that I made a while ago, right?
488 - > So that introduces kind of this next aspect of my choices, um, 489 - > which says, oh, maybe I can actually now uh go in and you 490 - > know take a look at the best of breed for the stuff that I'm 491 - > doing. 492 - > And we and I'm seeing I think that a lot, right? 493 - > So I think just to just to kind of uh add on to Sharish's point, 494 - > like you know, you may think that, hey, this specific 495 - > workload is best suited on Cosmos DB on Azure, right? 496 - > So I'm going to actually go in and for a lot of this uh data, 497 - > and we have seen this in some you know manufacturing areas and 498 - > so forth, right?
499 - > They're like I the this this is just like sheer petabytes of 500 - > data, and I want to like just you know push this um in a 501 - > scalable way. 502 - > Um so we so we are we are seeing kind of that as the second part 503 - > of it. 504 - > And then there's hey, I'm a new uh kind of startup, I'm a new 505 - > young uh company, and so what are my choices? 506 - > Um and that's where I think cost becomes also an interesting 507 - > part.
508 - > And I do see that many of them will naturally gravitate towards 509 - > flexibility, open source, uh relying on the community for 510 - > support. 511 - > And that's where you know choices like Postgres make a lot 512 - > of sense. 513 - > And I and at stream ourselves, when we had to make that choice, 514 - > right? 515 - > Right, you know, uh very earlier on, we said, well, you know, 516 - > let's go uh and we looked at JavaDB uh and we looked at 517 - > Postgres, right?
518 - > And we started off uh and then as we evolved, you know, then 519 - > customers kind of came in and said, Well, look, I'm already 520 - > running, you know, one standard potentially for my mission 521 - > critical workloads. 522 - > Uh, could you guys also support that? 523 - > Right. 524 - > So then we started moving into kind of you know supporting 525 - > additional RDBMS engines, uh, but we started off with 526 - > Postgres.
527 - > That was kind of the main point, right? 528 - > I don't think there's like a one size fits all uh at all. 529 - > Um I am seeing a lot of uh shift, I think, in terms of the 530 - > movement towards newer types of engines, especially uh on the 531 - > cloud side. 532 - > I think this attitude that I'm going to manage kind of like you 533 - > know my own schemas and my own tables and worry about the space 534 - > management, worry about kind of the backup restore part of it.
535 - > I think that is is becoming somewhat kind of outdated. 536 - > I think I do see even very, very strict customers in the 537 - > financial arena and so forth now changing their eight, nine years 538 - > ago, they were like, no, we'll we'll never get to the cloud. 539 - > And now, you know, you would almost uh that's a joke, right? 540 - > I mean, I think you'd probably not be in your role very long if 541 - > you have that kind of an attitude.
542 - > And I think so that that is a strong shift. 543 - > So we are seeing a lot of workloads sort of now, you know, 544 - > either migrate or or or split uh across sort of like the legacy 545 - > and the newer types of systems. 546 - > So I think again, the no no magic bullet here, but I do 547 - > think it really depends on sort of what journey are you in at 548 - > the time that are you inheriting a workload or are you creating a 549 - > workload from scratch? 550 - > To you know, what are the cost choices that come into play?
551 - > Um, and then finally, you know, you know, is this are the 552 - > existing systems not adequate where I need to go invent my 553 - > own, right? 554 - > I think so those are, and which are in the minority at this 555 - > point, I would say. 556 - > SPEAKER_02: So that was a long-winded answer, but you 557 - > know, John, hopefully I address, you know, you know, absolutely, 558 - > and you know, having the the broad kind of reactive view of 559 - > what's what's happening right now and and and what changes 560 - > we're gonna make to um to to to support all the dynamic shifts 561 - > in the market, but also kind of tying it back to okay, you know, 562 - > databases have been around for decades and they work the way 563 - > they do for a reason.
564 - > I mean, it's incredibly valuable to have that uh that perspective 565 - > on top of that. 566 - > Uh so I think that's that's that's always gonna be valuable 567 - > for uh for executives and and engineering leaders to apply 568 - > that in their in their thinking when they do choose, when they 569 - > do architect their next generation uh applications. 570 - > So speaking of next generation patterns, I'm I'm I'm gonna ask 571 - > about model context protocol, uh specifically for databases, just 572 - > to define it quickly, model context protocols, a really a 573 - > simple wrapper for LLMs to interact with APIs for uh for 574 - > either interacting with applications or retrieval from 575 - > uh warehouses or databases.
576 - > Sharish, I wanted to ask you uh what for MCP for databases 577 - > specifically, what value and patterns are you seeing from 578 - > early AI applications that are using it? 579 - > SPEAKER_01: Yeah, I so I do think that MCP has um real 580 - > potential to become the missing bridge between the large 581 - > language models and enterprise databases. 582 - > There are certain ways to get the large language models 583 - > interoperate with enterprise databases, but it requires a 584 - > very intended, intentional uh sort of access pattern from the 585 - > user without MCP.
586 - > Uh and there's just a lot of like sort of repetition in terms 587 - > of the implementations. 588 - > So you know the tension has always been how do you let an 589 - > LLM understand your data structures, policies without 590 - > actually giving away the keys to the world, right? 591 - > Uh MCP solves that. 592 - > And apart from all the work that you need to do, sort of every 593 - > server needs to be implemented in a different way, uh in a 594 - > unique way, in a bespoken fashion.
595 - > MCP solves that by creating a standardized sort of like a 596 - > policy aware layer between the model and the databases. 597 - > Um and instead of like the model scraping schema or free forming 598 - > SQL, uh, the server on the database side, in this case, can 599 - > expose um what is safe and what is like sort of clear metadata 600 - > information, could be schema fragments, curated tools, 601 - > whatever. 602 - > They can all be put together. 603 - > Um and they can be most importantly, uh can be filtered 604 - > by permissions, governance, and cost awareness.
605 - > And that's important because the filter really can help you in a 606 - > big way. 607 - > Um if you kind of zone in on the security part, it's essential 608 - > because the principle of least privilege can be applied to LLM 609 - > database interaction. 610 - > Um, and that can safely unlock the LLM-driven apps while 611 - > keeping compliance and trust and those kinds of things. 612 - > So, you know, some of the patterns that come up when I 613 - > think about the AA apps using MCP uh with our databases.
614 - > Uh so there are a lot of folks who are thinking about 615 - > policy-aware introspection. 616 - > Uh, models don't need to see the full catalog, don't even see it. 617 - > They get a sort of business-friendly, least 618 - > privileged kind of a view with sensitive fields, masked and 619 - > like with row-level rules baked in so that the the applications 620 - > models basically can take advantage of the data and 621 - > databases safely. 622 - > Um, there's also like a constrained execution.
623 - > Uh, you don't want to just instead of running anything, the 624 - > model can have uh higher level verbs that they can get um in 625 - > terms of let's say, hey, get me some customer metrics or like 626 - > explain this query or whatever. 627 - > The MCP server can then translate it, validate, and 628 - > force uh, how do you really, you know, just give the data that 629 - > the model wants? 630 - > Um, and filtering really is important. 631 - > Um, yeah, and you know, finally, I would also say there's like 632 - > auditable sessions.
633 - > Um, access is like time boxed, read-only by default, every 634 - > action is logged. 635 - > So that's also important. 636 - > But just generally the power of MCP is immense, but you need to 637 - > be very careful in terms of uh how you're auditing it, what are 638 - > the tools, etc. 639 - > There have been lots of cases in the recent times where some of 640 - > these LLM models, apps, et cetera, can go and do quite a 641 - > lot of damage to your database.
642 - > So you've got to be very careful about it. 643 - > Um, but once you've taken care of that, then uh you know the 644 - > design patterns in terms of like building a semantic layer, um 645 - > all the easier ways that the models can give to LLMs, those 646 - > are very powerful. 647 - > Definitely. 648 - > Definitely helps them for LLMs to be deeply database aware, 649 - > understand all the details without without again 650 - > compromising on safety or compliance.
651 - > That's an important shift. 652 - > Use the power carefully, but there's a lot of power here. 653 - > SPEAKER_02: Absolutely. 654 - > And and Alok on your side, you know, what's really required to 655 - > make model context protocol work for databases in a way that 656 - > enterprises can deploy this, not only with the right scale to 657 - > handle their workloads and also the right governance.
658 - > SPEAKER_00: Yeah. 659 - > I mean, great question. 660 - > And again, look, I think number one, uh it has significantly 661 - > simplified how you're exposing, you know, all kinds of engines 662 - > now. 663 - > Um, you know, you don't have to sort of in a proprietary way 664 - > redesign every single time you're kind of you know uh 665 - > working with a newer type of uh of an application, uh APIs and 666 - > uh and engines, right?
667 - > So MCP is super useful for that. 668 - > I think in terms of databases, I think some of the things that uh 669 - > that Sharish mentioned, right? 670 - > Um, around right from uh privileges and roles uh to kind 671 - > of the access to the data itself, um, these are 672 - > significant constructs that always seem to you know you have 673 - > you you must address. 674 - > Now I view this as just yet one more workload, right?
675 - > Um like agents coming in and then over a protocol 676 - > conversationally trying to just simplify what all of us really 677 - > are trying to do at the end of the day, which is, you know, 678 - > let's okay, you can think of it in queries and engines and query 679 - > plans and optimizers, but fundamentally I'm trying to 680 - > answer a few questions that currently are are interesting 681 - > for me for let's say operational analytic reasons. 682 - > So the conversational style is key, right?
683 - > So I do think that that's the power here. 684 - > Now, as you start having conversations, we've seen this 685 - > in human conversations also. 686 - > Some people are super interested, really good at 687 - > probing uh and getting information out, and some of 688 - > them are not so good, right? 689 - > And so with the right kind of agent, you could leak 690 - > information, you they could get into some of the you know uh 691 - > sort of not so protected parts of your database and cause havoc 692 - > and damage.
693 - > So, one of the ways in which I think um just like a traditional 694 - > workload, right? 695 - > And you Shirish mentioned the word read-only, right? 696 - > So that's why I kind of went down this path. 697 - > We've seen in the past where there was uncontrolled um taxing 698 - > of resources on a production system.
699 - > And, you know, being a database uh person, you know, an easy way 700 - > to address that is to say, well, why don't I actually try to have 701 - > maybe a server farm where you know I could have the rights 702 - > being routed in to one location, and then I have sort of a 703 - > replica form. 704 - > As long as the latency is good enough for the questions, uh, 705 - > you know, it's it's great, right? 706 - > So I think that's one of the ways in which one of the 707 - > patterns that's emerging um or is to expose kind of a database 708 - > engine you know over MCP.
709 - > Um, and then uh so let me take that back, database engine you 710 - > know, with some replicas, and then now having one of the 711 - > replicas being exposed over MCP to an agent. 712 - > So that's kind of your first go-to. 713 - > And then if something really warrants kind of an actual 714 - > transaction where it's not conversation, but I'm actually 715 - > making a transaction, then I think there's a rerouting back 716 - > to back to kind of the production system. 717 - > So I think that is going to be one of the interesting patterns 718 - > here.
719 - > Time will tell whether we'll all move in that direction. 720 - > But we saw that implemented very widely globally for some of the 721 - > like most critical workloads. 722 - > And agents kind of spawning more agents and having that kind of 723 - > uncontrolled, that does put a lot of load onto the production 724 - > system. 725 - > So, how do we mitigate against that?
726 - > Um, I think that's where kind of some of these new patterns are 727 - > are going to uh to emerge. 728 - > And at least and specifically at Stream, obviously, we're we're 729 - > we're right in there, right? 730 - > Where we are trying to say, hey, you know, we want to make the 731 - > actual intelligence that's exposed by not just one engine, 732 - > but by a combination of engines simpler. 733 - > And that's sort of like where we think MCP can be used to take 734 - > portions of a database and then cleanse parts of it for 735 - > governance and et cetera, and then expose that via maybe 736 - > another system.
737 - > And that could be a low-cost system. 738 - > It doesn't have to be the same cost as the original system. 739 - > And now you can actually have an interaction over MCP with that 740 - > system. 741 - > And this is sort of some of the ways in which you can try to 742 - > protect uh, you know, some of the challenges that Sharish 743 - > talked about, as well as you know, try to broadly just solve 744 - > like a distributed replication, distributed database problem, 745 - > you know, over MCP.
746 - > So that's kind of like those are some of the early thoughts in 747 - > the in this area. 748 - > Absolutely. 749 - > SPEAKER_02: And you see these patterns for real-world agent 750 - > deployments. 751 - > And it's so clear that the key to success there is in the agent 752 - > implementation, giving, creating these deep agents that have 753 - > autonomy to solve the business use case and have the right 754 - > layer of indirection where you have the infrastructure to be uh 755 - > set up to handle these deep agents and any subagents that 756 - > are created in the process to solve these problems.
757 - > So, you know, we talk about the popular use cases like customer 758 - > support uh agents that can go query the transactional data, 759 - > maybe make, you know, like you were saying, Oloc, actual 760 - > transactions on the database itself. 761 - > But doing that in a in a protected manner through read 762 - > replicas, which minimizes latency and minimizes risk. 763 - > Because you don't want the agent implementation to be concerned 764 - > about governance, right? 765 - > You want the governance to be solved outside of that and let 766 - > the agents do what they're going to do, but have the right 767 - > guardrails to effectively protect uh you know any any 768 - > issues that could cause.
769 - > So I do think this is one of those situations where the um 770 - > the innovation is certainly new with uh with uh model context 771 - > protocol, but the patterns are are very similar to having 772 - > analytical read replicas, right? 773 - > So uh the the actual infrastructure that teams are 774 - > are gonna roll out here is very similar to what uh how database 775 - > architects have addressed this in the past. 776 - > Like you said, Alog, it's it's yet another workload.
777 - > And how would you handle another workload? 778 - > Well, uh, you know, there's there's there's ways to do this. 779 - > Um so very exciting. 780 - > I think there's I I think this is a much more nuanced takeaway 781 - > for for the listeners.
782 - > I know that when when Model Context Protocol first came out, 783 - > you know, if you looked at the like the official docs for it, 784 - > or at least one of the sort of docs on it, I think the 785 - > suggestion was just to have a uh a read-only role for for your 786 - > agents on on a on a production Postgres database. 787 - > And people obviously ran into the obvious issues that can stem 788 - > from that. 789 - > Um, you have to be a little bit uh go a little uh deeper into 790 - > that problem to solve it effectively.
791 - > So I I I do think the listeners can can take away uh some some 792 - > really actionable uh design patterns from from uh what you 793 - > both shared. 794 - > So thank you. 795 - > I wanted to move on to uh to Postgres and and chat about that 796 - > a bit more. 797 - > It's it's always a very popular topic on on this show.
798 - > Uh we've had several guests who've who've gone deep into uh 799 - > uh really unique uh cloud scale Postgres implementations. 800 - > Sharish, first I wanted to ask you, I think Postgres is a 801 - > superpower's extensibility and the community behind it. 802 - > You know, what are some of the big architectural bets you would 803 - > like to see the community rally around in in Postgres going 804 - > forward? 805 - > SPEAKER_01: Yeah, so you know, I totally agree that the power, 806 - > superpower for Postgres is definitely extensibility.
807 - > It's extensibility. 808 - > And you could add new data types, new indexing methods, 809 - > even runtimes. 810 - > And that's why it's as popular as it is today for many 811 - > developers. 812 - > Um the shadow side of that is complexity, as you said, right?
813 - > Uh it is definitely something you can end up with. 814 - > Um you can easily end up with a patchwork of extensions in your 815 - > workloads. 816 - > Um, they may not really all interoperate in the way that you 817 - > want it to be, uh, in the in the cleanest way. 818 - > Uh, and that can make it harder for Postgres to scale in 819 - > production.
820 - > Um, so you know, to your question about what are the few 821 - > things that the architecture best that I'd like to see the 822 - > community rally around. 823 - > And by the way, we as Microsoft are invested deeply. 824 - > Uh, we ship many extensions ourselves. 825 - > Um, but I I a couple of things that I really like to see.
826 - > Firstly, I think it'd be good to see like a unified story for 827 - > scale out. 828 - > Um, we have, you know, we we have a few solutions uh uh and 829 - > we invest in CITES, we really believe in it, et cetera. 830 - > And we're trying to push it, make it easier for developers to 831 - > shard, not just on the uh uh a role level or a table level, 832 - > even schema level. 833 - > So we have different kinds of abilities to short, but um 834 - > having a native consistent way to scale a workload out uh would 835 - > really make Postgres uh uh the vision of sort of being a 836 - > default database or whatever.
837 - > I mean, again, I don't believe in being a default database, but 838 - > to the extent that developers want to keep it as a default for 839 - > most scenarios, um, that idea of unified scale out, and I know 840 - > different people have taken different uh paths here. 841 - > Uh so that's like you know, that's really important, I would 842 - > say. 843 - > The second thing, back to extensions, I do think that you 844 - > know, standardizing uh the packaging and lifecycling, I 845 - > know there's this work that's happening here, but installing 846 - > and upgrading extensions, it's a bit of wild west.
847 - > Um if if there is a if there's a way that the community could 848 - > converge on a model where extensions are like clearly 849 - > versioned, certified, they pay all well, play all well 850 - > together. 851 - > Uh that'll help the community. 852 - > It will definitely unlock confidence for enterprises. 853 - > You know, you're looking at adopting Postgres.
854 - > Um, it's one of the challenges, right? 855 - > So you could really go solve that. 856 - > Um then I would say, you know, the third one, I I think there's 857 - > a lot of work on AI primitives in Postgres, but I would love to 858 - > see like them done in the Postgres way in a native way. 859 - > Uh and this is where like, you know, we are definitely leaning 860 - > in in terms of bringing in some of the cool things like disk 861 - > ANN.
862 - > Um I know there's PG vector extensions, which is very 863 - > popular. 864 - > Uh, but things around vector search, embeddings, retrievals, 865 - > etc. 866 - > I think the community needs to spend a little bit more time on 867 - > agreeing on standard operators, indexing strategies, governance, 868 - > and how they really operate with the runtime capabilities 869 - > outside. 870 - > Um, there are a lot of you know competing sort of ways to do a 871 - > few things, and it's not it needs to be deeply thought 872 - > through.
873 - > Uh, more love is needed there. 874 - > Uh, I think those are a few things. 875 - > Maybe, maybe finally, I would just add one more. 876 - > Um I think operability as a first class concern uh is an 877 - > important thing for Postgres.
878 - > Uh, given the amount of engagement that it is seeing, 879 - > uh, the workload isolation, resource governance, uh, those 880 - > are the kinds of things that are really important when you are 881 - > trying to deploy this in uh at scale, uh especially in cloud, 882 - > the core and the extension ecosystem, they need to embrace 883 - > this thing natively. 884 - > Uh I think that'll really go a long way. 885 - > It'll reduce the cognitive load for anybody running at scale.
886 - > So Postgres has already won on extensibility. 887 - > I don't think there's anything, there's nobody debating about 888 - > it. 889 - > Um, but I think the next frontier I would say is about 890 - > coherence. 891 - > It's about really making sure that you scale it up to the 892 - > enterprise grid, think about operability deeply, and make AI 893 - > like you know, first class instance and make the extensions 894 - > more certifiable, extensible, and that kind of stuff.
895 - > The community is just um it's it's a brilliant community. 896 - > They care about all these pieces very well, deeply. 897 - > Um, and you know, we are here at Microsoft certainly engaging 898 - > with them, uh, contributing. 899 - > We do have uh very deeply wetted, deeply embedded sort of 900 - > uh contributors and commuters as part of our team.
901 - > Uh you know, we are very engaged and we look forward to 902 - > partnering with the community towards these goals. 903 - > SPEAKER_02: Absolutely. 904 - > And that's one of the most powerful parts is you know, 905 - > through community, you get standards and you have less uh 906 - > you know duplicative problem solving between companies. 907 - > Uh and Aloka, I wanted to ask you, you know, you're you're 908 - > also working deeply with uh Postgres through through through 909 - > stream and and database modernization and innovation 910 - > projects.
911 - > You know, what have you seen in terms of scale and adoption of 912 - > uh Postgres for uh for particular enterprise use cases? 913 - > SPEAKER_00: From a horizontal perspective, um, I think we're 914 - > seeing a lot of adoption of Postgres. 915 - > Um and let me just clarify what I mean by horizontal, because I 916 - > mean, with with specifically uh you guys, it can mean one thing 917 - > if you're in the engine level versus sort of like a like a 918 - > general level.
919 - > What I mean by that is um broadly from uh from an adoption 920 - > perspective, as folks are thinking about adopting AI and 921 - > analytics and newer workloads, they are thinking of spitting up 922 - > a lot of their existing um workloads and applications. 923 - > So there's this whole journey to get to the cloud. 924 - > Um, and what we are seeing is number one, the question that I 925 - > get asked most often is hey, look, right now I'm running this 926 - > on like a massive Oracle Rack cluster.
927 - > Is Postgres, is it mature? 928 - > Right? 929 - > Is it going to be able to actually scale to these 930 - > workloads that I'm running this very, very mission-critical 931 - > application on, right? 932 - > And awkward question for me because I don't want to take a 933 - > stance.
934 - > Uh ultimately I am um, but we are seeing that question come up 935 - > more and more. 936 - > Um and I do think that some of the very, very large enterprises 937 - > are making that bet, especially with the with the hyperscalers, 938 - > right? 939 - > When especially when they have their own flavor of Postgres. 940 - > Because then I think they do realize that, hey, if there's 941 - > something is not uh up to what you know they they need in terms 942 - > of SLAs, in terms of their their roadmap for the future, they can 943 - > go in and hold somebody accountable for that, right?
944 - > So that I think we're seeing a shift there. 945 - > Now, in terms of the adoption of Postgres, I can tell you at 946 - > Stream specifically, I mean, we've had massive 947 - > implementations where we've done, you know, literally within 948 - > a few weeks, thousands of actual uh migrations uh into Postgres 949 - > from you know on-premise types of systems. 950 - > And these are not sort of lift and shift migrations, mind you, 951 - > right? 952 - > These are actual what I call zero-downtime live type of uh 953 - > migrations, oftentimes with bi-directional links being set 954 - > up, because you don't want to give up on an existing 955 - > on-premise system.
956 - > You want to actually have a way to move over and pull the plug 957 - > at your own leisure, if so to speak, right? 958 - > You don't want to say, hey, I'm gonna keep my entire um partner 959 - > community, user community hostage and you know, December 960 - > 31st at midnight and hope for the best, right? 961 - > That I think those days are a little bit tough to swallow 962 - > nowadays. 963 - > So we are seeing that adoption.
964 - > The other piece that also we are seeing is sort of this from the 965 - > cloud to maintaining, and this is largely cost reasons as well 966 - > as some SLA concerns, where we are seeing a reverse flow of 967 - > limited portions of that data into what I call like you know, 968 - > managed systems in Postgres that customers are managing 969 - > themselves. 970 - > So that's where they could have a version that now they're 971 - > maintaining, um, but they just want to have that for maybe 972 - > let's say in a geographically distributed retail system, it 973 - > could be per online store.
974 - > There's a Postgres that's running. 975 - > And so they do want the reverse flow also coming in. 976 - > And if to that degree, I mean, this sort of should sound 977 - > familiar to your audience, going back to maybe operational to 978 - > data warehouses to data mart type of uh of a pattern. 979 - > Now we're seeing sort of like this you know legacy to cloud to 980 - > sort of like you know, back, which is sort of like this, it's 981 - > not quite a data mart, but I would say it's a data product, 982 - > right?
983 - > Where a team wants to actually get portions of that data, but 984 - > they don't want to necessarily hit just the the core system on 985 - > the cloud. 986 - > And that's where, again, from a scale perspective, literally 987 - > I've seen single customers adopt you know 10,000 to 15,000 988 - > Postgres databases, um, you know, and and they're they're 989 - > they're pushing data from the cloud-based systems onto these 990 - > uh these uh self-managed Postgres uh systems.
991 - > So I think I think it has a, and Postgres has come a long way, 992 - > right? 993 - > From uh I mean Sharish would know better, like I think from 994 - > version nine till now 17, 18 coming up and so forth. 995 - > I think right from, I mean, they've caught up on on 996 - > partitioning and on replication, on security, on uh, you know, 997 - > just uh just developer productivity, supporting basic 998 - > operations like upserts and merges and all kinds of uh JSON, 999 - > JSON B data types and the in the vector vector types.
1000 - > So I think they've made a lot of progress, but I do think I want 1001 - > to pick up on one point that Sharish made, which is on the 1002 - > mission critical part of it. 1003 - > I think customers do have have have concerns there right now, 1004 - > where the automated uh part of, for example, um maybe just 1005 - > moving over from a primary to cutting switching over to a 1006 - > standby and back, uh, or having like a cascading system where I 1007 - > do move to a standby, but then all the workloads that are were 1008 - > relying on the primary get adequate notification and it's 1009 - > seamless.
1010 - > I mean, and I personally deal with, for example, Postgres 1011 - > replication slots, right? 1012 - > So I was very happy to see failover slots, you know, so 1013 - > that at least you're you don't have to reinstantiate uh if you 1014 - > just go to a standby, you're able to take advantage of that. 1015 - > I was happy to see that there's support in these decoding layers 1016 - > for more than one gigabyte of uh of redo uh for if a transaction 1017 - > spans that. 1018 - > So I think I think we have come a long way.
1019 - > And um, as the future seems very bright, and customers are super 1020 - > interested in it. 1021 - > Kind of that's sort of my take on Postgres. 1022 - > SPEAKER_02: I always think that's a very important 1023 - > perspective for uh engineers and leaders to hear, because I think 1024 - > there's there's sort of two sides of the coin of hey, you 1025 - > know, you can try Postgres either in a lab or if you're an 1026 - > early stage startup scaling incrementally, you know, over 1027 - > you know, over the span of a year.
1028 - > That's completely different from an enterprise migrating an 1029 - > existing production workload uh that touches thousands of 1030 - > employees and millions of customers with lots of revenue 1031 - > on the line and lots of business risk, and and then migrating to 1032 - > a completely new database there. 1033 - > So like the the risk calculus is is is completely different, but 1034 - > it's also very promising to hear that there's certainly adoption 1035 - > and examples of wins and maturity uh there.
1036 - > So I want to ask about modernization broadly, and and 1037 - > maybe Sharish, we can start with with you here. 1038 - > So for teams that are modernizing data applications on 1039 - > Azure specifically, what's your practical playbook for success? 1040 - > SPEAKER_01: Yeah, it's a great question. 1041 - > You know, at Microsoft, we basically uh we we definitely uh 1042 - > have a lot of evidence and starting to see a developer 1043 - > story emerge around uh around you know the Azure AI and it 1044 - > comes down to unifying three levels that that I think used to 1045 - > be separate uh for a long time.
1046 - > Uh, there is a developer experience that we go after with 1047 - > Copilots and VS Core. 1048 - > There is the data foundation, which is where Fabric and Azure 1049 - > databases, Microsoft Fabric and Azure databases play. 1050 - > Um, and then there's an application runtime in between 1051 - > the developer experience and the data foundation, which is 1052 - > basically what we have with Foundry for agents and AI 1053 - > workflows. 1054 - > So that's sort of like the three-piece stack of Microsoft 1055 - > that we go with, and that's the the stack.
1056 - > Like there's a developer experience on top, application 1057 - > runtime, and then there is a data foundation. 1058 - > So that combination is powerful. 1059 - > Developers don't have to juggle like multiple tools anymore. 1060 - > They can just open the VS Code, use co-pilots to design prompts, 1061 - > queries, or really go deeper if you are a pro developer, or you 1062 - > could use like GitHub Spark for like, you know, just getting 1063 - > started with VibeApps.
1064 - > Um and then you you you get powerful extensions where you 1065 - > can see the governor, the data governed right there. 1066 - > You could push to production easily, uh, you could monitor, 1067 - > but you have all these key pieces from Foundry coming in to 1068 - > make that stack really, really helpful. 1069 - > Now, if you're a team modernizing data apps on Azure, 1070 - > specifically since that's your question, um, the playbook that 1071 - > I would generally use is like uh a couple of steps.
1072 - > The first step I would say is like think about the developer 1073 - > UX. 1074 - > You know, you you obviously tools, tools, tools. 1075 - > You gotta start there and you'll really make it easy. 1076 - > Uh, and in the world where there are so many tools to do a lot of 1077 - > these things, it's really important to standardize on a 1078 - > few things.
1079 - > I would suggest VS Code, GitHub, Copilot, Azure AI extensions, 1080 - > and such. 1081 - > You know, these are very well proven. 1082 - > Uh, the developers should test prompts, build workflows, debug 1083 - > database interactions in the IDEs, get confidence because it 1084 - > kind of really helps them. 1085 - > Um they have to get familiarized because it's a new era for 1086 - > everyone.
1087 - > You know, it's like honestly, 99% of the developers haven't uh 1088 - > ever built workflows in this way. 1089 - > Uh so you have to lower the barriers and build the momentum 1090 - > quickly. 1091 - > So betting on some of these tools like VS Code, GitHub 1092 - > Copilot, AI extensions in Azure, et cetera, really good starting 1093 - > point. 1094 - > The step two, I would say, is like you bake in your 1095 - > operational guardrails.
1096 - > Uh, and and this is like data in specific, I would say. 1097 - > Uh, this is where like fabric Azure databases with built-in 1098 - > governance, row-level security, workloadized, all those things, 1099 - > um, wrapping it with Azure policy, et cetera, comes really 1100 - > uh handy. 1101 - > Um every AI generated query, you gotta have a very clear path, 1102 - > you know, what is the MCP strategy or what is the 1103 - > governance strategy, what is the observability, et cetera.
1104 - > So those guardrails need to be very clearly defined, rate 1105 - > limiting, you know, from day one, for instance. 1106 - > So those those that is a step two, I would say. 1107 - > Then the step three, I would say, is that you've got to 1108 - > choose your workload wins uh that prove the point before you 1109 - > really go and scale. 1110 - > Um this is perhaps one of the most important mistakes that 1111 - > most of the customers do.
1112 - > Um, you wanna you wanna like really get something going, a 1113 - > clear pain point that you can go build and then establish your 1114 - > playbook. 1115 - > Uh we do see patterns like copilots for analytics teams 1116 - > that accelerate SQL and BN, Fabric, Agentic style apps in 1117 - > Foundry that automate things like customer support, IT ops, 1118 - > et cetera, with clear SLAs, um, high volume, cost-sensitive 1119 - > queries for databases. 1120 - > These are all a few workloads that you could get get going 1121 - > with, right?
1122 - > You sort of like win them, prove a point, uh, and then then scale 1123 - > about it. 1124 - > So I think the net effect that I'm trying to say here is that 1125 - > developers stay in their flow. 1126 - > Um, ops need to get the guardrails that they want. 1127 - > And then on the other hand, business leaders want to see a 1128 - > quick proof of performance, cost saving, reliability.
1129 - > You could totally achieve these three things that I touched on 1130 - > the developers, the ops, and the business leaders. 1131 - > Uh, so I think this is kind of really the playbook that I 1132 - > generally discuss with many of our customers. 1133 - > Uh, then you can go from experimentation to uh to true 1134 - > modernization on Azure um at a large scale. 1135 - > So we got all the tools, you know, we have all the things 1136 - > that you need at every step of the way.
1137 - > Um, but being very clear about the stack and then going these 1138 - > in in these orders can really simplify. 1139 - > I think that's a that's a good playbook and in my mind. 1140 - > SPEAKER_02: Alok, I wanted to ask you specifically, you know, 1141 - > we we we worked with so many uh companies that look at the data 1142 - > foundations that Sharish mentioned in the in the 1143 - > modernization process. 1144 - > What's your recommendation and advice strategies to to leaders 1145 - > who are going through major modernization?
1146 - > SPEAKER_00: I mean, I think from my perspective, I think the last 1147 - > point, right, about business leaders trying to do it 1148 - > piecemeal is very important. 1149 - > Because often I think that um, you know, they could take sort 1150 - > of like a like a, hey, let's just do this thing. 1151 - > And then if something doesn't work, then they start 1152 - > re-questioning and going back and saying, hey, was this the 1153 - > right choice or not? 1154 - > So, and I've seen multiple folks be successful, large uh 1155 - > companies that have literally 50 to 100,000 databases be 1156 - > successful in this is by identifying workloads that could 1157 - > take advantage uh of some of the capabilities, uh, let's say, um, 1158 - > on the cloud, either services or infrastructure, uh, storage, 1159 - > compute, uh, elasticity, scale.
1160 - > For and and so then the question is like, you know, what do what 1161 - > do we do with the data, right? 1162 - > So I think it's it's key to make sure that if you're trying to 1163 - > deliver a new contained workload in the cloud, you identify sort 1164 - > of like, hey, what is that workload, number one? 1165 - > And and second, what is the data that you're going to to need? 1166 - > And that's where uh at least my conversation start, right?
1167 - > That we're giving a lot of choice to these to these uh to 1168 - > these businesses to say, look, you could actually take subsets 1169 - > of this data and safely go in and try out uh and test out the 1170 - > the newer uh service, the new platform, the new 1171 - > infrastructure, storage system, compute, whatever, however, what 1172 - > whatever choices you made, but still allow your workloads to 1173 - > run there and see whether they meet your requirements and your 1174 - > SLAs or not.
1175 - > And I think that is what the modernization journey looks 1176 - > like, you know, to me, right? 1177 - > Because I would recommend that. 1178 - > That, you know, oftentimes if you do a lift and shift, I think 1179 - > it works for test and dev type of workloads, it may work fine, 1180 - > but for production, it's a little risky. 1181 - > And shutting down a business in sort of like 24 by seven global 1182 - > economies is no longer an option.
1183 - > So, what that looks like is to have a sort of a proper testing 1184 - > methodology towards getting there. 1185 - > And and I think loosely to be very specific, you could just 1186 - > think of this as either database or platform modernization or 1187 - > application modernization, and then having a path to that by 1188 - > saying I'm going to actually keep maybe concurrent systems, 1189 - > and then that gives me that choice. 1190 - > Uh, so I think that should be a key part of this.
1191 - > Um, along the way, I think there what we do see in modernization 1192 - > is a tool set that's needed because rarely do we see people 1193 - > say, okay, I'm running a workload, I'm just going to go 1194 - > lift and shift and run it somewhere else. 1195 - > They do want to take advantage along the same time to say, hey, 1196 - > all these choices we made 15, 20 years ago, maybe there was not 1197 - > um you know instant alerting available at that time, uh, you 1198 - > know, based on some trigger.
1199 - > Um, I do want to address that in my newer system. 1200 - > So along the way, do I want to actually now revisit uh some of 1201 - > the design choices for potentially my schemas? 1202 - > Um do I want to um, you know, go in and denormalize certain 1203 - > things or normalize certain things along the way, right? 1204 - > So now you're introducing sort of like the the requirements for 1205 - > a tool set that allows you to address this in more of a 1206 - > comprehensive manner rather than, hey, you know, leave it up 1207 - > to individual developers to kind of go in and do that.
1208 - > So, number one, introducing those changes, number two, 1209 - > tracking the state of those changes and the metadata around 1210 - > that. 1211 - > So you're aware that, hey, there's two multiple systems 1212 - > involved, but they are running with different types of 1213 - > configurations, but somebody knows about it and that's 1214 - > queryable. 1215 - > So that's kind of like one part of the modernization uh that 1216 - > kind of comes in uh here as well. 1217 - > So I think those are two key pieces.
1218 - > And it gives you, at least from from the from the stream side, 1219 - > I'll tell you that we have as we've revisited this uh the 1220 - > modernization journey, we've also seen customers say, Well, 1221 - > look, I am actually doing a lot of you know what I call instant 1222 - > dynamic alerting, dynamic decisioning by through polls, 1223 - > right? 1224 - > And so there's this ability to introduce that in the path, 1225 - > especially if this is going to be a multi-year journey, right?
1226 - > So then if you sort of put that in the path, in the in the 1227 - > pipeline, that's where you sort of add this whole modernization 1228 - > intelligence uh to the whole system, where you're saying 1229 - > you're getting into now some of the interesting concepts like 1230 - > the difference between historical deep analytics versus 1231 - > streaming real-time analytics, um, doing alerting based on 1232 - > push-based events as opposed to poll-based events. 1233 - > So that's also part of modernization because that's if 1234 - > you take a look at the consumers today, right?
1235 - > Um, if you talk to anybody who's under 20, um, they don't quite 1236 - > understand this concept of getting an email the next day 1237 - > and they almost laugh at it, right? 1238 - > And we laugh at it too. 1239 - > And that's what it is, because you're seeing uh, you know, some 1240 - > devices where people are dynamically saying, oh, let's 1241 - > change this route, order this, get my food delivered by the 1242 - > time I'm there, my Uber is there. 1243 - > That is the mindset.
1244 - > And so I think there's a great opportunity to modernize if you 1245 - > also are trying to say, let's take advantage of these newer 1246 - > types of consumer-oriented aspects within my modernization, 1247 - > modernization journey. 1248 - > So long-winded thing. 1249 - > So I would say that there are four phases here. 1250 - > There's sort of like the, hey, I just want to adopt the cloud and 1251 - > I want to just like migrate to the cloud.
1252 - > Then there's, hey, I want to along the way re-platform 1253 - > certain things, right? 1254 - > Redesign certain things. 1255 - > Then there's the third piece, which is, hey, can I revisit 1256 - > some of the choices I'm making in terms of, you know, where am 1257 - > I doing my analytics? 1258 - > And the fourth one, obviously, is like, hey, the readiness 1259 - > towards, hey, the AI workload uh in the cloud.
1260 - > So what's needed for that? 1261 - > And I didn't touch on that earlier, but an example of that 1262 - > would be if I have a lot of unstructured document I've 1263 - > collected over the last 30 years as a healthcare organization, 1264 - > maybe it's per a perfect time to actually create vector 1265 - > embeddings and stick them into the newer system so that the AI 1266 - > services can take advantage of that through semantic searches 1267 - > and so forth through rack patterns.
1268 - > And I think that is what the modernization journey should 1269 - > look like, those four pieces, in my view. 1270 - > SPEAKER_02: I know we mentioned vector indexes a couple of 1271 - > times. 1272 - > It is a lot more powerful than people give it credit for. 1273 - > It is almost magical.
1274 - > I mean, my anecdotal experiences it's it's always been sort of 1275 - > this magical thing that allows you to interface with. 1276 - > With data in ways you didn't imagine possible. 1277 - > You know, I know it's it's comparable to to search and you 1278 - > know uh uh I like and and you know regex-based matching, but 1279 - > you know, it can really be powerful when you uh lean into 1280 - > its its capabilities uh for for classification and 1281 - > summarization, sentiment analysis, things along those 1282 - > lines.
1283 - > So like and I think it is good advice, like you know, whether 1284 - > you're you know an enterprise that's taking this 20-year-old 1285 - > militarized on-prem database that can't connect to anything 1286 - > and just runs its workload. 1287 - > Now you're moving it to the cloud, you know, as you're going 1288 - > through that journey, yes, think like what's the art of the 1289 - > possible here? 1290 - > Because now this now this database can scale horizontally. 1291 - > And I have, you know, uh VS Code and copilot and all these 1292 - > amazing developer tools to do so much more with the rich 1293 - > operational data that this database is handling in real 1294 - > time, right, through uh through embeddings or streaming 1295 - > analytics or AI agents with MCP.
1296 - > And it's good to kind of get ahead of that in the in the both 1297 - > in the migration process and the modernization design and what 1298 - > the future looks like. 1299 - > Uh, because uh the the the main thing you get there is is just 1300 - > speed of execution and getting the ROI of the modernization so 1301 - > much faster rather than just saying, hey, we were doing this 1302 - > on-prem, now we're doing it in the cloud, and you know, almost 1303 - > like an apples apples comparison.
1304 - > Um, you know, it's it's it's it's really uh a lot more 1305 - > powerful than a lot of companies can can imagine, in in addition 1306 - > to getting a much better developer experience and 1307 - > internal velocity. 1308 - > So definitely uh uh great advice there. 1309 - > And then, you know, for for for companies that are already in 1310 - > the cloud, you know, uh I I still see this sort of you know 1311 - > uh segregation between like the operational database workloads 1312 - > and then the analytical AI workloads.
1313 - > And I think this this it was really great talking to both of 1314 - > you about this because they really have to be converged. 1315 - > I mean, when I I think the magic is really when you combine the 1316 - > operational workloads with the AI, and we're just scraping the 1317 - > surface of of agents. 1318 - > And I and Sharish, like you said as well, uh MCP is sort of the 1319 - > bridge between these AI agents, these deep agents that have 1320 - > autonomy and can solve problems with tools, uh, and get them to 1321 - > interact with the with the rich operational data, whether it's 1322 - > the customer data, the user-facing application, uh, and 1323 - > and and get the full value there.
1324 - > Because that's that's how companies go from having, you 1325 - > know, the uh and Alokio elaborated on this too, like the 1326 - > you know, the uh the the modern experiences that people expect, 1327 - > uh, you know, whether it's it's a chat interface, because you 1328 - > know, chat is taking over user experience in in so many 1329 - > applications. 1330 - > Taking your your you know, your legacy application, which might 1331 - > have been built in 2018, uh just like five, six years ago, and 1332 - > turning it into an AI native application.
1333 - > I think that's where engineering leaders, executives really have 1334 - > to look at combining the database with the AI. 1335 - > SPEAKER_00: Yeah, and maybe John, I'll add one quick thing, 1336 - > which was you know, we were working on some some interesting 1337 - > stuff uh that has to do with governance and validation and so 1338 - > forth. 1339 - > Um, and this this is my personal request to Sharish also so we 1340 - > can actually think about. 1341 - > So, one of the things is then imagine that these systems are 1342 - > SQL and Postgres and Cosmos DB, right?
1343 - > And there are there are subsets of data that are that are 1344 - > supposed to be identical. 1345 - > So, one of the interesting things is you know, we have this 1346 - > validation capability now that we're adding to say, go in and 1347 - > compare these things. 1348 - > But I don't want to stop there, right? 1349 - > I want, you know, uh Sharish, you to expose some uh you know, 1350 - > engine or MCP, but I can also say, hey, these things don't 1351 - > look the same.
1352 - > Were you doing some uh were you responsible for replication 1353 - > between these systems? 1354 - > Uh and if so, can you tell me about these keys? 1355 - > Like when was the last time you had visibility about this stuff, 1356 - > right? 1357 - > So this is truly now taking sort of like this isolated 1358 - > transactional view within the context of one system to 1359 - > conversationally sort of going across these multiple agents and 1360 - > truly trying to answer these kinds of questions, which are 1361 - > super interesting for forensics and troubleshooting and 1362 - > debugging and all of these things that you know our 1363 - > engineers spent so much time on today.
1364 - > Absolutely. 1365 - > Yeah, so that's I think that's what I'm really excited about. 1366 - > The about that. 1367 - > SPEAKER_01: Yeah, no, especially on the cyber tech, for instance.
1368 - > You know, these kinds of questions really matter a lot. 1369 - > SPEAKER_00: Yeah, yeah, absolutely. 1370 - > SPEAKER_01: Cybersecurity, I meant to say. 1371 - > Yeah.
1372 - > SPEAKER_02: Sri Stota, Alok Parik, thank you so much for 1373 - > joining this episode of What's New in Data. 1374 - > I think it's extremely valuable uh for all the listeners today. 1375 - > So thank you for generously sharing all your insights and 1376 - > experience and and and what you're seeing uh through the 1377 - > work you're both embarking on right now, working with all 1378 - > these very innovative, uh scaled-out companies uh 1379 - > launching their next generation applications.
1380 - > So thank you so much for joining. 1381 - > SPEAKER_01: Oh, thank you. 1382 - > Thank you, John. 1383 - > Thank you, Alok, for giving me this opportunity to join you 1384 - > guys today today.
1385 - > I I really enjoyed going deeper, talking about many of these 1386 - > technical aspects. 1387 - > Uh, it's an exciting time, and uh and I'm um once again thank 1388 - > thankful for giving me this opportunity to talk to you all. 1389 - > SPEAKER_00: Yeah, and and thank you, John, again, for having me. 1390 - > And Sharish, it was a pleasure.
1391 - > Thank you. 1392 - > And I look forward to uh multiple future conversations 1393 - > with both of you guys. 1394 - > Thanks.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.