Engineering Founders · 2026-07-30 · 40 min
Key moments - from our scoring
Substance score
65 / 100
Five dimensions, 20 points each
Rajat Monga traces his evolution from scaling deep learning at Google in 2011 through founding Google Brain and co-creating TensorFlow, to launching Inference IO and now leading AI infrastructure at Microsoft. The episode unpacks his decision-making framework for major technical pivots - specifically the "what if we started again" conversation that led to rebuilding TensorFlow from scratch rather than incrementally improving the existing system. A key insight covers why open-sourcing TensorFlow proved strategically superior to Google maintaining it internally, with striking examples like farmers using TensorFlow-powered cucumber sorters and disease detection systems in Africa. The discussion moves into his startup lessons at Inference IO, where Monga emphasizes the difficulty of customer discovery through real conversations rather than search or surveys, and the challenge of translating technical solutions into understandable products that users actually want. His philosophy centers on making intentional pivots early when conditions allow, and constantly reassessing whether current direction remains optimal given what you now know - lessons applicable to scaling infrastructure, founding companies, and managing organizational change.
Jeff Dean recognized that Google's past pattern of publishing papers on internal systems like BigTable and GFS led to inferior external recreations becoming industry standards, forcing Google to support bad APIs. Open-sourcing TensorFlow allowed Google to build the standard itself rather than play catch-up.
Inference IO focused on data analytics - understanding why metrics in production systems spike or drop when even with data access, users couldn't identify root causes through charts and graphs alone.
Monga explored multiple startup ideas including IoT and edge ML, but ultimately landed on data analytics. He handed off TensorFlow leadership at the end of 2019 when the project was thriving, believing that was the right time to leave rather than waiting for decline.
By 2014, GPU deployment in Google's data centers was increasing and TPU work was underway, while models were evolving beyond simple CNNs to LSTMs and RNNs. The existing system couldn't easily support these hardware and architectural changes, making a rebuild necessary.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains solid substantive insights about infrastructure decisions, organizational incentives, and product-market fit challenges, but these are often delivered through extended anecdotes rather than dense information. While Rajat shares real lessons about the hidden incentives blocking tool adoption and the decision-making framework for pivots, much of the runtime is devoted to storytelling that reinforces rather than layering new ideas. The rapid-fire section and some of the edge/inference discussion pack more density, but overall the episode could deliver more insights per minute.
the biggest thing that I saw was there were people all over the world who were now starting to use these tools
does that mean that team has less power? I mean, does that mean that we need fewer people in that team or at least not grow that team? Right. Which is how the leader for that team would think about their perspective and their value in the company.
While Rajat offers some genuinely useful frameworks - particularly around hidden organizational incentives blocking adoption and the 10x thinking methodology - much of the advice recycles familiar startup wisdom: pivot when the odds decline, talk to customers, focus on product-market fit, and balance cloud vs. edge tradeoffs. The insights on incentive misalignment stand out as fresher thinking, but the overall framing leans toward established best practices rather than counterintuitive or first-principles arguments.
if something does get automated, what does it mean for the team that is deploying that? And if some of the data analytics gets automated, does that mean that team has less power?
It's hard to do with everything. But for that tool to be successful, for any tool to be successful, that's um, a super important thing. Right. The people to think about are, uh, the users who is going to start the adoption for that tool.
Rajat Monga is highly credible: he co-founded TensorFlow at Google, was a founding member of Google Brain, founded and led Inference IO through a real product-market fit struggle, and now leads AI infrastructure at Microsoft at scale. He has hands-on experience shipping production systems, making real pivot decisions, and navigating both massive-scale infrastructure and edge computing. His seniority and direct operational experience across multiple phases of AI infrastructure development are substantial.
Rajat Manga is responsible for enabling an efficient AI stack at Microsoft from cloud to the edge. Before joining Microsoft, Rajat was founder and CEO of Inference IO, a smart analytics platform powered by AI. And during his decade long tenure at Google, he co founded and led TensorFlow and was a founding member of Google Brain.
I was at Google already in ads. Ads was a great place, by the way. I learned a lot the kind of scale that Google works at
The episode lacks concrete metrics, named companies (beyond Google, Microsoft, OpenAI, Anthropic), specific revenue figures, or quantified outcomes. While Rajat mentions real examples - the cucumber sorting machine, the cassava plant health app, the Inference IO pivots - these are illustrative rather than data-driven. The discussion of inference systems and edge challenges stays relatively abstract without benchmarks, latency numbers, cost comparisons, or deployment specifics that would ground the claims in measurable reality.
somebody in Japan, uh, an engineer, like a hardware engineer or whatever, built this cool automatic sorting machine for cucumbers for their parents farm
we were training all of this in CPUs. Now we don't even think about using CPUs for these. They are accelerators and GPUs of course, but back then it was all CPUs and that was great. That allowed us to really grow this and get this into products.
The host asks generally good directional questions and shows genuine curiosity about Rajat's journey, particularly around decision-making frameworks and hidden incentives. However, follow-ups are often affirmations or gentle redirects rather than sharp pushbacks. The host rarely challenges claims, ask for quantification, or probe contradictions. For example, when Rajat discusses why Inference IO struggled, there's no probing on whether the product itself was flawed vs. the go-to-market, or specifics on what the pivots actually tested. The conversational tone is warm but lacks the edge needed to extract maximum substance.
This question that you asked of if we started again, what would we do? I feel like it's such a powerful question
I love that criteria in that framework. I think that's great.
Computed from the transcript - who did the talking, and the words that came up most.
Rajat Monga joins the podcast to discuss his leadership and founder journey, from Google Brain / Tensorflow to inference.io and back to Microsoft. He dissects what it means to refound vs. start from scratch, the value of the open source community, and strategies for discovering what problem to solve when going the startup route. We also cover how to determine your users’ hidden incentives and what that means for both product development & marketing, along with navigating the balance between a product’s usefulness and consumers’ willingness to pay for it. Additionally, Rajat shares about what he’s currently up to at Microsoft and the emerging ML / AI technologies he’s most excited about. ABOUT RAJAT MONGA Rajat Monga is responsible for enabling an efficient AI stack at Microsoft from cloud to the edge. Before joining Microsoft, Rajat was founder and CEO of Inference.io , a smart analytics platform powered by AI. During his decade-long tenure at Google, he co-founded and led TensorFlow, and was a founding member of Google Brain. He’s built out and led many engineering teams, and designed large scale distributed systems including web scale crawling and eBay’s search engine.
Transcribed and scored by The B2B Podcast Index.
Speaker A: This episode is brought to you by Sidero. If your team runs Kubernetes, chances are upgrades can feel risky. That's what the Telos platform is for. It's built on immutable minimal OS designed for Kubernetes. Immutable means clusters cannot drift so they stay identical and upgrades are a non event. Minimal means there is almost Nothing to attack. 50 binaries, no shell, no SSH so it's secure by default. Upgrades get boring, security gets boring. And that's exactly the point. Check it out at ah siderolabs. Com. That's S I D e r o labs.com you have to be a little
Speaker B: delusional to be a startup founder. It's fun for sure, but. But you need to believe beyond reasonable things that yes, you can make it work, right? Even with everything against you, with making a bet on some X thing, that if A, B, C, D happen then this will be amazing and it can go to some level and that's great. And sometimes it doesn't go all the way, it just goes a little bit, but it's still fine, you figure out the right and so on. But at every point in there it's important to think about like okay, we are here, we wanted to be somewhere else, maybe, maybe we are better, maybe we're worse or whatever. But now where we are, knowing what we know now, what does that best case and um, like a few options look like? And at some point it's like okay, you know what, is that still true? And is that the right thing to do for me, for the company, for the people here? And it seemed like no, it wasn't anymore.
Speaker C: Welcome to Engineering Founders, the show for engineering leaders making the daring leap to start their own company. Rajat Manga, CVP AI Frameworks at Microsoft joins us to deconstruct his journey as a builder across the frontier of AI infrastructure. And this journey takes us from scaling Google Brain, co founding TensorFlow, navigating major startup pivots at Inference IO to leading AI infrastructure and edge inference at Microsoft. We talk about his re founding framework for rebuilding core infrastructure in the strategic decision to open source TensorFlow. Rajat also shares candid lessons from his founder journey, explaining why human habits are often harder to change than the technology itself, and how hidden organizational incentives can quietly derail the adoption of automation tools. Plus, we look ahead at the future of edge AI and the shift from experimental cloud models to production ready applications optimized for cost, latency and privacy on local devices. Let me introduce you to Rajat. Rajat Manga is responsible for enabling an efficient AI stack at Microsoft from cloud to the edge. Before joining Microsoft, Rajat was founder and CEO of Inference IO, a smart analytics platform powered by AI. And during his decade long tenure at Google, he co founded and led TensorFlow and was a founding member of Google Brain. Enjoy our conversation with Rajat Monga. Uh, Rajat, I'll just say thanks for being here. So much of your journey has been about being a builder and that's across a lot of different elements in terms of defining some early AI infrastructure with Google Brain and TensorFlow to your own startup. To now what you're doing at Microsoft across Edge and Inference and building a lot of really, you know, to undersell it, really cool things. So we're going to be going down a journey, uh, across a lot of those and extract some of the insights that folks in our community can apply across a lot of the different things that we can do. So this is all to say we're going to talk a lot about your journey and we're going to find some really, I think, really fun things to get into. So maybe it's best if we start at the beginning. Bring us into, I guess, the origin story of your time with Google Brain.
Speaker B: Yeah, so this started back in like 2011. I was at Google already in ads. Ads was a great place, by the way. I learned a lot the kind of scale that Google works at, that was interesting and super fun. But like you said, I've always enjoyed building things and building things from scratch is always fun. So, you know, a couple of years into ads, I was thinking, okay, you know what, I really want to build something and something new. I ran into Andrew Ng, who was there at uh, Google, just doing a talk about deep learning and how it's been very interesting in the lab. Turns out he was also talking to Jeff Dean at the same time trying to build up a group to see if he could scale deep learning up at Google scale. Right? Going beyond one or two PCs and going to something big and what would that do? So sounded super exciting. I had played with neural nets back in my undergrad days or read about them. Actually did very little, mostly read and it seemed like a promising thing, but it was early and I was like, you know what, why not? These are great folks to work with. Seems like an interesting space. Two, three engineers and myself along with Jeff and Andrew got us sort of started that group. We were like, okay, here's the goal, right? We have all these models out there, can we just scale them up, throw a bunch of Data at them and see what happens. So that's how things started. Over the next year, we started what was called disbelief back then, which was basically taking these networks, which is also called deep belief networks in some. Some terminology, so distributed relief things, and seeing can we go from that one machine to a few to 100 to 1,000 and so on. And it was an amazing journey. Right. I mean, it started in 2011 and I guess continues even now in many different ways. But of course things started to work well. Early days, I think we started with a couple of different areas. One was speech, one was, um, vision. Over the next year, we had two interesting papers. One was on the speech side, I guess, where the speech team was also doing some work, and the other was on the, you know, what became known as a CACT paper. I guess we had this unsupervised learning approach, et cetera. So fun days.
Speaker C: It's so wild to think about that because now it seems like such the formative time of what's kind of emerging when it comes to different AI releases and capabilities and stuff. And so to think about, like, those early moments, like, did you know that at the time, like, did you think that this was going to be sort of a formative or foundational, you know, technology or experience? Like, what was it kind of like. Like kind of being present to that?
Speaker B: I'll be honest, it feels like. I, um, mean, yeah, it'd be amazing. I knew everything. No, I mean, it can be amazing if it works. That was true. I think that has to be true for some of these things because you're making a bet on something new and there has to be some potential that if everything goes right and maybe something amazing will happen. But that's all we knew. It was a cool problem. There were real challenges to be solved. It wasn't like anybody could just go out and do it. Nobody had done it so far. So there were lots of real things. We came up with new algorithms, new ways of doing things, applying existing things in a very new way, et cetera. So lots and lots of innovation, but with some belief that, you know what, let's put all these right things together and see where that goes.
Speaker C: So I know this is going to kind of jump ahead a little bit, but, you know, part of this story that's interesting is how Google brain transitions or like the jump to then tensorflow. So can you talk a little bit about the bridge there? So at what point did this lead to TensorFlow and what was that leap like?
Speaker B: Yeah, so I mentioned starting to scale Those uh, machines up, we went from one to a few to like tens of thousands in fact 10,000 or so. And the reason was we were training all of this in CPUs. Now we don't even think about using CPUs for these. They are accelerators and GPUs of course, but back then it was all CPUs and that was great. That allowed us to really grow this and get this into products. And over the next two, three years. What happened was pretty much every product at Google was using disbelief and deep learning in some manner or the other. Whether it's vision, whether it's speech, whether it's search was starting to use it. Uh, they were like all kinds of things that were using it. What we also realized was a few things, right one, GPUs were getting interesting, so there was external stuff. By 2014 we were starting to deploy GPUs in our data centers. We were working actively on TPUs. They weren't out yet, but there was active work happening. So we knew that was coming. The models were evolving. We were moving from just vision models that go straight up or convolutional nets, et cetera, to uh, more LSTMs and RNNs, more recurrent networks for things like speech and text, et cetera. We didn't have transformers yet, but you know, and we wanted to try out all kinds of crazy things which our existing system was just finding it too hard to do all of these. So we were basically trying to change each of the pieces altogether and things were going crazy. So at some point we're having this conversation and Jeff's like, he writes up his talk saying, okay, you know what, what if we start again? What would we do? And there's some ideas. And we started talking about that and spent a little bit of time, like I said, a large number of products using us actively. This was an internal only project, slightly easier than an open source project in some ways, but we needed to support them as well. And we were a small team. We had like 10 people or so at that time. And so we were like, okay, how would we do this? What would we need? And after a couple of months of back and forth, et cetera, we were like, yes, let's go ahead and do that, let's figure out what we earn. And then um, we started getting a few more people. We decided, okay, a handful of us would continue to support the existing ones, but we want do a lot of new things on the old system and we'll start building the new system. So over the next year, I would say mid 2014 to late 2015, I guess. When we launched TensorFlow outside in mid 2014, we started a few early users inside as well. We were basically heads down building a whole bunch of things and supporting whatever's needed.
Speaker C: This question that you asked of if we started again, what would we do? I feel like it's such a powerful question and I've heard people apply this. Maybe it's like different moments where they have to make major pivots in the company. Can you talk to us a little bit about, like, what was that conversation like to really dig in and have this almost like refounding of this product or like this concept idea? Uh, what was that conversation like when you were working through that with everybody?
Speaker B: There's always the question of when you have an existing system versus starting from scratch. You're always giving up something as well, which is hard. So there's that side of the story where you're like, okay, if we keep doing this, here's how we can change all of these and we can accomplish what we want. And then you're like, okay, let's stop that for a second, where it's saying, okay, I don't have all of this. If I'm thinking afresh, uh, what does that allow me to do? What is that right place? Where do I want to be? What does that look like? And then it's an interesting thing to see. Okay, can we evolve this enough into that? Can we do. We should be really pivot hard, pivot, et cetera. We were a small team and like a startup in many ways. I mean, with a lot more resources in most startups, but still making that choice clearly and early is super helpful. And having gone through a startup later and looked at things differently, I feel like that's super important, super hard at times. But if you can make that pivot, it just enables so much more. Like in our case, knowing now what we went through with TensorFlow and what we could do, there was no way the old system would have been able to do that. Right. And we would be hobbling along.
Speaker C: This may tease like future stories, but like, what I really appreciate is these structures or these modalities that you methodically work through with with the team that then drive like speed and impact in a lot of other ways, but like having these really intentional conversations shape then a lot of the direction and then the speed in which you're able to go in different ways. I know we'll talk about kind of other patterns of this later on. But I was just like, uh, you know, as we're talking about this, I really admire that. There's one product decision that I thought was really interesting was like, you know, TensorFlow being able to be made available externally. Like, what was it like coming together and making that product decision? Like, what was kind of, what was the context around that?
Speaker B: Yeah, so that was interesting. And again, I would, um, give Jeff a lot of credit for sort of pushing that a bit and really helping us think about that the way we were, you know, starting when we started out, we were not thinking about open sourcing. In fact, because the existing stuff was internal, we were like, okay, let's start building. Then a couple of months through that, we're talking about things again. And the conversations about, you know, what Google's done all these amazing systems internally. Before that we had things like. Org and GFS and uh, some others as well. And every time, you know, Google would publish a paper after that system's built, there's some not so good copy of that product that gets made by somebody else. Eventually a lot of the world starts to use that. In fact, in one case, and I think this was, uh, bigtable, probably where Google had to number. By that time Google Cloud was there as well and we were starting to grow that Google Cloud had to support this not so good API from the external world on the system that was the original system because that was the standard now. So the question was, instead of doing that, why don't we build that standard and why don't we make it available? At least the initial thought was, okay, let's make at least part of this available and then we'll see what makes sense. Do we keep some, do we share everything or not? Uh, although in the end we ended up sharing pretty much everything as well. But there was a lot of value in sort of thinking about that and going through that. There were some conversations with, you know, at that time, Sundar as well, on okay, should we or not? And in the end I, uh, would say largely Jeff had a strong support for that, which definitely helped. Clearly. I mean, looking back, it was the best choice we made, I guess.
Speaker C: Well, and you were also telling me about like some of the amazing things that that unlocked like within the open source community, like some of the different projects that people were able to build on top of that. Maybe like as you're reflecting, like some of those projects you look back on, you're like, wow, I can't. Like, I'm so happy that we did that, so that these things existed in the world. I don't know if there are any projects that you reflect on from sort of the open source world.
Speaker B: Yeah, I think there were a couple of things. So the biggest thing that I saw was there were people all over the world who were now starting to use these tools. And it's not like TensorFlow was the first open source tool to do deep learning. There were a few, but mostly research projects, et cetera that came out, et cetera. And suddenly there was this tool that could do interesting things and available to all kinds of random hobbyist developers. And we found these projects where one was like after the first year or so, somebody in Japan, uh, an engineer, like a hardware engineer or whatever, built this cool automatic sorting machine for cucumbers for their parents farm. So his old mom used to sort them by hand and he built this cucumber sorter using like a simple classifier, et cetera, using Tansy flow. Just amazing, right? I mean actually changes somebody's life. Similarly, you know, people, somebody in Africa was using it for understanding like the cassava plant, are they healthy or not, when to do art, et cetera, uh, with just some app that you could run on the phone and do some interesting things. So just the kind of things that enabled we would never think of really was uh, amazing.
Speaker C: You know, big shout out to uh, bootstrapped uh, farmer developers because I feel like that's like the ultimate use case of constrained resources because farms are so hard like in terms of labor, like the dynamics are such that they have to really operate it leanly. Like my brother in law is a, is a rancher and so like I just am kind of a part of the secondary conversations about the concerns of running the ranch. All to say shout out to farmer developers like that's such a, those are such cool stories to see like that, that blend and like how that impacts families. I just, I love that.
Speaker A: This episode is brought to you by Sideroom. If your team runs Kubernetes, chances are upgrades can feel risky. One CVE patch can put a whole fleet in doubt. What should be routine takes too much time and often ends with your best engineers fighting fires. Instead of shaping a roadmap, the root cause is underneath a general purpose operating system that was never built for this. The moment someone runs a package manager or hot fixes a box at 2am you have a drifting node. And worse, the OS comes with a large attack surface that you never needed. That's what the Telos platform is for. It's built on immutable minimal OS designed for Kubernetes immutable means clusters cannot drift, so they stay identical. And upgrades are a non event minimum. Means there is almost Nothing to attack. 50 binaries, no shell, no ssh. So it's secure by default. It handles the fleet's lifecycle for you. Provisioning, upgrading and retiring machines automatically. That design matters most when the stakes are highest, including the edge with no one on site and AI clusters where drift is expensive. Upgrades gets boring, security gets boring. And that's exactly the point. Check it out@ah, siderealab.com that's S I D E R O labs.com what's interesting
Speaker C: then there's like, there's a part where there becomes a time to, you know, sunset your collaboration within like the Google Brain and TensorFlow World and Transition to something else. And so I was wondering if you could bring like, you know, this broad question of when is the time to hand things off and do something new. And then you jumped into what your startup was. So talk to us about like that transition and then getting involved with inference
Speaker B: IO, you know, at the back of my mind I always wanted to start a company. I'd been at several startups before, so I was used to building early things. Brain was sort of like an early startup, uh, in a very large company. So very, very different. I guess so. So that was at the back of my mind. But there was never a good time. I mean, in some ways I was having a lot of fun at Google, building TensorFlow, et cetera, so that wasn't really a friend of mine. But then after a few years and we were getting to this place where we were about to launch TensorFlow 2.0, things great, lots of things happening. I was like, okay, would I do this? Is this a good time? Should I do some other time? It's hard, especially leaving something that's been going on for a while that's going great. But on the flip side, I guess nothing like leaving things when they're going great because you believe that they can continue like that or whatever rather than, oh, they're not doing great, let me just hand off to somebody and throw it out or whatever. So anyway, I started thinking about the right timing and that timing seemed good. Exploring some different ideas on what to do next. This is around, um, end of 2019 or so, I'm exploring that. That's around the time I actually had handed off things. And then I'm like, okay, let me go off and spend some time to figure this out. There are a couple of different ideas that I was looking at. That I can talk about as well. But yeah, it was an interesting transition at that point.
Speaker C: So I want to get into. I think there's kind of a couple phases of the inference IO story. So I think one is this like problem discovery phase, because you were sort of working through a couple pathways to go down. So what was it like and how did land on the problem that you pursued?
Speaker B: So there were two or three big areas that I liked. The first one was Iot. So there was a lot of stuff going on there. I believed and still do that. I think a lot of ML needs to move to the edge as well. There's more and more interesting things that we can do at the edge. In fact, there are a lot of companies doing that where the products, devices that you use, which have some form of machine learning in them. Uh, so that was a goal. So that was sort of one idea. The second one was on the data analytics side, which is where I ended up. In fact, it was broader than what I was doing. But, you know, we see data all the time in every business. Whatever we do, we're looking at charts and graphs. So they're in fact, not just business, but even as a developer, if I look in looking at production systems, trying to understand what's going on, there are all kinds of charts to understand, okay, is everything okay or not? So my thought was, okay, can we do better? I've had cases where even with the data for TensorFlow, et cetera, we would look at it and we're like, oh, this seemed to go up or down and why nobody knows. I asked that question. Sometimes I get an answer. Usually you don't even get an answer. And so how do we solve that?
Speaker C: The element of this that I want to get into as well, because I think there was a lot of parts of that journey that taught you a lot. And I know that then what you're doing now, a lot of how you're thinking about is all these new ways of like, product strategy and how these things are sort of evolving. So like, for me, I'm like, okay, cool, let's get to the origin of some of these insights and lessons here and then see how it's manifesting. So I guess bring us into the product market fit journey and some of the insights that you gained along the way that were really interesting or shaped a lot of your lessons around building things.
Speaker B: There are lots of conversations. It's not like, okay, just talk to ChatGPT. I mean, there was no ChatGPT, but even Google Search, right? You can't just do Google search and figure this out. Yes, you do a lot of that, but you need to talk to real people to understand that, right? It's lots and lots of conversations and they're great. Uh, I mean, a lot of conversations are like, okay, I know most of this, but then there's like one nugget, one new thing that you learned from that. You take that and you build it and integrate and then go to the next one and so on.
Speaker A: Right.
Speaker B: And so that got me to. Okay, let's focus on this. Kicked off the actual company around data analytics, the inference IO. During that time, my co founder came on board as well. And fun times around Covid, of course. But, um, it was an interesting journey. You know, once we had the stuff, we were working with a few design partners to help us build the right thing and see if that actually solves their problems, etc. And we did like, we iterated on the product with them, found that there were some interesting pieces and there are a lot of interesting things. As you know, we build a product that we think is useful, great. Then you go show it to somebody who wants to use it. They're like, oh, yeah, this is okay. I mean, maybe this will help here a little bit. But what is this? I have no idea. Right, so you sort of iterate on just that UX or whatever it is, right, and understanding, et cetera. But we found a couple of places where, okay, we were like, okay, yes, yes, this is useful. Great, now are you going to pay for it? Uh, how does that work? That, that was, you know, this is where now you have a product that seems useful, but is there a really a market for it? Are people willing to pay how to do that? And that was an interesting one. Lots and lots of conversations and you know, people are always like, yes, this is a problem that we face. Every week or two they would see something in the data or some exec will ask, oh, here's the problem, go find me. And it take them a couple of days to go figure it out. And if somebody could solve it automatically, that's awesome. Now we build that and we show it to them. Great. Super interesting. What do we do with it? And this is where, uh, it starts to get interesting. Where, how much is it worth it to them at this point? Ah, what are they trying to do and where does it fit in their entire workflow? So people have a certain way of doing things in the data analytics side. They have their own things. They were used to certain dashboards which they don't I mean they want them as is because somebody's asked for this chart and they're used to seeing it. And then let's say they want to solve a problem. Are they going to switch between that and this other tool that we've provided or can they replace it fully with this new tool, which means we have to basically compete with the likes of Tableau and others on day one, which is probably not the right thing. So lots of hoops that we jumped through. We looked at a lot of different verticals as well, in some ways looked at marketing. That was the first thing that we were thinking about then product. Talked to a few companies using things for product. In fact later on we talked to somebody using things for experimentation. Right, you want to do some analysis there, et cetera. And a whole bunch of different ones. There was interest, there were a few folks who actually signed up and paid us for this, etc. But by and large it was a tough sell in that they had some product that they were used to, this was a problem. But they had other problems that just seemed m way more painful to them at this point. They were like, yes, it's painful but we're used to doing something for it. We know how to cope with it, I guess. And so it's a lot harder to transition them over to something new.
Speaker C: I've been kind of reflecting on that part because I think right now in terms of the analogy here is how people are adopting different AI tools that I'm drawing. And to me oftentimes the patterns are human behavior driven or habit driven and the adoption is more dependent on can you tap into the psychology of somebody and their pattern of learning a new flow? And ah, so as you're talking about this like the, you know, people are used to this product and kind of navigating that part, would they use a different tool? I think is really interesting. There's a couple other pieces that, a couple other insights I thought was interesting. I was reading, I was reading your write up, reflecting on the inference IO piece and 1 of the things you mentioned was this observation of like, does this tool give this user more power? And like however they define power within their job function. Um, and I thought it was really interesting to sort of look at power and incentives and like is it helping them achieve like even a broader goal than just like, like accomplishing that like specific problem? Talk to me about that observation.
Speaker B: Yeah, yeah, that's an interesting one. Right where in any company people have different incentives. While you could argue that you know what, everybody should be Aligned towards whatever matters to the company kind of there. But everybody has their own little incentive as well. So I'll give you an example. Right. We were often talking to the data teams because that's the team that's going to be using or at least deploying this. Right. Even if their users adopt this too. But the, uh, idea with automation and something that we're seeing with AI today as well is that, ah, if something does get automated, what does it mean for the team that is deploying that? And if some of the data analytics gets automated, does that mean that team has less power? I mean, does that mean that we need fewer people in that team or at least not grow that team? Right. Which is how the leader for that team would think about their perspective and their value in the company. Uh, so this sort of is one of those examples of misaligned M incentives. Right.
Speaker C: It's so interesting to think about the hidden incentives and how that shapes tool purchasing. Like, as you lay that out, like I was like, that is not a first thing that I think of as you, as you're starting to lay out sort of this pattern of behavior. I think it's just, to me, it's wild.
Speaker B: Absolutely. I mean, look, the reason I started this was, to me it was like a very clear problem. Here's a problem that is there. If I have a product that can solve that problem, I would love to use that as an end user. I think I would still love to use that tool. But just the whole process of getting a tool inside, actually making it available to more folks inside the company, how do individuals adopt that? And now I'm seeing that, of course, with the broader generative AI transformation with ChatGPT and all of these other tools and see the same thing everywhere. Right. And there are some people who will adopt these things very, very quickly. Others will take a long time, and then others who will try and avoid that as long as they can because it feels like too risky for them to do it. And usually it's the wrong thing to do because it's actually the other way. If you don't adopt it, that's, that's riskier in the long term, even though maybe you get six more.
Speaker C: Okay, so thought, thought exercise here is like going to the marketing side of this, like, could you market to those hidden incentives? Like, is that a, is that a way to kind of tap into some of that psychology in a way that better speaks to like, your tool's value? Like if you have a clear idea of what some of those invisible things Are like, do you think marketing is effective in terms of how you package the product that way? And maybe this depends on the tool. I'm kind of thinking of like some specific things. Maybe I'd be hacking on on the side, like, oh, like how does this align with sort of like psychology and these hidden patterns? I don't know. I don't know what your thoughts are. This is more thought exercise than anything.
Speaker B: Absolutely. I think it's hard to do with everything. But for that tool to be successful, for any tool to be successful, that's um, a super important thing. Right. The people to think about are, uh, the users who is going to start the adoption for that tool. You have these whole ideas of bottoms up and top down, et cetera. And they're both good ways. In fact, later on we were looking at some more bottoms up as well. It's fine. But in both cases there are different Personas that you're going after. Who will be looking at that in the bottoms up thing, it might be an engineer, it might be a data analyst or somebody. Do they have access to the right pieces? Does it align with their incentives? Does it help them personally in some way that they want to do it? And top down, it's more the buyer. Does it make them look good? So in the case of a leader in the company, does it help achieve their growth, personal growth incentives as well, does it make them look good with their bosses, with the rest of the company, et cetera, like a year from then, will they be lauded for that or will they be kicked out because they made that mistake? Right. On that extreme, there's always this. If you bought IBM or Microsoft or any of these, you're always fine. Right? There's you. You can't do any wrong with that. Right. That's an easy one. So how do you, as a, uh, startup, how do you get over that as a hard one?
Speaker C: Yeah, I love that criteria in that framework. I think that's great. So you and I were talking about how sometimes the decision to stop is oftentimes way harder. You went through a specific exercise. So going back to what we were talking about, the pattern of revisiting key decisions in more of a, uh, structured way, you went through an exercise to identify and work through this decision with you and members of the. I'd love to know, like, what that looked like, like, what did that, what it was like, the framework of the structure of that conversation. Were there certain signs and signals that you were starting to look at to help you work through this decision to
Speaker B: stop it was a tough one. Even now, when I think about it, it's hard. It's a journey. Not just as a founder, not just invest your time and effort, but also the team. And so there's an extra responsibility as well. So there are a number of things to think about, right? So when starting something, it's like, okay, you know what, we're making a bet on some X thing that if A, B, C, D happen, then this will be amazing and it can go to some level, and that's great. And sometimes it doesn't go all the way, it just goes a little bit, but it's still fine. You figured out the right exit and so on. But at every point in there, it's important to think about, like, okay, we are here, we wanted to be somewhere else. Maybe, maybe we are better, maybe we're worse or whatever. But now, where we are, knowing what we know now, what does that best case end up? Um, like the few options look like. If we were going back to our earlier conversation, if we were to start with where we are now, what would we do and how would we do it? And then does it make sense? Right? What is the right thing to do at this point? Uh, so there are lots of conversations around that. Uh, we were looking at a few pivots around that time as well. I mean, there were pivots that we had done in terms of the verticals that we were going after and focusing on different verticals. But now we were thinking about it from the marketing side, right? Going from the top down to more bottoms up thing. And so we were building something around that as well. And at some point, you know, my thought was maybe, but it's a long shot. And maybe if you were starting with that two years ago, may have been an okay thing to try, because we know whatever we know, but where we are, all of that, the odds of success are way lower than where we started from as a startup. I mean, the odds are never great. You're making it back because you really believe in it. At some point, you know, there are both sides. You have to be a little delusional to be a startup founder. Uh, it's, uh, it's fun for sure, but. But you need to believe beyond reasonable things that, yes, you can make it work, right? Even with everything against you. And at some point it's like, okay, you know what, Is that still true? And is that the right thing to do for me, for the company, for the people here? And it seemed like, no, it wasn't anymore. So, I mean, at that point it was like, okay, you know what, what should we do?
Speaker C: Uh, as you talk about, I can hear how hard it is to kind of work through in a really structured way, like to get to those types of conclusions and then the downstream of what that means. So not to dismiss the weight of that, but I do want to pivot to an area that I know you're really excited about. You are a builder at heart. It is so evident through all of these different phases you are now at a point where you're building again within the context of Microsoft and on a couple of problem areas that you're really excited about. So could you bring us into that transition to what's going on with what you're doing at Microsoft now?
Speaker B: Yeah. So after my M startup I was thinking about again the AI space, where it's going. A lot had happened since I, uh, had moved out to Google. ChatGPT was out. A lot of interesting things were happening. Wanted to be closer to where I could see things taking shape. There are a handful of companies, uh, doing things like that. And I thought Microsoft was in a good place to really explore. Happened to know somebody there. So I had a conversation and felt like, okay, yeah, there's some interesting things, let's go figure it out. And so the area I started with, okay, I'm familiar with building infrastructure like you said and seeing. So this team there was looking at inference for Microsoft, both really across the board. And we have the stack that runs on all kinds of devices. In fact, it runs everywhere, including the cloud as well. When I came in I was like, you know what, the stack is great, but we have a small team, let's focus and there's a place that we can do really, really well. Um, and that's the edge dev. So this includes laptops. We ended up partnering very closely with Windows and it now powers ML in Windows, basically the Windows AI platform itself, but also phones. And you can run apps on Android and iOS and other places as well. Right. So that was sort of one part of the charter. The other was the other extreme. Right, where we have the largest models in the world. We partner very closely with OpenAI and uh, we offer those same models to our customers and our internal partners as well, internal products as well. So if you think about Office 365 et cetera, they are using some of these best models to really offer the copilot experiences and other things to our customers as well. And our team really focuses on how do we make them really efficient, how do we run that at scale so that it's effective for us and our customers. Right. You make it faster, you make it cheaper. All the good things about that. So I've been looking at these two problems together.
Speaker C: Basically it seems like two wildly different problems like those on their own could be almost their own charter. Uh, I know we're kind of like bounce back and forth between the two of them because I think they're so critical. Can we dive deeper into the edge strategy pieces? Walk us through that and some of the problem space and why that's been so interesting for you.
Speaker B: The problem with the edge is we are seeing more and more interest. I talked about my Iot idea of running ML at the edge. That's still true and it's over time actually gotten more real in a variety of ways. If we think about we're using this solution today, but if you use something like team teams or any of the standard video conferencing things, et cetera, there's this blurring kind of features, et cetera. All of those, there's something happening on the device, right? So there's so many different things that we don't even think about which are happening on the device that it made sense to think about. Okay, how do we focus on this? What is important here? Given that the stack was already there, the key part of what I brought into the table was, okay, really rethinking the strategy of what's a good focus area, how do we think about it, how do we grow it, how do we measure success, et cetera. And that's like to, you know, I mentioned that now this runs as it's always been used on Windows by a lot of people, but now it's part of the core Windows ML thing that's offered that just powers that as well. And those kind of partnerships, given that we are Microsoft, that made a lot of sense. It helped us avoid building two stacks in parallel for the same kind of thing and doing better. Uh, and then has allowed us to focus more on some of these other platforms as well, like Android, et cetera, for the mobile and other devices that people want to run as well.
Speaker C: When you think about some of the trends in dynamics, I think it's really cool when you're connecting it back to the original problem you were exploring around ML on the edge. When you think about some of the trends that are emerging now and maybe engineering leaders building in that space, are there certain reflections that you have, trends that you'd want to share to help people building in this space as this trend goes where models get better, more gets run on the edge versus the cloud. What are you starting to see or anticipate? What does the world look like as, uh, as months go by?
Speaker B: I think we are at this place where models keep getting better all the time. So we are like, you know what? Yeah, let's just use the latest model and oh, that's amazing. Which is great. When you are trying out new things, that's the right thing to do. But once you have, uh, a certain kind of scale or you figured out the product market fit, so to speak, that you know that these kind of models allow you to build a product that's useful and do amazing things with it. Now you think about a few different things, right? You're looking at cost, you're looking at other parts of the product, right. You're. Are you getting the right latency? Can you change the experience if things were different? Right. And sometimes it's about, you need some of those things on day one. For example, privacy for certain things where you don't want the data to go to the cloud, in which case you'll start on the edge itself. But it's this mix where, uh, in some cases, like privacy, you start on the edge because that's the only option. It's either that or you don't use ML. I guess you don't use AI at all. The other side is okay. Yes, I want to use the latest models because that allows me to, to experiment and try and see if this works at all. But once it does, what's the best way to deploy? Sometimes it's still the best model because there's nothing that comes close. Sometimes it's less about that last bit of quality versus the lower latency, lower cost, or something that's really in the hands of the user, which really makes sense. And so that's how to think about it, right? I mean, there are different aspects of the same side, and I think we'll see more and more both in the cloud as well. There's this notion of let's get the best model or maybe the second best. Whether it's in the cloud itself is okay, because that solves my problem. It's way smaller, so it can be way faster, and so I can do a lot more with it rather than just trying to get these the best one.
Speaker C: Are there certain products at the edge that you're excited to see come up, like categories or devices you see that can become more powerful? I think. Are there some things kind of emerging that you'd be really excited to See, become real.
Speaker B: The whole robotics space is super interesting. Right. It's still very early. I mean the biggest robot of course we see is like Waymo or these cars running around. They are their robots too and they are running a whole bunch of AI at the edge in the car. We don't think of it the same way, but we have to actually do that. Their latency is sort of the biggest constraint. You cannot go back to the cloud and come back. We'll see more of that with things like robots for sure. Another one is the glasses that we see now from Meta, but the others as well. I used to have this Google Glass like 10 years ago or so and very early geeky thing that you have. But it's amazing. There are privacy issues there that we have to tackle. Uh, and do. Right. But it's amazing powerful technology. So we will see that continue to improve and have more of. Right. You know, I'm a geek at heart and I like to try the newest gadgets, so definitely want to try more of those.
Speaker C: What's interesting is so thinking about things around like the edge and focusing on like edge challenges is almost like a different problem shape than also thinking about like the inference, uh, like in the inference challenges there. So I was wondering if you could talk a little bit about I guess that problem space, like what are you working on when it comes to large scale model inference, the problems in that space and some of the dynamics. Because I think there are engineering leaders that are probably wrestling this in different ways. So I'm trying to peel back the curtain and get a window into your world there and what that's like, uh, what engineering leaders can learn from that.
Speaker B: Totally the key piece that you asked about, right. How do you decide what to use, et cetera. That remains the same. Do you use the best model versus what's the best to scale with, et cetera. So that problem is the same from the technical challenges, et cetera. You're right. Completely different thing. This one in the cloud is much more of a technical challenge. The other one is much more of a product challenge and how to fit and get the right pieces, et cetera. So in some ways uses different parts of my brain and I have to flip flop between them half day each. Uh, but it's fun, it's been fun in the cloud side. Look, the systems that we are building today. I've always loved distributed systems and their complexities there. It's almost like these inference systems are going from. Oh yeah, you used to run this model on a single GPU or a single machine to. I mean yes, training was huge, but now inference itself is going to a few GPUs, to multiple GPUs across machines, to really managing this entire cluster of machines with all kinds of GPUs, maybe more than one kind of accelerator and managing all of them. So almost like um, you know, from a systems problem perspective it's like a distributed operating system that you have to think about and coming up with all kinds of challenges there, which is super fun for me and you know, intellectually, uh, challenging and fun from that perspective from um, a product and the, you know, the user's hat on I guess. And we'll see where some of this goes. But feel like we are still at this point where a lot of companies are early in their journey in adopting AI and there it makes sense to try out the best models and people are today continue to move to the newest model, whether it's from uh, OpenAI or Anthropic, which are ones we offer at Microsoft and some of the others as well, other vendors as well. But then there's still a fraction of the population that are using these open source models, maybe fine tuned for their case etc. Sometimes they're not as good, sometimes with the right help they can actually be at least as good for their specific use case because of how things work. And we'll continue to see this m mix but as people scale up, the question for them would be is the best thing for them to continue with the best model out there or do they switch to something that's more manageable and gives them other things to play with? Right. And that's going to be an interesting journey here at Microsoft. I think about both and how do we enable customers and we offer all kinds of options of course and the ecosystem is continuing to evolve very rapidly as well.
Speaker C: I think what's incredible is like as you're laying this out, it almost is like the expression of all of these different elements of other projects that you've worked on are kind of getting expressed right now. So thinking about customer enablement, thinking about distributed operating systems, then uh, thinking through like the technical and product challenges like all at once it sort of seems like this is like you get to, you uh, express a lot of the different areas that you love enjoying solving problems and building around. Uh, that was just like the sense that I got as you were kind of describing all of the dynamics that are at play.
Speaker B: Absolutely. It's been a lot of fun thinking about all these different things and making them work. So Definitely a lot of fun there.
Speaker C: I love it. Rajat, we've got some rapid fire questions if you're ready to go.
Speaker B: Go for it.
Speaker C: Okay, perfect. First question. What are you reading or listening to right now?
Speaker B: Epic disruptions. I think it's got somebody, um, forgetting, but it's about all kinds of disruptions that have happened over time and. Super interesting.
Speaker C: That's great. Next question. What is a tool or methodology that's had a big impact on you?
Speaker B: Yeah, maybe I would say 10x thinking. So something I would say I learned probably from Jeff back at Google, where you're like, okay, you know what, we can do this A, B or C. Well, what if we could do it 10x bigger or where could we go if this was really successful, etc.
Speaker A: Etc.
Speaker B: You know, and that, that I think really opened my mind to a very different perspective.
Speaker C: Well, it's interesting because then it also, what I appreciate that too is like, it also surfaces constraints or helps like, bust myths around constraints too, when you start to really press some of those, those 10x questions. Next question. Rajat, what is a trend that you're seeing or following that's interesting or hasn't hit the mainstream yet?
Speaker B: Well, it depends on what you call mainstream. I guess in the Valley, mainstream is all of tech. Nothing seems new, but I would say robotics and, uh, some of the applications of AI to the biology side, etcetera, are definitely super interesting to me. Uh, something I follow very deeply. I guess they are still maybe a few years away from hitting mainstream, per se, but super exciting.
Speaker C: Last question. Rajat, is there a quote or a mantra you live by or a quote that's been resonating with you right now?
Speaker B: Oh, yeah. Uh, this one's easy. I love this quote from Steve Jobs. I guess. Stay hungry, stay foolish. Always think about it in all kinds of ways and keeps me excited about new things.
Speaker C: Love it. Rajat, thank you so much for taking us down your journey, uh, and I think especially for the folks in our community that consider themselves builders, uh, I know that they see a lot of themselves in your journey and pathway, and I think just appreciate the generosity of sharing your lessons, uh, across all of the different stages and scales.
Speaker B: So thank you, thank you for having me. It's been a fun conversation. Really appreciate that, the opportunity.
Speaker C: If you're listening to this and you're wondering, how can I connect with other engineering leaders in my city? Pull up your phone right now and go to elc.community click our chapters page. You can see that on the menu on the left. Find your local chapter and click Join. We're hosting virtual and in person events all the time and this is the best way to help you get involved, expand your network in your city, and support your leadership and career growth. So pull up your phone, head to Elliot llc Community, Join your local chapter and get involved. A huge thank you to all of our local leaders who make Community happen, and thank you for listening to the Engineering Leadership Podcast.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.