
People of AI · 2025-08-28 · 56 min
Muhammad Farooq's journey highlights a fundamental shift in how developers work with AI systems. Two years ago, he built Local GPT by hand - researching, coding, and testing piece by piece - to enable enterprises to run RAG systems locally without exposing proprietary data to external APIs. Today, rebuilding Local GPT 2.0, he's taking on a different role: as a technical architect who writes detailed specification sheets and manages agentic workflows using tools like Gemini, Cursor, and Ollama, rather than writing every line himself. This shift mirrors a broader evolution in the developer landscape where deep technical expertise now means understanding how to orchestrate AI agents effectively rather than pure coding ability. The infrastructure has matured too - Llama.cpp has become more stable, quantization formats like GGUF have standardized, and coding assistants can now handle frontend development, documentation generation, and test coverage at scale. For operators building with LLMs, Farooq emphasizes that RAG isn't just about retrieval; it's a foundation for agentic systems that automate search-driven tasks like report generation and enterprise automation. His 200,000 YouTube subscribers and 21,000-star GitHub project demonstrate how documenting this evolution in real-time builds both credibility and audience.
RAG is a technique that brings private enterprise data into LLMs so they can answer questions based on fresh, business-specific information rather than just their frozen training data. It's critical because most enterprise data is private and inaccessible to public models like Gemini or ChatGPT, but RAG enables businesses to harness LLM capabilities on their own data.
He built Local GPT to enable enterprises to run RAG systems using local models in air-gapped infrastructure, so organizations uncomfortable sharing proprietary data with external APIs could still leverage advanced capabilities privately and securely.
Two years ago developers had to do all research, coding, and testing manually with limited tool support; now they can use agentic tools like Gemini and Cursor to handle specific focused tasks, write tests, build frontends, and read documentation at scale - shifting developers into a technical management role rather than pure hands-on coding.
He's using a combination of Gemini, Cursor, and other agentic tools alongside mature infrastructure like Ollama, vllm, Llama.cpp, and MLX for local model deployment, with MCP support for contextual information.
His YouTube content on technical topics and open-source work (especially Local GPT) generates visibility and credibility, which directly feeds consulting opportunities as industry practitioners discover his expertise and reach out for implementation help.
Computed from the transcript - who did the talking, and the words that came up most.
Meet Muhammad Farooq, a machine learning engineer and a YouTuber on the channel @engineerprompt, where he explains AI concepts and builds his own projects. In this episode of People of AI, hosts Ashley and Christina chat with Muhammad about his journey in tech, his open source project localGPT, and how he's helping organizations deploy scalable AI solutions.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Don't limit yourself in terms of the capabilities both like from a human perspective and these systems can bring. Right. For me, uh, thinking about what's possible, uh, and looking at the future, what is going to be possible with this system, I think that that's going to be huge.
Speaker B: Wow, Christina, what a great conversation with Mohammad Farouk.
Speaker C: I love Mohammad. I think he's fantastic. He has such a great story, right? Like he has over 200,000 subscribers on YouTube and he got them in a relatively short period of time. But what I think is so impressive is that when he started his channel, which was just kind of because his friends were like, we want to know more how these systems work. And they're like, well, you should share this with everyone. You know, it blows up because he happened to be, you know, right at the perfect moment when um, everyone was starting to care about, uh, LLMs. Um, and he has the capability of being able to, you know, really inform people how they work and teach people stuff. It's incredibly cool.
Speaker B: Yeah. His story about building Local GPT, which is his first very successful open source project on GitHub, is so interesting to talk to him about how he started building it two years ago and how he was so hands on and really building it piece by piece. And he is now launching local GPT 2.0 and built it in a completely different way.
Speaker C: This whole season is about builders. And I think that Muhammad is like the cosmic builder, right? Because he's, he's, he's a creator, he's a builder and he's doing the things and we can all see what he's doing both as he's documenting his journey on YouTube, as he's walking through this is, you know, how these systems work. Um, and also just looking through the code itself. Right? Because he's made it open source.
Speaker B: Definitely. You said it. He is the consummate developer. So let's, let's go meet him. This podcast is sponsored by Google. Any remarks made by the speakers are their own and are not endorsed by Google. We're so excited to have you, Mohammad. Welcome to the studio and welcome to the podcast.
Speaker A: Thank you, thank you.
Speaker B: Like with all of our guests, we're going to start by reading your bio. Uh, so Mohammad Farouk is a PhD trained machine learning engineer who specializes in information retrieval and retrieval augmented generation systems, or RAG for short. Through his consulting practice, he has helped organizations deploy scalable AI solutions that harness large language models and autonomous autonomous agents. An open source champion, he created and maintains local GPT, a widely used RAG framework with more than 21,000 stars on GitHub. He also shares hands on tutorials and industry insights with more than 200,000 subscribers on his Prompt Engineering YouTube channel where he combines deep technical expertise with a passion for making cutting edge research accessible and useful to builders everywhere. Welcome.
Speaker A: Thank you, thank you.
Speaker C: Welcome Mohamed. We're so glad to have you here.
Speaker A: Uh, thanks for having me. Uh, it's my pleasure.
Speaker C: So we want to kind of get into your story a little bit and how uh, you know you got into um, as Ashley mentioned in your bio, you have A um, your PhD trained. Um, but um, I think we became aware of you because of your YouTube channel which as uh Ashley mentioned you have over 200,000 subscribers. So just kind of getting into this um, I guess how did you first kind of get into tech and then what led you into uh, starting up your YouTube channel?
Speaker A: Uh, yeah. So um, as you said I have a PhD so that was kind of a more academic uh, path. Uh, I did a lot of research on um, uh, sensors and machine learning systems specifically focused on biomedical engineering. I always had interest in um, natural, uh, language processing because, because that's a very interesting idea of uh, um, processing natural language and make it machine readable or understandable. For most of my career I have been working in the smart uh, sensors domain which is not kind of related to NLP. But uh, towards the end of 2022, um, when ChatGPT was released, this was the first time that um, I started interacting with chatbots. And um, so I have a lot of friends who work in the tech industry but are not like in the machine learning. They started asking me questions and I would create these cool demos using these external APIs and start sharing with them. And uh, one of the way I found was useful was uh, to just create videos and start sharing those videos with them. Uh, and then somebody suggested why don't you just make it public and started YouTube channel. So initially I started that uh, just for sharing with my friends, but turned uh, out there was a lot of uh, interest in it and people started watching it and I've uh, been doing it for the last couple of years now.
Speaker B: What was the video that you first launched where you were like oh, I've made it.
Speaker A: Oh okay. So for me it was a little different uh, than I guess like most uh, content creators. I released this video on creating uh, AI avatars. Right. So this was like very early days uh, of animating uh, uh, an image using uh, audio scripts. Right. Uh, and that went. So I think I released it in the first or second week and was just like showing how to put these things together. Right. And it went viral. So within the first week I think it had over a million views. Uh, at that point I was like oh okay, this is probably something serious. Right. Uh, because people are watching it. Right. And uh, I started creating more and more content but I think I have found a more niche of technical content. Uh, right. So I still do it as a part time. Right. It's not a full time thing. Uh, but, but I feel like my audience are now more interested in, in, in, in technical content.
Speaker C: Very cool. So yeah, I mean, I mean that's so interesting. I, it's always so interesting to hear from people like what like kind of their first breakthrough video was. Um, you started your channel to, to connect with, with friends and uh, because they had questions about emerging tech and then um, you uh, know they encouraged you to share it with others. Do, do you know how your, your videos were discovered so broad other networks. Was it just kind of like the, you know, the algorithm gods, you know, looking upon you like what do you think it was that kind of got you, you know, maybe surface? Uh, for a lot of folks.
Speaker A: Yeah. Uh, so I mainly shared it on YouTube and I didn't do uh, any like cross platform posting. I actually got my uh, like X account probably a few months ago. Right. So it wasn't really active on anything else. But seems like there was a lot of interest in the topics that I was covering. Right. Uh, so I think that that helped me a lot. And at that time when I started my YouTube channel, uh, I think there were probably hardly any content creators who were covering uh, technical aspects. Right. Uh, so now like with AI trends that are a lot. But at that point I think uh, the timing was right. Apart from the topic that those are, um, that I work, I was covering.
Speaker C: How would you say right now you say you're a part time creator, you have, but you have this big channel. How would you say that the types content that you cover on your channel now how does that align with the work that you do as a consultant or otherwise? Um, is there like a lot of crossover there or are you still kind of seeing them as separate things that you're doing?
Speaker A: So for me specifically I focus as I said on the technical contents. Uh, so uh, I would divide it into two parts. One is the latest model and news part which drives in a lot of interest from people who are just interested in learning about models or new capabilities. Uh, but the second focus area for me Is uh, my open source work. So uh, that's like building rack systems. Right. And then content around it. And that has been um, a really good source of getting uh, to know people in industry and through those getting um, uh, some consulting, uh, uh, as well. Right. Uh, so I have tried my best to geared um, it towards more technical uh, stuff that is used in industry and that has been extremely helpful in getting to know people.
Speaker B: The bulk of your work seems to be around retrieval augmented generation systems, rag, as you say. Can you maybe for a layman's person explain a little bit of what that is and also why this is so important in the space of AI and machine learning specifically?
Speaker A: Yeah, yeah, that's a very good question. So if you look at the LLMs in general, like Gemini or any of the models out there, they are trained on a finite amount of data and when you train them you kind of freeze them in time. Now in most of the cases what happens is it's going to be trained on publicly available data, but for enterprises or businesses, usually the data is not exposed, so it's private data. So how do you leverage these LLMs, which are uh, essentially chatbots that you can talk to, interact with. Right. So retrieval augmented generation is a technique that you can bring in your private data and expose it to these systems. Right. So uh, rather than them being fixed, uh, in terms of the training time or checkpoint, uh, now you are bringing in fresh data. Right. Uh, so that's like one way of doing it. Right. And I feel like uh, it has a lot of applications in industry. Right. Uh, because it enables you to interact with your ever changing data and at
Speaker C: the same time you can still use what those larger language models have been trained on and what they're good at, but give them the additional context of things from your private information.
Speaker A: Right, exactly right, exactly right. So because when we train these models, these are extremely capable models. Right. It's just the freshness of data and I would say like your business context that is the most critical part. Right. So the models have uh, capabilities that you are leveraging on your own data and that's why it makes it very powerful. Right. So it's not just the model, the data itself, but leveraging, uh, the model capability with your data.
Speaker B: Yep, yep. I can see how that would be really useful to have. And are you in your consulting business? Is this something that you're getting a lot of questions about? Do you find your time mostly consulting around implementing RAG systems or does it range in other areas as well?
Speaker A: Yeah, so uh, it's like flavors of retrieval. Right. So the way I look at it, uh, um, uh, RAG is a search mechanism. So there are components of uh, agentic systems where uh, you build autonomous systems but essentially in most cases you are doing uh, search. Right. And search in itself, uh, I think it's a very important problem uh, in enterprise. So most of my consulting is how to build better search engines uh, for your data and then leverage that uh, search results to do uh, subsequent tasks. So it could be like report generation uh, or automation. So you leverage uh, RAG as a basis but then you empower agents within. So that's, that's how um, I help businesses.
Speaker C: So let's talk a little bit about Local GPT which, which um, is one of your open source projects. It has almost 21,000 stars on GitHub. And I think uh, if I remember correctly this was one of the earliest local uh, um, uh, you know, kind of tools that was, that was available because I worked at GitHub at the time and I remember watching to see, you know, what projects would get a lot of stars. I remember this one being one of the early implementations for people who for whatever reason wanted to be able to uh, run the Rack systems locally using local uh, models. Can you tell us a little bit about what the impetus was behind uh, Local GPT and um, how you built it, uh, to begin with?
Speaker A: Yeah. So the reason behind or uh, the philosophy behind Local GPT was to leverage local uh, models. So even if you look at today, um, a lot of enterprises are uh, not ready to share their Data with proprietary APIs. Uh, that's a big concern. Right. Uh, so the idea behind Local GPT was like okay, the models are getting better, we have better retrieval system. What if you enable businesses uh to use these local models which they can uh, run in um, airtight or air gap infrastructure.
Speaker B: Uh.
Speaker A: Right. And bring in the same capabilities that a proprietary models can bring. Right. So that's how uh, I built the initial version. And the goal was everything is supposed to be local and private so you can trust. Right. Um, and uh, it got a lot of traction on GitHub, uh because there were a couple of uh, uh other projects but at that time local um, GPT like one thing which I was focused uh on was to make sure it uses uh, GPUs, um, and hardware that actually makes it useful. Right? Um, yeah. So that was kind of the start of Local GPT.
Speaker C: And then you documented the process of building it on your YouTube channel as well, right?
Speaker A: Yes, yes. Yeah. So uh, when I was building Local GPT, I was creating content around it. Right. So everything that I was doing, uh, all the learnings are well documented. Uh, so I think that was another reason as well. Probably like a lot of developers were interested in it because, uh, there were a couple of iterations and I was interacting with people, figuring out what worked, what didn't work. Right. So it was overall a community approach, uh, which made it better in every iteration. And, uh, that YouTube content creation, I think played a part in, uh, making sure that people are aware of the capabilities, uh, and the project itself. Right.
Speaker B: It's like you're showing people how to make the sauce while you're making the sauce and experimenting with it.
Speaker A: Exactly, yeah. And learning from people experiences as well.
Speaker B: Right, right.
Speaker A: Incorporating those.
Speaker B: Yeah, yeah, yeah, yeah.
Speaker C: Did it surprise you how quickly Local GPT took off? Was that a surprise to you?
Speaker A: Yes, yes. Yeah. Um, like it started as a weekend project. Right. I saw a project and I was like, okay, here are the things that I can improve. Right. And I started building it and uh, I posted it and actually, um, so it was trending, uh, as number one project for a few days. And I didn't realize it. Somebody actually sent me a link that, oh, it's trending on GitHub. I was like, oh, okay, good. Awesome. Because as I said, I was m. Mostly focused on the YouTube, so I was probably looking at the views on the video that I created. But then it was trending on GitHub and I think it got noticed on, um, uh, Twitter as well.
Speaker B: I was just thinking, because you've built Local GPT two years ago, um, the landscape has changed enormously since. I was wondering if you could give us a little bit of a preview or a summary of what it was like to build it then. It sounds like it was a very collaborative, piece by piece effort. If you were to build it today, I'm sure things would be very different. How so?
Speaker A: Yeah, as I'm releasing, uh, what I'm calling local GPT2, uh, point zero. Right. So congratulations. Thank you. It's very different than the initial version. Right. But coming back to your question, like at that time, um, so we were talking about before any of these, um, coding assistants. Right. Even though it's like just two years, but I think a lot has changed since then. Uh, so much has changed. Yeah. So like, I had to do all of my research, right. I had to write all my code, uh, and then test it to make sure everything works. And then, uh, with these local models, you need to make sure. There is cross platform support so for every operating system. And initially most of the bugs or issues were like oh it doesn't work on Windows or it doesn't work on Linux or uh, my gpu. So it was I think a lot of back and forth now with the second iteration, um, I have worked more as a uh, technical manager and I came up with design uh, specs of how exactly I want things to be, uh, and then used a combination of different uh tools like Gemini Cursor and it's not just one tool but a combination of different tools to do my research and then code implementation testing. So this has been uh, I think a lot bigger um, effort compared to what I did in the first iteration. But it has been much more quicker uh and I hope the project itself or this iteration is going to be much better as well in terms of capabilities.
Speaker C: That's really interesting. Can you talk a little bit more about I guess what has changed? Obviously um, agentic workflows, um, have come out in the last couple of years. M but there have also been new GPUs, new, new um, you know, chipset drivers, um, you know Max went from you know, uh, not having kind of a great way of kind of dealing with GPU things that used to be for local stuff anyway we would only have to deal with like Nvidia, you know, drivers and that would basically be it into having you know, MLX and better stacks there. You know some of the AMD stuff has gotten better too. Has that made it easier for you when building this second version? Like uh, are you able to maybe offload those tasks in kind of your way? You're treating this kind of as a management task or are you still having to deal with some of those implementation issues too?
Speaker A: Yeah, no, I think that these tools have matured enough. So when I started local GPT, um, the only way to run a local model was uh, Llama cpp. Uh and even within LLAMA cpp uh the support for models was very limited based on their architecture. Uh and at that time uh, I think gguf, uh format uh was just getting started. There was AWQ format. So these are different quantization methods that people use. So supporting all of this was really hard because it's like the way I look at it, uh, an airplane is flying and we are building it together at the same point. So it's really hard. Things break now. I think uh, these tools have matured uh a lot. Right? So like okay, uh, people use Ollama vllm right? Uh Those are normally local deployment with llama, CPP and then uh, mlx, uh, Jax I think is also adding support for a lot of different models. Right. So overall I think it has matured a lot which makes development a lot easier. And second thing is even the coding models, right? Or the um, uh, assistive coding tools, the way I want to call uh, them, they have also matured a lot. So like when the first iteration came, uh, going back to the training of the LLMs, they were kind of limited or frozen in time so they didn't really have the ability to look at these new updated packages. Uh, and things change. Right? But now most uh, of these agentic uh, tools have the capability to do web search. Right? So look up the documentation or maybe like MCP supports. Right. So that has I think uh, improved the workflow and even the development time for me. Right. So I've been leveraging all of these different tools and they've been really great.
Speaker B: You talk about this concept of being a technical manager. Can you say more about that? I mean you have described how you're sort of managing these different workflows together, which has been a big change. You know, two years ago you were in the weeds, coding, doing the hands on work and now it's almost like you're outsourcing different pieces to various like LLMs or Gemini or Cursor, um, and you're wearing sort of a different hat. So I'm curious to learn more about this shift and also where you see kind of your role evolving.
Speaker A: Yeah. So uh, with the coding agent or tools, right. It's very easy to build uh, software. Now the problem that I'm seeing is uh, a lot of people are coming in without technical uh, background. So you start building on technologies that you don't know and the LLMs or these identic tools are going to build software the way they want. Right. So let's say if I ask Gemini, here's a project I want, right. Can you suggest a tech stack? Right. So it will suggest a tech stack and then build on top of that. Right. Now that tech stack may or may not be uh, like well optimized for the task that I want. Right? So this is where I think the technical expertise comes into play. Uh, and you want to be very precise on how you instruct these systems, uh, because they are going to go and uh, uh, add a lot of features that you don't want. Right. So in my role, the way I see when I'm building software is uh, I come up with a spec Sheet like okay, here are the specs that I want, right? Here are the list of features that I want, right? And then break it down and then you can really use these coding agents if you make uh, them very focused. So for the focused tasks in uh, iteratively building with these tools. So I think that that's going to be how uh, the future workflow is going to look like for uh, software developers. You um, need to have still like deep technical expertise uh, to effectively use these systems, but it's going to be more on managing uh, uh, these tools effectively rather than uh, writing your code.
Speaker B: Thanks for sharing your secret sauce.
Speaker C: And have you found like with the second version of Local GPT, um, where you've been able to kind of give it these requirements and stacks, um, have ah, you been able to implement features that maybe you would have shied away from implementing before, um, because you had the awareness of what you want to do, but maybe it would have been too time consuming or otherwise. Has there been anything, I guess, um, have you been able to do things that maybe users have requested that you might not otherwise have committed the time to? Has anything with that changed?
Speaker A: Oh yeah, yeah, yeah. So like I consider myself uh, as an ML researcher or maybe a backend developer at most. Right. So I never had uh, um, the background to build front ends. So that's been like one of the things. Like even the initial version that I created was a simple streamlit app, right? So uh, it's functional but not really uh, pretty. Right. But with this new version, uh, since these tools can actually build front ends and really nice ones, uh, I was able to actually implement ah, those and it's like all right, especially with the multimodal ones, it's amazing because you give it a template, this is what I want, and then you iteratively build with those. So it's been amazing. Some of the things that's just like one example, right, of something that even today I can't do it. But with these um, uh, agentic coding tools, uh, we're able to build new and new capabilities. Uh, and that's just as I said, one example. There have been a lot of other things, um, like the amount of documentation, for example, um, uh, I'm able to read through the system is just amazing.
Speaker C: Yeah, it makes um, testing easier too, right? Because you were mentioned before when you're building the first version, you're having to write all those unit tests yourself. Uh, there were some uh, AI coding systems, but not to the extent they are now. And so I'm sure that you can have wider test coverage too now um, when you're building stuff out um which is also going to make for a better result.
Speaker A: Yeah definitely. Um, especially I think it's a normal problem with developers uh when you're developing uh you don't really have enough time left to even have full coverage. Right. But with these tools definitely it's a ah great value addition uh and it's been um, really great uh in that sense as well.
Speaker B: Can we go back to. You said you're now able to read more documentation than ever before. Is that what you said?
Speaker A: Yes.
Speaker B: Say more about that. What do you mean?
Speaker A: Okay, so normally let's say if I'm working on developing uh, or integrating a certain API, uh I'll go and look uh up the documentation for that. But the problem is um, you'll probably find an older version of the documentation which is not even relevant anymore. So you have to sniff through all of those things. But now with uh these agentic tools you could do a wider search. Right. Uh, uh if you use them properly they will really uh narrow down the search area and can actually figure uh out or um show you more uh relevant documentation even especially something like deep uh research. It's been extremely valuable because that can find uh things that you would miss when you're doing like a human search.
Speaker B: Right, right. Was there something in when you were using these agentic coding tools that really you were like wow, I really see the power of these agents and I'm really going to lean into it and leverage it.
Speaker A: Yeah. So one um, example would be uh, like going back to a prior conversation integration with these different uh uh packages right. So going back to the LLMs like if you want to support uh Mac os. Right. And then it's like there are uh Apple Silicon versus like the, the older uh intel. Right. And then a whole bunch of different uh GPUs from Nvidia. Right. Or TPUs. Right. So for me uh, it would be like an impossible task to do it in a managed time. Right. But with leveraging these tools you can easily build that. So that's been pretty awesome uh and a big unlock especially for me. And then certain times you run into bugs um in some of these search tools will be able to find sources which you uh probably won't even notice.
Speaker C: So obviously uh there are people in the community have contributed to the project uh too but this is primarily been like a solo project ah of yours um, how much would you say like and I don't know what your open Source experience was like before Local GPT, had you managed um, or maintained any um, open source projects before Local GPT?
Speaker A: No, no, this was the uh, first one.
Speaker C: Okay. So I mean uh, again I'm somebody who's like, comes kind of from the open source world and that can be a lot uh, to undertake. This is one of the reasons why I asked if you would expect it to be as successful as it was. Because on the one hand that can be really exciting and on the other hand that can be overwhelming when you have that many pull requests or issues and that many people, you know using your stuff. And then now you're like oh do I have to support this? How do I maybe integrate these things into it? Um, especially if you're doing this as like a weekend project and like as a solo developer, um, I don't know if you could maybe talk a little bit about what I guess is, you know, the last two years and even the last you know, year or so or even six months as changed or has anything changed that has made it easier for you to manage a large project like that as an individual developer?
Speaker A: Yeah, so it's definitely been a struggle uh, to be honest. Uh, as I said it was my first endeavor into uh, open source project. Didn't expect uh, the success it had and then it comes with all of these issues and pull requests, things like that. So initially the first version was very hectic because I work full time content creation so it was much harder. But now for the second version, even during development times I have been leveraging some uh, of the AI tools so there are automatic code reviews. Uh, I've been uh, using a combination of different tools from different companies. So I hope, I don't know yet when the second version is released. I hope it's going to be a lot smoother journey. Uh, but one thing I want to uh, point out is I leverage these agentic tools uh, a lot. But you need to be very careful of how you use them.
Speaker C: Let's talk more about that actually because I think that that's a really important thing for people to know about. What are the things that you need to look out for when you're using these tools and you're building these systems?
Speaker A: One of the normal thing that we have seen specifically with Wipe coding, people build software and it's great for demos right. So you can, within a weekend you will be able to build something really cool without even any coding background, which is awesome and it's enabling new people to take uh, on development. Uh, however, um, if you start building more scalable solutions that you're going to uh, support, let's say thousands or hundreds of thousands of people. Then uh, you start seeing holes in the implementations, right? So like a simple example I would uh, give it is when you start building features, uh, with these um, uh, agentic code systems, right? They were going to start uh, forgetting about it, right? So let's say I build a feature, right? I, uh, built three more features on top of it, right? Now I want to reuse that feature. And all of a sudden you are going to see that um, the agent starts reimplementing that feature again because it forgot about the prior implementation, right? No, yeah, right. Or a second uh, uh, issue that I'm currently running into, uh, is this right? So let's say uh, it wrote a function and that function is being used or called in 10 different places right? Now you want to add some more functionality. We change the input outputs of that function. What ends up happening is, uh, because of the limited context, it's going to update the input outputs and let's say five of those function calls, it will forget about the two other ones, right? And now it's like, all right, everything works because you test it. You run the uh, initial tests and it might be just using those two function calls, but when you start putting this into production, you're going to see those cases, uh, where it starts failing. So that's why it's very important to understand your code base and uh, review every change that it makes, uh, because you're going to miss out on those edge cases. So those are just a couple of concrete examples that comes to mind, uh,
Speaker C: when you're talking with, uh, either with some of your consulting clients or people who watch your YouTube channel who might be, um, junior developers, um, and it might be, you know, comfortable with the AI tools but maybe don't have as deep of a background. What advice can you give to them, I guess to maybe use these things effectively and not, you know, run into some of these prattles.
Speaker A: Yeah. So, uh, the way I'd see it is these are great learning tools as well, right? So, uh, when you are building them, um, uh, with these, like, the good way of learning is ask it why did it choose one implementation with the other against the other. So that gives you a really good idea of kind, uh, of peeking into the mind or uh, thought process of these systems and then leverage that for your own learning to build better systems. Uh, but I still feel like the traditional system design concepts, uh, are very important. So you still Want to have uh, that basic understanding what are those concepts. So you want to see uh, or you want to make sure that you are putting these system uh, together uh, to be scalable, uh, and how different APIs are going to interact with each other. You don't have to know the coding aspect of it, but you want to make sure that you understand how to put these systems together. So that's I think a most important thing that needs to happen and just to um, uh, uh, add to it. There is another trend that I'm seeing which is I think really unfortunate. Um, most of these systems are data driven. Right. So whether you are building a Rack system or any Agentix system. Right. The trend that I'm seeing is people do not want to spend time looking at their data. Uh, yeah. Right. And they expect magic from these systems. Right. Which is uh, uh, like it's not a sexy job or uh, thing to actually spend time on looking and understanding your data. Right. But it's something that needs to be done. Right. And if you don't do that later on you're going to run into issues.
Speaker B: Yeah, it's super important. I mean we talk a lot about in sort of AI principles. I mean the quality of a product starts with the value of your data and whether it's properly cleaned, it's been you know, organized, it's being rid of biases. Um, yeah, data seems to be the biggest place to start building a product that's going to be really good to all of its users.
Speaker A: Yeah, totally. Especially like in enterprise or real world data is very messy. Right. So you need to make sure that you uh, properly pre process it or anyone like in some cases you want to have uh, post uh, processing or safeguards. Right. But that's a trend that unfortunately I think not especially like the new developers are not paying too much focus to.
Speaker B: Yeah.
Speaker A: Which is critical. Right.
Speaker B: How are you addressing that? As you said, it's not a sexy job to go look through your data. So what, what incentive do you give them that sort of switches that light bulb and they think oh I actually this is really crucial, I need to do that.
Speaker A: Yeah. So the way I look at it is if you start looking at your data, uh, you're going to actually reduce your development time. So a simple example is like again going back to rag systems. Right. So people, when they build RAG systems, so you need to worry about how do you chunk your data, uh, what is going to be the chunk size, what type of embeddings to use and so on and so forth. Right. And what I have seen in industry is people are going to try 10 different chunking methods and see, okay, which one performs better. Right. But a better approach is just look at the structure of your data and based on the structure of the data, uh, uh, chunk it the way a human would look at it. So if you, an example would be if you're looking at a research paper, usually you have an abstract, uh, there is introduction, there are subsections, uh, that is probably a really good way to chunk your data and that will require understanding, looking at your data and that will essentially, if you um, uh, start with your data, understanding the rest of the pipeline becomes a lot more easier. Uh, rather than praying to LLM gods that okay, they're going to figure it
Speaker C: out, right, that it's going to do everything for you and it's going to not overwrite something, um, uh, or give you the wrong response. What are some of the differences that you're seeing? Uh, I guess right now, especially when it comes to comfort with AI tools between maybe uh, more junior devs and more senior devs or are you seeing any differences?
Speaker A: Yes, I think uh, I see two extremes. Uh, one is junior developers look at these AI tools and they want to go all in because they can uh, really generate code at a rate that was not possible before. Uh, when it comes to more established senior developers, uh, a trend that I have seen is they probably tried a coding assistant or a coding tool six months ago and based on their initial experience. Okay, I'm not touching this at all because these are so bad. Right. But now I feel like as a developer in general you want to make sure to leverage these tools irrespective of whatever um, um, your uh, experience level is. Actually before that, let me point out there was a very interesting paper that was released uh, last week from a company called Metr M E T R. So they looked at um, it's a very small sample size. So it's I think around 16 developers if I'm not wrong. So they looked at 16 developers uh, who are contributing to open source projects. And the study premises, uh, was how much uh, speed up LLM coding assistance can bring. Uh, so initially uh, the estimate was we're going to get around 20 to 25% speed up in code development. Right. At the end of the experiment they found that with experienced developers there was um, around 19 or 20% reduction in the amount of quality code that will actually build. So it turned out that in that study it's actually slowing people down rather than uh, speeding Them up. Right.
Speaker B: Oh, I've been hearing about this. Yeah, sorry, keep going.
Speaker A: Yeah, but then uh, I don't want to generalize it. Right. So it's just a single study, uh, uh, very small subset. Right. But I think we need more and more studies, uh, like that. Right. To actually see the impact. Right. And then um, another thing is like uh, we have yet to figure out what is the proper uh, interface or UI UX for these coding tools. Right now everybody's just focusing on uh, chatbots. Right. But that may not be the right um, interface. If you look at some of these LLMs, um, or actually like the LLMs in general, uh, every one of the LLM or coding agent has its own personality. So you really need to adopt to that personality when you're using it. Uh, so I feel like we still have to figure out a lot about uh, these systems. But uh, in general, uh, my recommendation is like yeah, we need to adopt these or include these in our workflows.
Speaker B: What do you say to the junior developers and what do you say to the senior developers? I mean you've kind of given a bit of an overview, but if a junior developer is coming to you saying yeah, this is everything, I want to do everything with this, what do you say? And then if a senior developer says no, I'm done, I'm not touching this. What is your way to sort of bring them in the middle, uh, and get them to see how to use these tools in a way that is a bit more balanced.
Speaker A: Yeah. So for junior developers my recommendation would be do specs driven development.
Speaker B: What does that mean?
Speaker A: So you need to have very specific uh, uh, specification of your product. Right. So normally when we are doing a product development we have uh, PRD that drives the product. Right. Uh, so we probably want to have a very similar specs for software development as well. Right. Especially when we're using these tools. And then at each step you have to review, you can't blindly trust what the output is. Uh, I think another piece I would add is iterative development. Uh, so incremental feature, uh, development rather than giving these systems. Here's my wish list. Right. Go and implement it.
Speaker B: Execute.
Speaker A: Yeah.
Speaker C: So instead of trying to one shot it, you're like actually let's multi shot it. Let's do this step by step.
Speaker A: Exactly. Right. That's how we as humans build. Right. So of course, yeah. And that should be how we uh, interact with these systems as well. Right. Uh, the current iteration I think uh, is going to require a lot more uh, handholding Than like, okay, just one shotting everything. Right. Uh, so that's for junior developers. Right. For senior developers, I uh, think we uh, need to look at the capabilities of these systems. Right. And then figure uh, out how um, they can be used. Like one example would be test, ah, coverage. That's something that can easily be done with these systems and it actually saves a lot of time. So uh, going back to my example of being a technical manager, so for experienced developers I think it's a lot more easier because they have a better understanding of leveraging um, these tools with their experience. So that would be the role that I uh, would recommend for people who
Speaker C: might not be able to keep up with everything that's going on. Um, do you have any advice for how they can start to uh, experiment with these systems?
Speaker A: So one way of doing it, I guess I started doing it was I looked at some of my uh, older projects that I have built and tried to rebuild them with uh, these AI tools. So that will give you a really good idea of both the capabilities and limitations. Right. And since you already know the code base for example. Right. So you have a very better sense of what goes wrong. Right. So that's, that's how I would uh, approach it.
Speaker C: When you've done that with your other um, with some of your older projects, have you either been like impressed with some of the solutions that have come up with that might have been different from what your implementation was or, or um. Has there ever been anything that's kind of surprised you where you've gone? Oh, okay. I did consider doing something this way.
Speaker A: That's a good question. Um, like it's, it definitely has improved the code writing speed. I don't think I have done this with reasoning models. So reasoning models may come up with a better solutions. When I was testing initially I was definitely impressed with the uh, uh, implementation. I don't know if I can say that for the code quality. Right. Uh, sure, yeah. Right. But like the reasoning models, some of the newer ones definitely can uh, uh, come up with solutions which are out of the box. Uh, I use some of these to brainstorm ideas. Uh, now one thing which I don't like about LLMs, they are very um, agreeable. So you're like, all right, here's the feature that I want to implement. Right. Okay. And they'll be like, oh, this is the best thing in the world.
Speaker B: Yes.
Speaker A: Okay. So you'll be like, all right, here's what I want to do it. Like, this is how I want to do it. Right. They'll be like, yes, this is the perfect solution. Right. And, uh, excellent idea. Exactly.
Speaker B: Well done on you. Yeah.
Speaker A: Yeah. So, like, from that perspective, I don't know, like, if I can trust, like, the design choices to be honest at this point.
Speaker C: That's such an interesting thing to think about that, like, they're too agreeable. Kind of need it to be like the stereotypical, like, you know, grumpy coder, um, who's going to be pushing back. Right. Like that more established senior dev who's like, no, no, no, no, no, this isn't right. Um, as your code reviewer, that's really interesting to think about.
Speaker A: That would be a great product. I would use it.
Speaker B: Yeah. The subjective LLM.
Speaker A: Yeah.
Speaker C: Yeah. Judgmental. Right? I like that.
Speaker A: Yeah.
Speaker C: Judgmental. I'm going to judge your code, uh, and roast it. I like that idea, actually. Someone please come up with that. Someone listening. Please come up with, Please.
Speaker A: With.
Speaker C: With the Persona for these things to roast and judge, um, and give like, critical feedback on these things.
Speaker B: I wanted to shift. We hadn't kind of planned this, but I wanted to know a bit more about which of our Gemini models do you use, how do you use them, what's been the most helpful part of them, and why did you choose them? Yeah.
Speaker A: I have been, uh, using Gemini more and more, um, kind, uh, of going back to that ideation. Right. Um, I think the long context, uh, for the Gemini models, it's definitely, um, really, uh, good. Right. Uh, like some of the other, uh, model providers do offer long context right now, but, uh, Gemini. I think the, uh, Gemini team actually, um, figured it out. Right. I think it was one of the earliest ones.
Speaker C: Can you just explain for our listeners what the importance of the context is? Um, this goes along with some of the. Or similar to some of the rag things. But can you just explain a little bit why, um, having more context is more useful?
Speaker A: Yeah. So LLMs have this idea, uh, of how many, uh, tokens or how many words of information it can process. Right. So every model that gets released has, uh, a certain amount of. Of, uh, information that it can process in a single shot. Right. So we call it like context Window Gemini. Uh, I think when, uh, 1.5 was released and correct me if I'm wrong, like if I'm off. Right. So they were offering 1 million context window, which was, uh, uh, like industry, uh, the biggest one in the industry. Right. Most of them were.
Speaker C: I think it was almost 10x, you know, similar to what the other models at the time were doing.
Speaker A: Yeah, yeah. Right. And especially like for coding tasks, right. If you're working with uh, large, uh, code bases, right. Uh, the context window plays a very critical role. Right. If the model is just looking at part of the code, it's not going to have a global understanding of your code base. Right. So you want a model, uh, that can look at larger portions of your code base. Right. And for programming I have found to be a lot more useful, uh, compared to some of the other offerings. Right. And that's specifically talking about the long context. But I do want to point out, um, the Flash and Flashlight models. So I'm a big proponent of you don't need the best or the biggest model for every application. You want to choose a model based on your needs. Right? So the Flash models, they offer the same context window of 1 million token, but they are much faster. Right. And uh, if you use Gemini 2.5 Pro, um, uh, as a planner to plan and then it's like, okay, here are the different features, uh, that you implement. Here's how you are going to implement. Right? You don't need the Pro model anymore to actually implement those. Right. You want a much faster model that can generate a lot faster code. Right. So for those kind of cases, I have found, um, Gemini Flash to be very helpful, very useful. Uh, yeah, so that's, that's kind of like m my personal uh, use case.
Speaker B: To all Those star beginner YouTubers out there, what advice do you have for them for starting their own channels?
Speaker C: Because it's obviously a much more crowded space now than it was when you were getting started. But would you have any advice that you want to give to someone who's wanting to start a, uh, more tech focused, uh, you know, coding focused YouTube channel today?
Speaker A: Okay. That's a hard one. Okay. Because I don't know whether I'm doing things right or not. But my suggestion would be figure uh, out a niche. So in AI space, if somebody's just covering news, it's not going to last long in the long run. So you want to figure out what actual value you are bringing in. And that's based on, let's say if somebody has technical expertise, being a technical creator. So that's, I think the most important thing that you want to differentiate yourself rather than like just hop on in this hype train. Right. Uh, and second thing is based on my experience, content creation is a lot harder than it looks like. Right. So, uh, like When I started YouTube channel, it was just for myself, right? Like, okay, um, I didn't have any plan and uh, I was just posting videos. Whatever I found interesting right now it has grown to a point at which you kind of start thinking about your audience as well. Uh, and you want to make sure that people, uh, who are spending time on your content, you provide them with valuable, uh, um, information. Right. So it becomes a lot bigger responsibility as you grow. Right. Because then you have to think about every video or every post that, okay, am I wasting somebody's time? Or it's actually valuable. Right. Uh, so when it starts growing, you need to consider that aspect as well, because I think, uh, for people that don't realize that. Right. Uh, it's a huge commitment.
Speaker C: Is there anything that knowing what you know now that you would have done differently?
Speaker A: I would have focused a lot more on the technical, uh, things. Like one thing which I have tried my best, uh, is not to hype up things. Right. Uh, but I might have done some of the videos, like, along the line. Right. But I think I have matured a lot now. Uh, but from the beginning, um, that's something that, um, I would love to do it again. Like, make it a lot more focused, uh, make it more, uh, uh, like I see myself as an educator rather than an influencer or entertainer. Right. Uh, and that's like, uh, if I had to redo it, I would simply focus on, uh, educational content.
Speaker C: That's. That's really interesting. And I think that kind of goes back to what you were saying before about, like, understanding your niche, understanding what differentiates you. And like, for you, it's the fact that you do have this technical background. You are like a, you know, somebody who, you know, you have a PhD, you've studied this stuff, you've worked in these systems, you've consulted for a lot of companies. You are an influencer in the sense that, you know, uh, people look to you and take your advice, but you see your yourself as you're wanting to educate, you're wanting to, uh, show, uh, people from a technical level what needs to be done.
Speaker B: Yeah. And you're bringing developers along the journey of you exploring and experimenting with things. Sometimes I think when you watch these videos, it's like, here's how you do it done versus the journey that you're on, you're taking developers on is here, let's build this together. And then you can see the ins and outs of how hard this is or how easy this is. And I think there's a lot of value in that storytelling in a way.
Speaker A: Yeah, yeah, yeah, definitely. And actually, like, through the process, I have learned a lot from Other people as well. Right. Uh, that's another aspect like the uh, uh, if you build audience, right. You're going to get feedback. Right. And there are things probably you didn't cover in your video. Right. People tried it and like, so it's a whole community effort. Right. And one thing that I do want to say is, uh, content creation has uh, helped me, uh, meet interesting people and network. So that has been like a huge value added. And if people start recognizing you, that's the best thing that can happen for your career as well.
Speaker B: Yeah, absolutely. Absolutely. Okay, I think it's time for some rapid fire questions. Christina.
Speaker C: Let's do it. All right, we're going to get you started. We. What was the last thing that you automated for yourself? Assuming that you did like maybe that you let either AI do for you or wrote a script to do something for you.
Speaker A: So I do a lot of data scraping for some of the project. Uh, so I think that would be one and um, uh, test uh, coverage that's pretty much automated now.
Speaker C: Love it, love it.
Speaker B: What was the last thing you asked? A large language model.
Speaker A: Okay. Um, so this morning I was testing different uh, things and one of the question that I had was uh, uh, I asked the LLM or a number of different LLMs that pretend if you could write a letter to humanity what you would say. And that was like, I would actually recommend everybody to test this on different LLMs and see their responses.
Speaker B: What were some of the interesting responses or surprising responses that you saw?
Speaker A: So one of them, uh, was, I think, a lot more human. And it talked about, uh, how humans treat other humans. So whatever is happening in the world right now, whatever we are doing to the environment. So it started that letter specifically pointing those things out, which was very interesting because I, uh, wasn't expecting that. But I highly recommend, uh, to test that.
Speaker C: That's awesome. That's awesome. What can you do now that you couldn't do six months ago?
Speaker A: Oh, uh, uh, I couldn't do web development. Web apps, front ends, I can do them now.
Speaker C: Amazing. Amazing. Yeah, I'm kind of the opposite. Like I'm a front end person and my backend is now a lot stronger than it was six months ago because of uh, the agentic systems. So I love that.
Speaker B: What are the next sets of problems you're trying to resolve?
Speaker A: Okay, so we are entering uh, more agentic era, uh, where agents uh, are able to take steps. Right. So I'm actually uh, thinking about how do you leverage those to not only aggregate information but do uh, things or automate things more effectively. Right. So that's kind of the direction I'm moving in now.
Speaker C: And then finally, what is your. And this is kind of a big one. Um, what's the biggest learning that you've had maybe in the last year about, about the future of work?
Speaker A: I would say don't limit yourself, uh, in terms of the capabilities, both from a human perspective and these systems can bring. Uh, that has been a huge unlock, uh, for me thinking, uh, about what's possible, uh, and looking at the future. What is going to be possible with this system? Uh, I think that's going to be huge.
Speaker B: Couldn't think of a better way to end. What a great answer.
Speaker C: Really, really fantastic. Um, and, uh, go ahead. And we will have a link to this in our show notes. But, um, people want to follow you on YouTube. Uh, you are YouTube.com engineer prompt.
Speaker A: Uh, yes, niereprompt. That's my YouTube channel. Um, uh, I also have a Discord server, uh, where I interact with people. And uh, you can also find me on X with the same Engineer with double Rs at the end prompt.
Speaker B: Great. And we'll have everything in the show notes.
Speaker A: Thank you. Thank you.
Speaker B: Thank you so much for being here. What a great conversation.
Speaker A: Thank you. Uh, I really enjoyed the conversation. Thanks for having me.
Speaker B: Thanks for joining us. For more information about our guests today, please check out the description.
Speaker C: And for more great conversations like this, hit that subscribe button. Until next time, thanks for listening.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.