
TestGuild Automation Podcast · 2026-07-07 · 44 min
Key moments - from our scoring
Substance score
62 / 100
Five dimensions, 20 points each
Amit Rawat shares his philosophy of agentic engineering and showcases Prompt Write, a desktop application that translates plain English prompts into working Playwright automation code. Rather than viewing AI as a black-box solution, Rawat advocates for meticulous upfront planning using tools like Claude and Copilot - starting with brainstorming, creating detailed plans visualized in HTML, and gathering feedback before execution. His approach emphasizes that success depends not on fancy prompts but on clarity of vision, taste, and curiosity. Rawat argues QA professionals are uniquely positioned for the agentic era because they already understand test planning, edge conditions, and domain analysis - skills that translate directly to steering AI agents effectively. He also discusses how technical depth and understanding of underlying systems (versus blind 'vibe coding') creates competitive advantage, and touches on loop engineering as a future shift from 'human-in-the-loop' to 'human-on-the-loop' testing where agents run autonomously on schedules rather than waiting for human prompts.
Prompt Write is a desktop Electron application that translates natural language prompts into working Playwright automation scripts. It uses AI agents (powered by GitHub Copilot SDK or compatible models like Moonshot) to understand the intent, execute browser automation in headless Chrome, stream a live screencast of execution, and generate deterministic Gherkin test scenarios from the observed actions.
Detailed planning ensures AI agents receive clarity and context rather than vague instructions, reducing token consumption and increasing execution quality. By brainstorming, building detailed plans, visualizing them in HTML, and gathering feedback before execution, the quality of the final result depends on plan depth - potentially reaching 90%, 95%, or 99% completeness.
Vibe coding is trial-and-error prompting with no underlying technical knowledge, relying purely on observing outcomes. Agentic engineering requires understanding the technical layers, design choices, and how systems work, enabling you to steer AI agents with intention and design rather than just guessing and measuring outputs.
QA professionals should emphasize test planning, domain analysis, edge-case identification, and understanding interdependencies - the 'brain' of testing. Technical implementation details like page objects and code frameworks will be commoditized by AI, so the differentiator is taste, curiosity, technical depth, and ability to identify what actually needs testing.
The record feature observes browser events in real-time using Chrome DevTools Protocol (CDP), then allows you to give AI a direction - such as 'generate multiple positive and negative test cases' - and it analyzes all recorded actions and creates derivative test scenarios automatically.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains solid practical insights about agentic engineering, planning before prompting, and the shift from technical implementation to domain expertise in QA. However, much of the content is demonstration-heavy and repetitive (showing PromptWrite features multiple times, explaining loop engineering from several angles), diluting the insight-to-time ratio. The core ideas - planning mode, memory layers, scope specialization of agents - are valuable but not particularly novel to operators already familiar with AI workflows.
So I don't write a prompt. Okay, go and build this. I start brainstorming
test planning and whatever the expertise testing we do is the brain of the testing and automation is just the robotic arms of the testing
Amit's perspective on agentic engineering for QA is somewhat original - reframing testers as well-suited for the AI era due to their domain understanding and systems thinking - but the underlying frameworks are increasingly common in AI discourse. Loop engineering, planning-first prompting, and memory layers are now widely discussed. The specific application to testing automation is somewhat fresh, but the core mental models lack true contrarianism.
why QA professionals may be the best suited people for this new agentic era
loop engineering is a glorified cron jobs
Amit Rawat is a practitioner with genuine depth: he's built PromptWrite, runs multiple AI agents in production (chief of staff, health officer, finance tracking), and has spent two decades in QA before pivoting to agentic engineering. He's not a pure theorist or influencer - he's shipping real tools and living his frameworks daily. However, he's not a marquee name (no Fortune 500 background, no massive exits), so while credible, the caliber is solid mid-tier practitioner rather than top-tier operator.
I've spent like two decades in qa
I created a telegram bot which is written which, which speaks his language, which is his native mother, um tongue Hindi
The episode includes concrete examples (PromptWrite's Gherkin generation, Mac Mini setup, GitHub loop agents running every 5 minutes, $200/month Claude Pro budget, Moonshot model at 1/10th Claude's cost), but much of the specificity is about Amit's personal setup rather than generalizable data or external case studies. The PromptWrite demo is visual but lacks metrics on success rates, token efficiency gains, or time savings. Claims about QA being 'best suited' lack empirical validation.
I have a $200 cloud code subscription
it's running every five minutes. It goes through all my repositories on GitHub
Joe asks decent follow-ups (on 'vibe coding,' on cost ceiling, on reliability of different agents) and pushes back on framing, showing genuine curiosity. However, many of Amit's answers are long, tangential monologues that Joe doesn't interrupt or redirect sharply. Joe doesn't challenge weak claims (e.g., 'testers are best suited for agentic era' lacks evidence but isn't pressed). The host allows Amit to control the pacing and depth rather than surgical questioning.
Okay, good, good, good. Let me know why and what you would call it
I don't know how you get it to be reliable. I've been trying openclaw and it always has issues
Computed from the transcript - who did the talking, and the words that came up most.
Amit Rawat is an agentic engineer who spent two decades in QA before shifting fully into building AI agents. He's the creator of PromptWright, a desktop tool that turns natural language prompts into automated Playwright browser tests, complete with screen recording, Gherkin scenario generation, and self-healing locators. In this episode, Amit and Joe get into what it actually takes to work with AI agents at a high level, starting with why the planning phase matters more than the prompt itself. Amit breaks down his own workflow, brainstorming with AI, building a detailed plan in HTML before ever executing, and why curiosity and technical depth still matter even as AI gets more capable. They also cover why Amit believes QA professionals, more than developers or DevOps engineers, are best positioned to thrive in the agentic era, how he tracks the ROI on his $200-a-month Claude subscription, the "Chief of Staff," "Chief Health Officer," and "Chief Financial Officer" AI agents he's built to help manage different aspects of his personal life, and how he uses a memory layer so those agents understand his preferences and become more useful over time.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Hey, what if I told you that in the next 12 months, most manual test cases won't be written by a human at all? Today I'm talking with Amit Rawat, an agentic engineer who built an awesome tool called Prompt Write that turns a plain English prompt into a working playwright automation script. And I asked him to share his screen as he shares how he automated a huge part of his own life, from creating a chief of staff AI that runs his LinkedIn and YouTube, to health tracking bot, to. To a telegram assistant he built for his dad in Hindi. We also get into how he plans before he writes a single prompt, why QA professionals may be the best suited people for the new agentic era. We also get into how he plans before he writes a single prompt, why QA professionals may be the best suited people for this new agentic era, and how he decides if $200 a month in AI tokens is actually worth it. You don't want to miss it. Check it out. Hey, me. Welcome to the Guild. Hey.
Speaker B: Hi, Joe. Thanks for having me.
Speaker A: Great to have you. I've been following you a while on LinkedIn. You're always posting really cool stuff, so I thought, hey, it'd be great to have you on the show. I think you even created something called, uh, Playwright a few months ago that I featured on the new show. So I thought it'd be great to have you today.
Speaker B: Yeah, I think that's prompt, right? Not playwright. But yeah, I think last year I created it, uh, uh, and the pace with which AI is moving, it kind of became obsolete. Though, uh, I realized it has 100 plus stars, so I decided to revamp it. And because you talked about it, so before preparing for this, uh, kind of discussion, I revamped it. Maybe we'll briefly show what I've changed in that prompt. Right. But yeah, it's prompt, Right?
Speaker A: So you said you had to revamp it. You know, we're moving so fast with AI, and you call yourself an agentic engineer. How do you keep up even with what you've created?
Speaker B: Um, yeah, so I think with the, with the AI agents, uh, I think that the most valuable thing is the idea, like what you want to build, uh, and how you want to build and the kind of clarity, uh, you need, uh, like for that idea, not at a very abstract level, but to a, ah, kind of a certain depth. And then once you are done with that kind of work, like it's AI agents can take care rest of the things. So that's why, like, uh, when I decided okay, prompt, right. You wanted to talk about it and I could see it's quite popular. It's at 100 plus stars. But then I realized okay, I built it like one year back at a time the capabilities were quite different and now we have like fable and clotsonit 5 so maybe I should be more ambitious. So I had some few other features and convert it to a desktop electron app. But uh, yeah, I think just to answer your question, I think it's more about the ideas. Now if uh, you're good with agentic engineering, uh, then how you want to execute them, you just need to fire a few agents and after like maybe a couple of hours you will see that your idea has kind of manifested into what you want it to be in a one shot most of the times.
Speaker A: So people hearing that might uh, be scared but you seem to embrace it. So are you seeing anyone with idea could do it one shot or does it take you up to like maybe 80% or have we gotten to the point where you could one shot it and it's probably 90%? I don't know.
Speaker B: Uh, so as I said right. So uh, how you want to, how you visualize the idea is really important. So like for most of my work I start with a plan so I don't write a prompt. Okay, go and build this. I start brainstorming and I'm a big fan of like Boris Chenry who is the creator of Claude Code and the Peter Stevenberg who is the creator of openclaw. So I've heard most of their conversations and everybody like all the pro kind of uses the plan mode first. So what I do, I start brainstorming with the AI I wanted to revamp. Prompt, right. This is what I built initially one year back. These were the capabilities, that's why I had these features. Now my plan is to be more ambitious and let's brainstorm. So I ask like AI to discuss with me. They throw some ideas, I discard few of them. I kind of ah, consider a few of them and then refine them together. And then I asked him as the AI to build a uh, detailed kind of a plan. And most of the people get overwhelmed when they see a big plan in a markdown format in a terminal or in a VS code kind of a copilot interface. I think that's a new trend. Now I use stml right. So I asked, okay, whatever plan you built, just build it in a fancy STML so that I can see it like a website and I can go through each kind of uh, uh, area of the plan and then with some notes. So in that HTML, uh, it will give me a placeholder where I can give a feedback about a section, why we are doing this, uh, follow up questions. And once we finalize the plan then I say okay, go and execute it. So it depends on the clarity of the plan, how much detail the plan was. Uh, that depends that we do the last mile. Kind of like kind of a feature completeness, like the 90%, 95th or 99th percent. It all depends on the plan, how detailed the plan is. Got it.
Speaker A: So it sounds like you're not just given blind trust in this. You actually are putting up all the upfront work before you commit to even doing a prompt. Which sounds like uh, is the recommended way you think everyone should approach this.
Speaker B: Yeah, I think that's that, that's that brings a taste element. Right. For your creation. Because if it's just about the prompt that go and build a CRM, then anybody can do it. Right. And it will build. So as they say like these uh, like AI models are probabilistic, kind of a, kind of a model. So they tend to go to the, they tend to average out the entire knowledge of Internet they are trained on. Right. If you want to do some deviation from that average, which anybody can do it, you have to bring your own test taste element, your own agent engineering kind of techniques. Uh, that's kind of a differentiate uh, you. And that's the reason I try to like start this kind of initiative around agentic engineering. Because I could see that that will become eventually a very like uh, a kind of a journalist kind of a skill for all kind of a knowledge work. And everybody should kind of focus uh, more on the agent tech engineering.
Speaker A: Got it. So what are some skills you think someone needs to make this transition? Because it is quite a leap. I think it's almost like going back in the day where you had to be a product expert or a domain knowledge expert. Um, I don't know if that's true or not.
Speaker B: Uh, well, so I think um, so it's very difficult to articulate like what exactly needed for this transition. But I can guess few things which kind of kind of worked for me. Uh, I'm a highly curious person. So like whenever I see something new I try to dig down and see how exactly it is working. And, and for me it's really important, uh, the excellence in the product, uh, not just getting things done. So like I've spent like two decades in qa, so I used to have Debates with my developers, not like for a smaller, like kind of a UI UX issue. So if you are a kind of a perfectionist person and you are quite ambitious to build things like shape of different features, uh, then it really helps. And another thing I would say, um, if you are a technologist at heart, I know there are like two camps on Internet. Some people say now you don't need to learn technology. And like they're like all the software engineers will go out of job. But I kind of uh, uh part of the contrarian, uh, uh, kind of a view where I think if you are more technical, if you understand the underlying wiring, how the things are working, then that gives you extra advantage because now you can be more creative, more innovative how you want to wire these things, how you want to connect all the different tools. Uh, because the AI model is as good as what tools, what all kind of uh, like hands, which we say uh, what all tools, mcps, whether it could be clis, it could be any other access to different kind of data. Uh, and also what kind of context, what kind of prompts you are uh, giving. So I would say your taste, your curiosity and like what kind of tools and what kind of uh, powers you have given to AI agents. So all these three things kind of uh, makes you really powerful if you kind of nail all these three things.
Speaker A: Awesome. So I want to get into AI agents, but I want to wait. Um, um, before we move over, um, can we talk a little bit more about prompt? Right. Can you show it really quick and maybe give people a little taste of what it is, what you've done?
Speaker B: Okay, sure. Okay. So uh, I think I started it around one to one and a half year back. At the time like the. There are different waves of kind of uh, AI agents which were trying to automate the browsers. They were like few popular were like Playwright, mcp. There was a browser use, there were one like CLI by the Vercel team. So everybody was trying to build and they're trying to drive the browsers. So I realized okay, maybe I should also try to build something where from a prompt, uh, you can kind of uh, like in a natural language you can give it a prompt and then your kind of actions get translated into actions using Playwright on a, on a real browser. And then that way you can easily automate any task, be it for exploratory testing or you want to kind of uh, automate uh your kind of a testing workflow. Uh, so let me. So, so, so what I did. So initially that was the initial uh kind of a vision and at that time one back the AI models were not so good. So I built a very very kind of um, immature kind of uh a UI which was built using Python but now I decided to use a Electron application which is a desktop application. Anybody can install it and I'll quickly show a few things. So here uh it is a dark mode, light mode as you can see Then we have a settings uh under the settings uh so under the hood it uses the GitHub copilot SDK. So which means uh I have not written any of the agentic kind of a harness. It relies on the Copilot intelligence, it takes care of everything. Now like Copilot smart enough Microsoft is behind it and they have uh memory management, they do compaction of your context they have lot many tools so I need not to reinvent everything. It uses the Copilot which is an open source kind of SDK. Then you can use your credits which you get from a copilot uh which you can use it. Otherwise you can bring your own kind of API key. So right now I'm using a bring your own API key. I have given Moonshot, it's one of the Chinese model, quite 10 times cheaper than Claude but kind of uh at capability level they are kind of uh matching with the capability of Claude so I chose this particular Moonshot but you can provide any provider which is compatible with OpenAI kind of API standards. Uh now let's come to the chat. So I'll run one of the example let's say let's do a simple form right? So navigate to this application and try to fill this form right I'm using a very simple application but you can do many uh like much more complex task here I'll say run test and what it will do now under the hood it has multiple agents so it will see that this query is related to a web automation it will use a playwright kind of uh agent uh you'll quickly see that one uh of the agent uh will come into action. Let me show the entire logs. Uh so now you can see uh it's using run command it running some tools on the right hand side it will open a browser view. So uh one thing I have taken care is sometimes like for these kind of tools the browser browser pops up over your tool right it kind of distracts you right So I kind of are running in a headless mode and streaming the kind of a screencast of the browser in this UI so that uh it's contained it uh, will not disturb you still you are seeing the live broadcast of the browser, what's happening in real time. But if you see into the processes, you will see a headless chrome is running on which the automation is happening. Uh, it will try to fill all these forms. Then there is a activity kind of a panel here. You can see every command. What exactly happened? What was my base prompt I have used? What are the different agents commands we are running under the hood. So this is for the troubleshooting, uh, just to know what exactly happened here. And we can copy it and we can give it to our AI to analyze what exactly happened. Uh, let me close this. So right now it's trying to fill uh this form. Let's, let's wait. And I have like so sometime if you think it's too noisy because trying to log every single run command you can go to key key steps. So it will just log the basic things. So now it says so so here like so what it has done, it has taken your prompt, it has tried to understand the intent, what you wanted to do and it tried to come up with like what will be the kind of a uh verdict for this scenario. Right? What would be the acceptance criteria? So if whatever you're trying to do, if it happens, it's kind of a test pass. Otherwise it will treat like yeah, it couldn't do whatever you intended to do. Right. Using a natural kind of a prompt. So in the end it says whatever I wanted to do, it has failed everything. It gave me a verdict here the confirmation and on the right hand side where, where you were seeing the live execution now it turned into a recording. So which means I can go and play it. So it will just show what exactly happened. Uh, you can pause it, restart it, download it later. Uh this is one part. The another important thing is it will show all the like tokens what how many kind of uh API time it is consumed, how much output, which model we use and okay, just let me remove the last part is so whatever the raw prompt you give from it, it drives some webflows. It tried to figure out what like navigation steps it has to do. In the end if you click on this generate refined test steps what it will do, it will give you a very fancy gherkin scenario out of this. Uh, so that uh, so you'll go away with the recording and probably a more kind of a concrete gherkin deterministic steps rather than your raw prompt which you have given uh, uh, initially. Right. So that's the whole point so what it will do, it will see the logs from the conversation and see what exactly it has done in real time. So it will not bias or rooted into your initial prompt, but it will be more grounded in the actual things happen in the real browser right out of that prompt. So let me. So if you see now we got a very fancy Gherkin scenario uh, which kind of expanded our prompt with the real data and some kind of a comments shows the kind of a locator so that next time if you give this particular prompt to any AI agent it will be even more deterministic and will consume lesser tokens. So this is one simple ah, kind of a feature. I'll show one more feature which I really like is about the uh. Like if you wanted to show AI like a particular webflow and then AI has to learn about it. What we can do is I will say okay, start observing now. So it's like a record play right? I think we both are from a QTP or Windrunner era. We know like how record play used to work so. But I tried to apply the same thing used with the, with some flavor of AI. So when I clicked on start recording it open a new kind of empty tab on a Chrome. I'll just do any, any kind of a navigation let's say google.com let's try to search about playwright and the first link comes. Let me click on it. Now we'll go back and say stop recording. So what was happening is it has observed all the events happening in a browser. It uses CDP protocol to see what all steps, what all events being fired. And now I can give some intent what I wanted to do with this recording data. Huh. I'll quickly say generate couple of test cases out of this workflow. Some positive, some negative. Sorry, I used whisper uh, flow just to convert my text to speech, uh, to be more efficient. So which means I gave a certain direction to steer it so that it knows what it has to do with the recorded data. And now what it will do, it will analyze the entire workflow. It could be pretty long. I showed uh, in the interest of time I'm showing you a smaller kind of a simpler workflow but it will go through all the events where we clicked what values we selected and then it will give me a couple of test scenarios. I said give me a couple of test scenarios out of this workflow. So it will give me uh, this particular test scenario file where you can see user searches for a playwright on a Google navigates to the official website. This is A positive handles intermediate navigation. So try to give me different kind of a scenarios out of this. And depending on that prompt, I said give me a couple of. I could be very verbose, give me all positive, negative edge whatnot. But whole idea is let AI observe, uh, what you are doing while you are manually excluding a website, and then let AI learn from it and write many more kind of, uh, derivatives kind of test cases, uh, out of that workflow. That was the whole intention.
Speaker A: Wow. All right, so you completely vibe coded this, is that correct?
Speaker B: Uh, I do not like the term vibe coding.
Speaker A: Okay, good, good, good. Let me know why and what you would call it.
Speaker B: So let me give you one analogy which I, like, love to give to other people. So let's say, let's say you, uh, like someone makes you sit in a, in a airplane cockpit, right? And you never studied physics. You don't know how aerodynamics worked, how gravity works. Uh, you are completely kind of coming from a, kind of, um, like a very old era. And now someone gave you a chatbot, okay, Use this chatbot. It's quite smart. And why don't you use the chatbot to drive this aeroplane? Right? So you will try different things. You will say, okay, now the altitude is going low. Can you, like, increase the altitude? I want to land. But the jargons which you'll speak will not talk. Uh, you will not use the same jargon which a typical pilot will use, or someone who is well versed with aerodynamics, or, uh, the underlying physics which is used to make the aeroplane fly. Uh, so to me, if you are not at all aware of anything, that's wipe coding because now you are completely judging your actions by the sheer outcome. So you try different things to AI and then you see whether the altitude is going high, then you'll say, okay, whatever you said, it worked. Rather than directly steering your, uh, AI agent to do the things which you really wanted to, rather than just giving the outcome that you have to reach to that outcome. Uh, so if you are technical, you know, the underlying layer, how these things work, that gives you extra advantage compared to a typical white coder who is just giving some instructions and measuring by, uh, just the outcome and has no say in the design, how he wants to layer the solution. So that's why I don't want, uh, to call it as a white coding. It's more like agentic engineering.
Speaker A: All right, So a lot of testers seeing this at, uh, least old school people. Like, we've been in testing for a while when, um, we run our qtp, they're so fixated almost on the coding and the page objects and all the glue that really had nothing to do with testing. And I think what you just mentioned is more like you need to know the language of your domain in order to get the most out of this. So what are you telling a tester nowadays or an automation engineer when they say oh my job's going to be replaced? I know you said it's going to take taste and you need your ideas but like what's anything concrete to get them out of this mindset that it was never about page objects in the first place?
Speaker B: Um, I think uh, the people who are too much attached to the technology or the way of working rather than the outcome or the impact their work is bringing into the, into the, into the domain or into the craft. Uh, so I'll give you like two examples. Let's say there was a one tester who was completely driven by the upliftment of the quality of the product by their work. So they are not worried about too much about like how I'll automate what will be my automation percentage. So I usually to my teams I give this kind of example that usually uh, uh, like the test planning and whatever the expertise testing we do is the brain of the testing and automation is just the robotic arms of the testing. Right. Uh, so whatever, the automation will just do whatever you planned for, right what, what, what you intended to do. But if your planning is slightly weaker, you never analyze the entire application in detail. You haven't taken care of your edge conditions, you never gone deeper into like interdependencies between the modules of application. Then automation will be as good as your kind of a planning which you did. So the people who are more kind of a uh, like identify themselves as a uh, as a kind of a uh, someone who kind of uh, brings direct impact into the quality of the product rather than how it has been done. Because I've seen people who are too much attached to the technology, they get lost into the pledge object model or it's a component screenplay model and how I'll write parallelize because most of these things will be now offloaded to the AI. They'll get commoditized. The only skill will be left is the taste and how curious you are, how deep you understand the domain, how you can correlate with the, with the testing. Like, which means how, how quickly you can find bugs, what interfaces you want to test, how you want to layer your testing like the entire sdlc, shifting left Shifting. Right. Uh, I think those people will shine more rather than the people who are very kind of a myopic and they were just focusing on, on the technical aspect because I know technique, technology is something which kind of drives a lot of people. It used to drive me, but, but now I realize like technology is something which can be easily commoditized. The, the. So the most important thing is what is your identity, how you identify yourself, what impact you want to bring.
Speaker A: So also I know, uh, you seem you're up on the latest and greatest. I've been hearing more about looping, uh, with the genetic AI and people talking about how now rather than being a human in the loop, they're going to be a human on the loop when it tests. Because you could just hand it off to agents. Do you see us going that way? Like what is, what is a loop first? Uh, maybe you could break it down. Does it impact testers at all?
Speaker B: Sure. So I'll uh, so I think, uh, let me. Okay, okay. So let me uh. I created few slides so I'll quickly jump to some of the slides just to explain the loop engineering. So I think so now I'm a big fan of Odi Osmani. I think he's a great guy. He introduced this loop engineering. Then everybody started talking about it, uh, including Andre Karpathy and like Peter Stevenberg. But if like we all come from like the era of UNIX where we have worked with cron jobs, right? So everybody has written a cron job using Crontab. And in the early days we used to verify what will be the frequency of a cron job. Right. And we used to use a website called Crontab. But I think in my opinion this loop engineering is a glorified cron jobs. So uh, the whole idea is. So right now we are so, so it's more like a one on one kind of a relationship between a human and the AI agent where we start, open a session and start kind of writing a prompt. Whether it's about building a feature testing application and then the entire efficiency of this particular relationship is dependent on the human availability, like how fast you can go and review the outcome of the AI and give the next instruction. So whole idea is to how to kind of overcome this kind of a bottleneck of the human time or availability. The uh, loop engineering, nothing is, it's like, okay, you are running these agents. Maybe it's multiple agents which are talking to each other and each agent is going into a loop. So I'll give you one example what I Do right now is let's say I'm working on prompt right? Uh, let's say I need to add a new feature. What I'll do, I'll use the AI agent. I'll say this is hosted on GitHub, we can use GitHub issues. I'll say OK, I want to add this feature. Let's brainstorm, let's draft the complete spec of this feature and put it as a uh, enhancement issue on GitHub against my repository. As soon as I post this, there is another agent which is running in a loop. What it does it, it is running every five minutes. It goes through all my repositories on GitHub which are private or public. As soon as it sees a new issue appears against the repository, it starts a fresh context. It start implementing it feature or whatever it is. It could be a bug. So it's a classic example of loop engineering. Uh, so if I need to add a feature I'll just write an issue into that particular repository. I need not to go and separately prompt and now go and implement it. There is a loop engineering come into play which will say okay, every five minutes I'll scan all your repository. If a new issue comes, whether it is a bug, new enhancement request, I will start working on it and then I'll put a comment on it once it is done for your review. So it will not merge it. But then only human bottleneck will be to review all those implementations by me rather than prompting agent again and again to implement it. I hope that gives you some idea.
Speaker A: Yeah, absolutely. So everyone says prompting was a skill that we needed. It sounds like that's no longer the case. Like being doing a good prompt, creating a good prompt. Am I wrong?
Speaker B: Um, yeah, I think so. Uh, like two years back everybody's doing courses on prompt engineering like and everybody's talking okay, act like a, like a financial kind of a consultant, act like ah, a Q engineer. So everybody trying to like kind of give that Persona to your agent and then they realize it's performing pretty well. Um, but so prompting is still I would say more or less agents like all these models have become smart enough they figured out a lot of things. You need not to be explicit about a lot of things. But still I believe the people who are more verbose, uh, right. So if you like give a requirement in two lines rather than you try to be more verbose and expand it to a paragraph though few, few places you are repeating yourself but at least you are not taking a risk to be more concise so that whatever is missing in your prompt, AI will take a liberty to assume things. Right. The only problem with a shorter prompt is like because whatever the kind of open ended things you kept in your prompt where you haven't given anything about any specific like any specific kind of requirements, AI will take a liberty to assume things and will take its own choices and start building the way it want it to be. So few tricks I will say is the plan mode which you already discussed. So another thing I always ask my AI agent as the last kind of instruction, ask me as many questions to clarify the requirements. There are few skills. I think there's a guy called Matt, uh, maybe I'll share the link. Uh so he has a very popular skills, he has around 120,000 stars on his skills. So he has a creative skill called Vill me. So if you install that skill so you give a requirement, let's say let's build a CRM together and slash grill me, then that grill me is nothing. It has a very detailed prompt inside it. But it will start interviewing you on all aspects of the CRM application like how many customers, how many users you're expecting, which tech stack. So I mean so either you have to be more verbose or you have to ask let AI to ask you as many questions so that before AI start and like start building something it has at least some decent clarity on your prompt. So these are the key kind of I would say nuances everybody should know around prompt engineering because if it is just about 2 line prompt then anybody can give. Right. Then what's, what's the differentiator you'll bring?
Speaker A: Absolutely. You also have a lot of uh, you have a few AI agents you created and one I think uh, is one that replaced you almost. So can we talk a little bit more about AI agents? Especially when it comes to testing. You said this is calling AI agents. Do you think they should be almost um scoped so you have like a performance AI agent or a uh, email AI agent. So you have these agents but they all have one job.
Speaker B: Uh yeah. Let me quickly show you how exactly I'm running my life. So um, I think the best part of uh this agentic engineering is it's like a upward productivity kind of a loop. Right. The more you uh, automate your life the more productive you become and it like ah a recursive like a self improvement kind of or productivity like upward spiral where you will see that you are becoming more productive day by day. So I'll quickly show uh, so, so One thing I did is in the January when, when like everybody was crazy about OpenCloud, I quickly went to the Apple store and secured a uh, Mac Mini. I know it's very difficult to secure it now with the, with the memory kind of uh, prices going higher now it has become Maybe I think $300 more expensive because of thanks to Micron and Apple. But I secured it and I started. Okay. I decided I need to be at the cutting edge because this is something of my interest. So what I do is I have a Mac Mini through which I'm like attending uh, this kind of a conversation. It runs 247 and remotely I control this Mac Mini through my telegram, through my phone. I use crowd code Android app and it runs my entire life. So everything is running on this Mac Mini which uh, is roughly $700 kind of um, when I purchase it. Uh, so what it is doing is there are three key elements to it. One is the cloud code. So I have a $200 promax plan of a cloud code which I try to maximize it in terms of value I can extract from this plan. And that's why I have automated a lot of things. Then I have an open claw for low hanging tasks where I don't need Claude, I use open claw which including my whatever the loop thing we mentioned. Right. I have a periodic jobs which run certain things. So I use the open claw kind of a clone uh cron job harness. Then I have a memory layer. I think this is third piece which people underestimate because AI models are kind of a stateless. They never remember what you told them last time. So you have to build the memory layer uh, around your work so that it understands you whenever, whatever work you are doing. It should understand your taste, it should understand what, how you want to approach things. So kind of it's your second brain, right? Whatever you would do in any situation they would know from the knowledge base that what is the best choice to be made here so that it should not come to you every time. Most of the things are uh, stored in the.
Speaker A: Is that, is that a skill or is that like an obsidian type of repository?
Speaker B: Uh, it's a, it's, it's a collection of markdown files you can think of hosted on my private GitHub repository. But I use a library called QMD. It's a quite popular library built by the uh, Shopify CEO uh Toby uh, and it is nothing. It's a, it's a kind of a rag, kind of uh, a miniature rag pipeline which you need not to worry about how rag works one library, it will take care of every everything locally. But you can do semantic search against your knowledge base. So let's say you have like 50 markdown files about your life or about whatever work you do rather than AI agent. Every time scanning everything, it will do a semantic search against your knowledge base to figure it out. What's the answer? Okay, so then I have a three agents. I'll quickly show you. So I have a like a chief of staff. My chief of staff runs my LinkedIn, my YouTube channel, my lot of things like it uh, it can read my emails, x.com whatnot. I have created recently chief health Officer which means I, I'm, I'm not subscribing to any of the paid app for my fitness. So I have a telegram bot. I just log what I'm eating. I take a picture and log it and then AI agent kicks in and keeps track of everything into a database and give me a briefing next day. How was my last day? Uh, and try to nudge me or whatever like if I need to work out more or I'm not eating too much junk. But I'll show you a quick uh, example then. Recently I did a chief financial officer which is more for my like income tax and other like my mutual funds and all. Uh, but they all are like very uh, running in a sandbox environment. They cannot talk to each other unless I want them to talk. So okay, so few things as I said. So one challenge I faced it because the pace with which AI is moving is really difficult to uh, kind of catch up because every day we are seeing something. So I decided, okay, I should automate this part. So every night my morning AI briefing agent goes and kind of scans the uh, Reddit, X.com, linkedIn, YouTube, whatnot and try to curate some 10, 15 biggest things happened in the last 24 hours. Every morning I get a uh, briefing I'll quickly show. So this is my uh, yesterday's. If you can see this is on the 1st of July. This is the AI briefing I received on my telegram and which says anthropic ship plots on it. 5 Loop engineering is the phrase of the week. Boris. So then I have a detailed markdown file. This is the highlighter. And then every morning I try to get up and go through it. Uh, this is one bit. The another bit is uh, while scanning this, if it detects something which it's worth doing a deep dive, right? So it finds something which I have given a curated kind of a uh, kind of a prompt to the AI agent that what excites me and what new things I wanted to learn. So if it finds some topic which I should learn, then it adds to my learning lab which is a separate project and it says okay, new learning topic added to your backlog. So that I have to kind of clear this backlog that these are the curated items my agent or my chief of staff wants me to learn so I can go and then go deep dive. So loop engineering came around three weeks back, surfaced by my agent to me and then I did uh, a deep dive. Uh, like what is loop engine? Same way I have to look into this role confusion kind of a prompt injection. And third thing, as I said, so my cho knows what I have done last day. So every morning I get a message like this which says okay, yesterday was a strong day. Everything about, about my food, my workout, my medicine. So it gave me a kind of a sense that how was my study in terms of my fitness and all. Um, so I tried to use Telegram as a control plane which means Telegram has uh, these Personas, separate bots, CFO Cho, chief of uh, staff and all and through which I can also talk to them. And also there are some loop engineering which means in the background they keep scanning my things and do work for me. Um, one thing I wanted to highlight is uh, recently like my dad has gone through my like a surgery. He was in India and he was a bit reluctant to use these kind of AI models to like know more about his medicines or about his condition. Then I created a telegram bot which is written which, which speaks his language, which is his native mother, um tongue Hindi. And, and and I gave a telegram bot to him installer. I said okay, now you ask in your language. You can do voice notes and ask anything which medicine or anything you have any confusion, it will answer you. And I will monitor the locks, uh, and I'll keep improving the experience. So. So I mean it's not about your work. I mean this agentic engineering can improve like uh, take care of lot of friends, uh, like for your life. It's not just your work thing. The uh, another thing I just wanted to talk briefly is I'm like an avid cyclist. So I cycle usually every day, 10 km or every other day. So I use while cycling. I try to do a lot of agentic engineering and I'll quickly show how I do it. So I have a kind of a telegram bot which is Saathi where I, what I do, let's say I need to start build uh, some New product today, right? In 30 minutes time. When I'm cycling, while I like start cycling, I go to this chatbot, I say start a new cloud code session and name it whatever, like my weekly newsletter, uh, or Topmate Priority or whatever. It could be anything. What it will do, it will go to my Mac Mini and start a new cloud code session. It will give a name to it, it will activate the remote control to it so that my Android app can see it. So as soon as this happens I will see a new kind of entry pops up into my clock code app which will have the same name which I give and then I open that chat through my phone while cycling and I use a whisper flow which is a kind of uh, a uh, speech to text, uh, kind of a service. And then I start brainstorming while I'm cycling, I'm talking to my agent and we are brainstorming in 30 minutes. Usually I kind of ship maybe smaller features so I just wanted to highlight that. Yeah, so this is the way um, the future work would look like. And I really got inspired by uh, like uh, Boris who is the creator of Cloth Code, who said like 70% of the work he churned out in Anthropic is through his phone. So that's why I thought I experimented with it and I realized if you have a good kind of um, a uh, speech to text kind of a service, you can talk to your kind of app then yeah, you can drive lot of things remotely on a, on a remote machine. It could be a Mac Mini or a, or a server in cloud. So I hope that gives you some idea on agents.
Speaker A: And this is very cool. I don't know how you get it to be reliable. I've been trying openclaw and it always has issues. I know a lot of people are going to Hermes. Is it more of people, uh, don't understand openclaw and that's why they go into Hermes or they think Hermes is just better.
Speaker B: I think there are many. They're Nano Claw and they are like by Nvidia also. Uh, so I think more or less all these products work on the same principle. It's more like okay, there is an agent which is running remotely and there are some channels. Openclose is kind of for more technical people. It has more layers. Uh, but yeah, I think Hummus is good if you are not very technical or you don't want all the components, the gateways and daemon and all. It's a uh, kind of a miniature version, a more leaner. But yeah, anything it's, it's. It's.
Speaker A: It's.
Speaker B: The best thing is to get started somewhere. Right. To start with if, if nothing maybe start your own agent. Simple uh using LangChain or, or any other library. Uh, but, but yeah uh, the, the most important thing is to get started.
Speaker A: Now I know you said you're using Claude. Not Claude. Was it Claude? Max? Whatever Your, your, your your uh.
Speaker B: It was uh. So I use. I have a $200 cloud code subscription and which gives me uh, which is unlimited. So best part is uh, there is no kind of a pay as you go kind of a model. They have some rate limiting. Every five hours I am only allowed to consume let's say a few million tokens. If I consume in the first hour then they will block me for next four hours till the fifth hour hit and I'll uh, reset my kind of a token. But I rarely hit that five hour limit. Uh I try to distribute my work across as I said with the help of Loop engineering. The night time lot of work gets done when uh. I know I'm not driving any of the, any of the agent. Uh but yeah the whole. So whole idea is I'm spending $200. I want to extract maximum value out of my $200. That's put me more pressure on me that how, how I can like do more work in agentic fashion so that I can uh like do more out of that particular limit which I have a ah, part of the $200 plan.
Speaker A: Right. I know a lot of companies are like Uber blew their whole uh budget in four months. Do you see that as an ongoing issue token cost and maybe that's going to limit us for being as awesome as you are with the gentic uh, like for you know we're going to hit a ceiling at some point because we can't go as fast as we want because it's going to cost too much money
Speaker B: usually when I talk to my team. So I think last year we talked about the capability of AI right. So people were burning tokens left, right, center, uh trying to build some toy applications, chatbots and showing see this is the capability. It can generate videos and we have seen OpenAI launch Sora just to generate like some social media kind of uh, short videos. But then they had to discontinue it. So I mean now I have start change the conversation with my teams from like show me, show me the demo or like show me your work to show me the value. Right. So eventually everything will boil down to like you burned. Okay 10 million tokens. What's the value you achieved? Don't show me any demo applications, just tell me what's the logged value, right? What is the realized value? Have you automated your manual test pack into automation with that by burning these tokens? And now even if you plug out the AI, your automation can be run. So I think provided you justify the value. So in my case I have written like. So I'm not sure I showed my personal website, but I can quickly show. So I've written a small uh, blog. Uh, okay, let me skip this intro. So if you go to my writing. So maybe people, if people are not sure whether they should invest 200 per month, I have written a detailed blog. Uh, like what, what kind of a decision making process I've gone through, how I'm trying to get the ROI out of my $200. Not for a short time frame, but maybe every year I'll revisit and try to uh, like do a math around it, how much total annually I spend, what value addition I got. So I mean few things. I always wanted to start a YouTube channel, which I couldn't. So now uh, I have a YouTube channel. So I mean uh, few are kind of a tangible things. Few are maybe your passion things which you wanted to do because of your time limitation you were not able to do. But now I have a channel where I post videos. So I mean it depends how you want to extract that value. It could be running your business like building some side projects, running a YouTube channel, running your post on LinkedIn or whatnot. It could be anything.
Speaker A: Awesome. Okay, before we go, is there one piece of actual advice you can give to someone, especially a tester, to make them move towards this agentic era we're moving into and uh, the best uh, place to find or contact you.
Speaker B: Okay, so I would say I think I've written a LinkedIn post last week, uh, and which got a decent amount of traction. So I think this agentic engineering is more about if you try to declutter the noise around it, it's more about automating your job. Right, how to automate yourself so that you can move up the value chain and do starting high impactful work. And if you are coming from a qa, especially QA plus some automation experience, you have to apply the same kind of a principles here. You need to understand what's the problem statement, what's the workflow you are planning to automate. But rather than writing a deterministic script in Python or in Selenium, here we are dealing with more of a agent take or probabilistic kind of a thing where you need to articulate what needs to be automated, how it needs to be automated and then it will get automated. It is highly resilient. You need to worry about, you need not to worry if locator changes or uh, API spec changes. It is smart enough to improvise and figure it out that your automation still works, though it will consume slightly more tokens. So if you are a QA with a good kind of a curiosity, good understanding of a QA domain and with some decent technical skills, you are very well positioned to excel in this agentic engineering era. Unlike to other professions like developers or a DevOps because you understood the larger picture of the product, you always worked uh, on a larger feature set, larger independencies of the modules, rather than a developer who works on a very narrow kind of JIRA story. So I would say just try to focus on these couple of things, some technical aspects and don't kill your curiosity. That's the key thing here. That's the fuel to drive your agentic. The more things you try, more things you will build, the more intuitions you will build. And this is something you cannot learn by seeing others doing it. You have to make your hands dirty and try it, struggle with it and then you'll start excelling and building your intuitions around it. And if you want some help, uh, I have a topmate, uh, uh, I run it. If you want uh, to book some time with me you can book. Otherwise I have a website which is amitrawad.dev uh you can reach me through this website or my LinkedIn or you can find all the, all my social coordinates on this website. Uh, my LinkedIn X. Uh and yeah I have a YouTube channel where I'll try to talk about agent tech engineering and try to publish one video every week. Uh, yes, that's pretty much awesome.
Speaker A: You can find a check out all the links for me down below. Thanks again for your automation awesomeness. The links, everything of value we covered in this episode head on over to test guild.com a595 and if the show has helped you in any way, why not rate it and review it in itunes? Reviews really help in the rankings of the show and I read each and every one of them. So that's it for this episode of the Test Guild Automation podcast. I'm um, Joe and my mission is to help you succeed with creating end to end full stack automation awesomeness. As always, test everything and keep the good Cheers. Hey, thank you for tuning in. It's incredible to connect with close to 400,000 followers across all our platforms and over 40,000 email subscribers who are at the forefront of automation, testing and DevOps. If you haven't yet, join our vibrant community@, uh, testguild.com where you become part of our elite circle driving innovation in software testing and automation. And if you're a tool provider or have a service looking to empower our guild with solutions that elevate skills and tackle real world challenges, we're excited to collaborate. Visit test guild.info to explore how we can create transformative experiences together. Let's push the boundaries of what we can achieve. With Lutes and Liars the bards began their song A tune of knowledge, A melody, uh, of code through the air it spread like wildfire through the land Guiding tester showing the secrets to behold.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.