
Future of UX · 2026-07-02 · 51 min
Key moments - from our scoring
Substance score
43 / 100
Five dimensions, 20 points each
Caitlin Sullivan, an ex-UX designer with marketing and research experience, shares how to leverage Claude Code for AI-powered customer research. Rather than viewing AI as a replacement for research rigor, Sullivan emphasizes starting with documented expertise - defining decisions already made, scope boundaries, segment definitions, and hypotheses in a Claude markdown file before running any analysis. She demonstrates a practical folder structure for research projects, including separate markdown files for context, segment definitions, and codebooks, which feed into Claude workflows selectively depending on the task. The episode walks through her approach to avoiding hallucination and misalignment by treating Claude like a team member who needs clear briefs, then validating outputs through spot-checking or structured evals. Sullivan's framework suits designers, product managers, and researchers who want to reduce manual synthesis work without sacrificing decision-making quality or introducing AI-generated assumptions into their findings.
Document your team's existing knowledge: what decisions have already been made, what's in and out of scope, segment definitions, a set of beliefs (separating known facts from hypotheses), and what good research outcomes would look like for the specific decision you're trying to make.
She creates a project folder with a Claude.md file loaded on every chat (containing research questions, decisions, and segment definitions), plus separate context files (deeper segment details, codebook for survey coding, research question history). Skills saved at the user level point to specific markdown files only when needed, so Claude doesn't load irrelevant information.
Sullivan recommends a middle path: implement either a spot-check workflow (validating a sample of outputs against raw data) or structured evals (designing and iterating the Claude workflow itself until it gives consistent outputs, rather than manually verifying each result).
She started experimenting with ChatGPT at the end of 2022 when it went mainstream, driven by concern about being left behind. She transitioned to focusing entirely on AI customer research after returning from maternity leave, when client requests shifted to the AI research methods she was posting about on LinkedIn.
Documenting prior decisions and beliefs prevents Claude from generating a new set of assumptions that might contradict your starting point, which can confuse teams about what to prioritize and lead them to defer to whatever the AI suggests rather than grounding insights in existing knowledge.
Our reviewer’s read on each dimension, with quotes from the episode.
There are genuinely useful tactical ideas - structured evals as rubrics, layered Claude MD file architecture, and the 'agents are just a series of files' reframe - but they are spread thin across heavy conversational padding and host affirmations. The core insights are surfaced but not developed to depth a senior operator would need to act on.
I think in a funny way I'm cutting fewer corners because I'm actually able to get AI to fill in all that. I would have deprioritized because I would have said, oh, I need to deliver this by tomorrow, I don't have time.
evals are just a set of measurements. It's like a rubric.
One genuinely fresh reframe - 'start thinking of agents really as just a series of files' - and the idea of running structured evals before live insights work are mildly counterintuitive; everything else (document your knowledge first, don't let AI decide for you, check outputs) is standard AI-workflow orthodoxy recycled in UX language.
start thinking of agents really as just a series of files
If I feel like I haven't thought hard during the week, then I realize I have to cut out an AI workflow
Caitlin is a real practitioner who has held a Head of UX Research role and built genuine Claude Code workflows, giving her credibility over pure thought-leaders; however, her experience is primarily solo/freelance at small-scale startups, she pivoted to this specialty partly by necessity, and she has no notable track record at significant organisational scale.
I was a head of user research in a SaaS company before I kind of went freelance
the thing I kept getting requests for was around the stuff I was posting on LinkedIn, which was all around how I was experimenting with AI for different research use cases
The WHOOP folder walkthrough is the episode's only concrete artefact, but the guest explicitly states all the data is fake and synthetic; there are no real client names, no measured time savings, no accuracy percentages from evals, and no before/after metrics - just named tools and described file structures.
this is a fake folder. To be clear, I don't work with woop, I never have. So this is entirely fake data
run the same process again and then check if all of the numbers, all of the counts for all of the multiple choice questions, like the measure is the counts from the multiple choice questions and whether they are all the same
The host asks useful clarifying follow-ups (how MD files reference each other, where skills are saved, what a pull request is) that help listeners unfamiliar with the tooling, but there is no substantive pushback, no challenge to any claim, and a high volume of unearned affirmation that prevents the guest from being pushed to greater depth or precision.
I love that. I love that because last year I also had this agent rabbit hole
How do you Reference the different MD files to each other
Computed from the transcript - who did the talking, and the words that came up most.
Most designers think Claude Code is a developer tool. Caitlin Sullivan makes the case that it's one of the most underrated tools for customer research and insight work. In this episode we talk about how she actually uses AI to get from raw customer conversations to insights you can act on, and why designers shouldn't wait for permission to pick this up. If you've felt stuck in the chat window, this one's for you. Key Learnings Why Claude Code is useful for designers and researchers, not just engineers How Caitlin turns messy interview and customer data into clear patterns Where AI actually adds value in research, and where human judgment still has to lead A simple way to start if you've never touched anything beyond a chat interface Find Caitlin Sullivan : Linkedin: Newsletter: Course: AI for Designers : 5-week Bootcamp opening September →
Transcribed and scored by The B2B Podcast Index.
Speaker A: Hey, Kaitlyn, and welcome to the future of ux.
Speaker B: Hello. Thanks for having me.
Speaker A: I'm so happy that you're here. Super excited to dive into the topic of AI customer research cloud, how to use AI. Before we are diving into everything, I would say do a quick intro and tell the listeners who you are and what you're doing.
Speaker B: Sure. So I have, uh, a. I have a mixed background. I usually start with that because I think it's important. I'm an ex UX designer, but before that I was also in marketing and digital strategy and then sort of ended up being more focused on research. But I've done research across all those roles and think it's probably a good thing that I can take all of those perspectives while doing research. For the last couple of years, I've been pretty much entirely focused on AI customer research or using AI in the best possible ways and getting better results when doing insights works.
Speaker A: Cool. I really love this diverse background and I feel that is always such a plus if you're working in design, with customers, with people, to have that empathy and to really zoom out and also working with AI, have this broad understanding. This is so cool. Um, how did you started to use AI in the first place? I'm curious.
Speaker B: I got really curious all the way back in 2022, the end of the year when ChatGPT first went kind of mainstream live for all of us to use. And I think like most people, I had sort of, um, a certain reasonable amount of apprehensiveness. Uh, but, you know, I very quickly kind of realized that this was one of those moments in time that is, or I would get left behind if I didn't start to grab the bull by the horns and try to understand it as well as I could from the start. So I did that pretty quick, quickly. I was still critical and careful about things and didn't immediately trust everything I was seeing and getting out of the results. But I, I knew that I needed to try to figure it out, understand how large language models work, where their kind of benefits and limitations are, what works consistently, what doesn't, and why, in order to kind of be able to continue doing my craft and working the way that I like to work in the future. And it was out of my own curiosity and feeling that I had to keep up. But the fact that I ended up now doing what I'm doing, focused entirely on it was not really by choice because as we were talking about before we started recording, I also had a babe, my first child, and I went on maternity leave for a shorter amount of time than a lot of people in Europe. But I came back and didn't have any client requests. So my previous way of working kind of retainer based with individual product or research teams was no longer an option. And the thing I kept getting requests for was around the stuff I was posting on LinkedIn, which was all around how I was experimenting with AI for different research use cases.
Speaker A: I think that makes so much, so much sense. Align and adjust a little bit based on also your circumstances. And then with AI, we also talked about it prior, the recording opens up a lot of new opportunities, especially for people who don't have the same amount of time and want to be a little bit more productive. What would you say are the areas, especially in AI customer research, where AI really makes sense, where AI really makes a huge difference compared to maybe doing things manually without AI?
Speaker B: I think the time aspect is kind of the easy one to point out, the one that's almost as much, maybe just as much of an impact for the ability to a more complete job than I think I was able to do before I was using AI because I've worked for a long time as kind of a solo operator. As a, I was a head of user research in a SaaS company before I kind of went freelance. Even there I was very often kind of like working on a project on my own because we had a limited number of people working in the team doing insights and uh, you know, and then as a consultant I would be brought in to lead a project, support a team with training, but also take on some new kind of figuring out who the actual ICP was and projects like that. Working on your own has so many limitations. You still have the same sort of tight deadlines, especially in startups and scale ups and you have to do so much in so little time that most of the time isn't really possible. And so now with AI, I think in a funny way I'm cutting fewer corners because I'm actually able to get AI to fill in all that. I would have deprioritized because I would have said, oh, I need to deliver this by tomorrow, I don't have time. And now I can say AI can pick up the slack. I can actually be more complete in what I'm delivering because I don't have to base this only on what I can do.
Speaker A: I think, um, super cool examples, right? To see the opportunities that AI brings and makes your work even better. What would you say or how does it look in practice? So when you're working on um, a research project, um, you do some customer research from the start to finish. What would you say? Um, where do you sprinkle AI in and how. Just get a bit on, like a bit hands on. I'm curious.
Speaker B: Have kind of a general principle. I think every project is going to be a little bit different because, you know, helping a team, understanding where onboarding is going wrong. It's very different from um, exploratory research. Like I've done a bunch of research with different teams testing whether they should launch in a new market, which is a completely different animal. Right. Than design decisions on a smaller kind of in a part of the product. So I think where you use AI is very different. Don. The amount of risk, the amount of uncertainty, the amount of existing knowledge that you have about the part of the product or the users that you're going after or the customer. But generally I say with most things that I do with AI for insights, I try to start from my own expertise or my own understanding or the team's understanding of the problem space or the customer. Get that down if it isn't already. Like document that, which I think ties into the rest of what we're going to talk about today, but document that in the first place so that you have your own expertise and the pieces from your own brain and your own knowledge in the first place before you even open up a chat in any way, shape or form or start an agent. And that's really crucial for me. So regardless of how, you know, where I put AI in which steps or you know, whether I'm creating agentic feedback loops for itself or whether I'm manually checking, almost less important than make sure we start with our own knowledge and understanding of like what we know about the customer, what we know about the product behavior and then also what would make this project good. Right. Because I think so many people still struggle with identifying what good research would look like in this case.
Speaker A: Yeah.
Speaker B: Before they could just let AI start doing it for them. And that's the source of.
Speaker A: So what would you say, what would you write in this document? You already mentioned, like your customers, your goals, maybe some problems that you work currently trying to solve, um, some questions. Do you write like everything that you have in that doc or do you have, I don't know, like those top 10 topics that you always mention in this documentation?
Speaker B: I think the more important pieces are things like what decisions, like what decision are we making but also what decisions have already been made, Things that we are not going to cover in this research scope Right. Like what's in versus what's out. Very clearly defined definitions of segments or who we think our segments are. Um, like those are to me, all of the fundamental pieces, but also a set of beliefs. I sometimes start with a set of beliefs and we can say either we can turn these beliefs into concrete facts because we can say, we can point at all of this evidence that already exists. Right. So we know that that's true. You know, a high, high level of accuracy or high level of confidence. And then there's other pieces or list, you know, items in that list of beliefs that say, are probably hypotheses, but at least having them in place before we say, hey Claude, go, you know, help us make this decision. And it comes back and uh, you know, tells us a whole new set of beliefs and we're now confused about what the priority was and what we should be paying attention to and sort of follow whatever Claude says instead of what we started with.
Speaker A: Yeah, I think this is such a good tip. And what a lot of people skip or what a lot of design teams skip, they write maybe a little bit of instruction, a little bit of background information, but I assume it's most of the times not enough. You probably did a little bit more. And especially I like what you mentioned there, focusing on the decisions that have been made and also on the scope, like what's in our scope and what's out of scope. At least from my experience, you get a little bit lost into all the opportunities and you get excited about the insight and then you like move out of the scope and get lost at some point. Right. So that can happen. Okay, so the first step was that this document could be like a text document, right? Or Google Drive document.
Speaker B: So when I start that is usually like a notes document. Like I jot down all of these things or I make a team jot those things down into a document. What's really important though is if we want to work with Claude code or any other kind of file based system, all of those things won't necessarily live in the same place. So if we want to kind of start turning our knowledge and the things that we believe Claude or any other non human entity is going to need from us to do the best possible job, then we actually need to think about where all of those things belong. Are they all things that need to be given to Claude every time we like open a chat in a specific folder, regardless of whether we're doing a design workflow or a decision, you know, support process or an insights workflow or Are they things that should be fed at specific pieces times? So we have to think a little bit more. Like that's the second step. Mhm. Honestly, I don't usually have to think about most of that from scratch. Especially if I'm working with a team that has kind of already done that with me and then we start a second project, we already have most of that in place. Or if I'm doing things for myself like mining insights from my own work, my own kind of coursework and so on, I have those systems in place that kind of always apply. So you may create that sort of thing that you have all of those, you know, segment definitions, things we already know, decisions that have been made, those are things that could apply to everything you do for the next six months or years. And then there are the things like the research questions, the decision we're making right now that will be more project specific.
Speaker A: M or the scope for example. Right. This could also be. Okay, cool. Um, and how do you feed this into Claude now? So we define the areas that, you know, like the background information and some of the questions that we maybe partially, uh, use for only parts of the project.
Speaker B: Let me share my screen.
Speaker A: Yeah, sure.
Speaker B: All right. I have a folder here, so do you see my folder that says WHOOP screen? Great. So this is a fake folder. To be clear, I don't work with woop, I never have. So this is entirely fake data and information here. But I use this as an example when I demo things because I think a lot of people know what WHOOP is. And I, um, have this fake project where I'm pretending that I'm exploring the idea of having a screen on the famously screenless armband. This is an example of what a
Speaker A: project for everyone who doesn't know whoop. Could, uh, you explain real quick what it is?
Speaker B: Sure. So WHOOP is in is like a fitness tracking armband. It measures, uh, heart rate and your sleep, you know, tracks your sleep and tells you how well you're recovered, uh, if you've done a workout the previous day, whether you should work out again, and so on. So it's a bit of a recovery and uh, activity tracking. And it doesn't have a screen, it's intentionally screenless. So it's not supposed to distract you from, you know, going about your day, doing your workouts and. But I'm exploring in a lot of my synthetic data. I have sets around things like WHOOP would add a screen and whether the kind of synthetic data that I've created is saying that they should or shouldn't. And I have a lot of Easter eggs hidden in there, buried to, to trick AI a little bit sometimes and see if it actually comes up with the same answers that I would. So just an insight into how I, how I test things and how I test workflows and whatever I do with AI. And so this is one of the kind of fake project folders that I have. But this is exactly what my real project folders look like and the ones that I put together with clients and students in my courses. What I have here is within one project. This could be, you know, an onboarding redesign or whatever it is you're working on. There are folders here. This is for kind of insights work, but you could imagine that there might be like a design drafts folder. So the CLAUDE markdown file is one place where some of the information we just talked about goes. But then we also have things like rules files, which I'll talk about in a moment. And then we also have a context file. So the CLAUDE markdown file here, the Claude MD for anyone who's new to this, is where the sort of essential, uh, information about this particular project, anyone who would work on this, you know, obviously it's AI, but anyone who would work on this project needs to understand. So we have the research questions, the decision, warmth, things like segment definitions, who we're kind of looking at data from in this project, uh, anything else you think would really be essential to understand about this whole project before working on anything in here. But then we probably want to have more details like we were just talking about, have maybe more kind of deep dives into who our segments are, how we define them, what they look like, background information and data we have about. Maybe we want to detail more about the research questions, where they came from, what we already know and why we're, you know, why we have these questions and not other ones. I also have things like a code book to start with if we're coding survey responses and things like that. So there can be tons of context in here. And so you don't necessarily build all of this from scratch if you're just kind of starting your first project and starting, you know, creating all of these files. But I think you can imagine that we might create at least these ones before we do any sort of, uh, analysis or insights work with a set of data. Because having research questions and segment definitions will already give Claude so much better background information to anchor to before kind of making assumptions about who our segments are and what they look like.
Speaker A: How do you Reference the different MD files to each other. So in the main, a Claude MD M file referencing the other MD files or the other folders. Or do you leave this to Claude to find this on?
Speaker B: So the Claude MD is something that every time you do any workflow in this folder. So I have my Claude terminal chat window open over here just because I've opened it. I started up this chat within this project based the chat. If, if this is new to listeners, you can start up a chat by just kind of opening a window and writing claude. And then this is going to access everything on my computer, which I don't always want it to do. I mean it, it will always have access to everything on my computer. But if I want to do an Insights workflow, I don't really want it like hunting through all of my files on my computer if it can't find the right one. Right. Whereas if I launch in one project folder like this keeps it a bit more contained. So once I've done that, if I were to do something like Survey analysis, which is a skill, I've set up like a hard coded workflow that will run on its own, this will automatically load this context file, but it won't automatically load all of these. And that's intentional because I want to give it the basic information that it would need for anything it does, which will be loaded every time. But if I want to give just this Survey analysis workflow access to the code book, for example, I don't want the codebook loaded every time I do like a design workflow in here. I only want it to load with this skill. So this skill has instructions that points to, that this doesn't need to. This one's just the high level overview. Hey Claude. This is what this project is about. This is what we're digging into, what decision we're making and so on. But every time a, uh, workflow needs more detail, more specific information to get that particular task done, then you can kind of use all of these files on demand. Does that make sense? Yeah.
Speaker A: So you basically mention them in the terminal or you use a skill where the MD file is connected.
Speaker B: Yeah. So in the skill, it points to this. The skill file has a set of instructions that say something before you start coding. Look at Codebook, Maryland.
Speaker A: And where did you save your skills? They're not in this folder. Right. So they are more, um, you save them somewhere else.
Speaker B: Right? Yeah, they're saved on the user level. So they're under here, um, on my top level, user level. And that's Where I tend to save everything just so that I can always access all the skills.
Speaker A: Yeah, I think this is also super interesting because I assume, I mean, not sure if listeners or viewers have used skills at some point. Um, if you are just starting out, you usually create your skills. You save it, um, in your cloud app, right. Like you have your skills in your cloud app, but you can also download them and then save them local computer, ah, on your user level. And then basically use these skills if you either use the terminal, if you use VS code or cursor or whatever. If you work locally on your computer. Right. So can use the same skills, but you can't access the skills that you have in your CLAUDE app.
Speaker B: Right.
Speaker A: You need to download them so you have them locally. I think this is important. And they're basically the same skills that you can create in the cloud app, right. So an MD file with references, with agents, whatever. Uh, super cool. Thank you for seeing that. I think this is always super interesting also to see how content and insights and data can be structured. Not that you have everything in one file, but I really liked how you have this brain file, pretty much the CLAUDE MD file that is loaded every time that you're working on the project where you have all the general information. And then if you feel, uh, I need to go deeper, you just tool you're using and then it basically gets the information manually. When, uh, when you need this kind
Speaker B: of metaphor, if, if you've ever been managed or have managed right, like in either direction, you've had an experience needing to tell, like, I need you to take over this project and hopefully give them everything that they need to do a good job. Or you've had a manager who gave you the things to do a good job in. Either direction, you've had probably at least one experience in your life where that went not so well. Like you didn't get the things you needed. You went back, your manager was like, you did a bad job and then you needed to go, okay, but I didn't have everything I needed or something was missing. Uh, when we think of it that way, like that's what we're putting in these files. Like, I always try to think of what did I actually really need to give someone who was working for me, or did I get given a design brief that really wasn't as good it should have been. Cause I was missing something. I was missing some of the insights that kind of drove the decision to do this thing. So like, those are the pieces that we have to build in somewhere. We want to be able to give Claude like a support, the full package that it would need to do without coming back to us and going, oh wait, you forgot this. Or like, oh, what about this? Or do you have an answer to
Speaker A: a hallucination, like starting to come up with things that you never mentioned, uh, you haven't mentioned and then it comes up with its own ideas. Don't want that. Right. Totally makes uh, makes total sense kind of insights. What IT need all went through the situation that you described where we didn't have enough information and we didn't really know what to do. So I think that makes a lot of sense to have that basically as uh, a mental um, map. When you create this Claude MD file as the next step, which is probably step number three only we are creating insights, we are using Claude to synthesize, um, um, doing some analysis for us and we are getting insights. What is our next step with these insights that we are getting? Maybe we create this in the terminal, we have our insights there, or we save it locally somewhere. And what is the next step? Maybe also the handoff to the, to the other department. Right. You're not working in a bubble as a researcher.
Speaker B: I think the next step that most people do is checking everything that came out of AI, to be honest. And I think this is really important because you know, one question I get asked often is how do we stop kind of triple checking everything? There's a lot of teams who are, you know, very used to doing research and know how to do it well, whether they're designers or product managers or researchers are still unable to immediately look and um, go, this is right or wrong because especially if you're doing, you know, new work, you're not testing with old insights or old data files, you can't tell. Claude and all of the other large language models give you something that looks very confident and they can be filled with so many issues, but who can tell what's right and wrong? So I think everyone is still, let's say people fall into two camps, either overly trusting or really second guessing everything and going. We spend, you know, a fraction of time getting the insights out now from the data. Uh, but we kind of eat all of that time that we saved back up because we kind of triple check and compare exact details and data and the outputs. So I think a next step is actually implementing a workflow where you have a little bit of an iteration process and that doesn't have to be that you, you know, the human dig through every detail of what came out Of Claude or other large language models, you need kind of a system. So either you need to say, okay, at an entry level, we are going to um, kind of do a spot check like of a sample. So we take a sample of the outputs we got and we look at them closely and we compare with the exact data points and we make sure we can find all of the quotes that the counts of responses to that multiple choice question make sense, add up and everything or what I actually do, uh, the advanced version that I think most people will not get to immediately after listening to this podcast, but just to plant a seed is structured evals. So if I run structured evals on a workflow I've designed, so I do like a survey analysis as a workflow, I design how CLAUDE should do that. I run evals on that workflow so that the system I've designed is iterated until it gives you consistent outputs every time. They will not be perfectly exactly the same every time, but I want to see them with the same sort of insights every time so that we're not second guessing, you know, like, well, it gave us these insights this time in a completely different set of insights. The second time we ran the same process, like that's not what we want. So actually before running any insights work evaluate, I do evals on my workflows to make sure they're reliable. And I really recommend that all teams start learning how to do that. Instead of going like this prompt or this set of instructions, work today and then not sure if it's going to work tomorrow. If you test like that more systematically and when you're running the insights process for real and you get those outputs, you will uh, automatically know whether you can be confident in those results and you'll know like which details to more kind of spot check or more closely look at.
Speaker A: And how do you set these evals up Just in a super simple, not in detail, just like that, we get a sense of how we could get started.
Speaker B: Evals are just a set of measurements. It's like a rubric. It's the equivalent of saying um, you know, for like after you output all or like from a set of results, from a set of survey response, you know, analysis results. Go back in the same, you know, run the same process again and then check if all of the numbers, all of the counts for all of the multiple choice questions, like the measure is the counts from the multiple choice questions and whether they are all the same. And so then the, the rubric sucks. If they're all the same, you Get a pass. And if there's one of the three that is different than the other two, then it's a fail. Okay, that's a very simple one. But you can design evals for pretty much anything. You just need to have tangible measures for things. Whether you're doing survey analysis, interview analysis, any sort of like design workflow. You want to make sure that you get the same kind of view with the exact style guide you gave it every single time. You can create an URL for that.
Speaker A: Okay. Oh yeah, Cool. Uh, basically thinking about parameters, certain measurements, how you could measure some. Define this, write this down and then see if it matches the result. Right. I mean like, just like super simple. Of course it's, it's more complex, it gets more complex. And is this something that you as a researcher would set up or someone else in the team?
Speaker B: I set it up and it's really just. It's a document like everything else that we've talked about here. It's another document. It's just a file. It's a file that basically says like, this is what we're measuring. We're measuring the, the accuracy of the counts of these multiple choice questions and we are determining whether you come up with the same answer or the same count for all of them across three runs. And you just clarify exactly how to measure everything and exactly what counts as a pass or a fail, whether it's a percentage or all three have to have to be exactly the same in order to pass and so on.
Speaker A: And would this be an MD file as well that you share with Claude?
Speaker B: Yeah, I have a few evaluations set up that are a skill.
Speaker A: Mhm.
Speaker B: So just like I would run a skill that's called survey analysis, or you could do skill wireframe, you know, app wireframes or something like that. And it runs a series of tasks. It's a file, but it's a skills file, a skill file that I can do like forward slash, eval 1. And it just takes a certain data set or like whatever data set I've told it to use and runs one process three times, or five times, or seven times, whatever I've specified in the instructions. But it's really just that one skill file that says, hey, here's the rubric, here's the eval system you need to apply, and here's how to run the experiment. Basically like run it three times, use the same data set, uh, ask me which day.
Speaker A: Okay, cool. Okay, that makes sense. But do you always use the terminal for that or also always the terminal?
Speaker B: I Know, it's like, for so many people it's like, ah, oh, cringe. Who wants to sit in the terminal? But I mean, former designer. I also appreciate an actually designed ui. I am getting really used to it because I find that there are just some things that come really natural when you're sitting in the terminal. Like you said, everything is just lives locally. You know where your file paths are. Yeah, you can do pretty much everything the same. The same way you can do all of the same things in the cloud desktop app in the uh, code tab. But, but I have gotten really used to the terminal, to be honest.
Speaker A: Yeah, fascinating. I mean sometimes I'm working with the terminal, but to be honest, it's a bit, uh, it's, it's not so visually and I also know a lot of designers who get really scared of the terminal. So it's, it's, it's great to hear about your experience and your workflow because I hope that, you know, it removes the scarcity a little bit for designers to just like experiment and try this out because like with all the commands it can be overwhelming.
Speaker B: Like a non technical person here. I'm really not a technical person anyway, so I can tell you you can get used to it. I have gotten so used to it that at this point it's like hard for me to use something else like Coworker or something that is designed to remove you from the terminal. I, I'm not sure if I'll make, be able to make that switch.
Speaker A: And I think this is also pretty interesting. You, you just mentioned Cowork. Cowork is, is basically the using the terminal, but in a nice way. So I think a lot of people think, oh, this is magic, uh, what all the agents are doing there. Yeah, it's pretty much using the terminal in a nicer interface.
Speaker B: It is like to, to clarify for people, is more sandboxed. Like there are more reasons why I'm not switching from the terminal to, to Coworks. There aren't so many arguments for me not to use the Claude desktop app because that is functionally the equivalent. But Cowork is sandboxed. You don't have the same exact direct access to like files and functionality and like, um, yeah, I mean there's just a little bit more complexity that you can do in terminal.
Speaker A: Do you think it's for safety reasons or why are uh, there not the same functionalities and coworkers that I don't know.
Speaker B: I can imagine that it could be for safety reasons, but I'm not sure. Like I have a hard time imagining where they're going with that. Uh, maybe it is for to try to get more non engineering teams to, to adopt it and be able to control a bit more permission wise what's accessible and not. I can imagine that. Yeah.
Speaker A: Interesting. Okay, next step. Now we validated our insights. We made sure that everything's, um, aligned with the content that we shared. So we are ready to basically share our insights. How does the handoff looks like?
Speaker B: I think I even have a visual I could show you for this.
Speaker A: Yeah, sure. Great.
Speaker B: So if we've got insights, uh, you know, imagine you have an insights team or your design team kind of doing the insights work. You know that that often starts the process. Right? Like we're starting with insights. When we're doing things as well as we can be, then we're starting with that. And that usually informs maybe simultaneously either product decisions and design work. Starting or starting with writing a PRD and getting the decisions around product changes done first before handing off to design. All of this can be done in synced files. Ideally when a team is able to work either with Dropbox, where you're actually syncing very consistently from, from the cloud to your own desktop, then it's not usually a problem. And otherwise another term or another platform that's going to make designers cringe a little bit. But we can use GitHub and do git pull requests. If you're working in the terminal, it does start feeling like we're getting more technical here. I realize that, but I just want to flag that. That's a really good system for making sure that all of the teams actually get the same documents more or less in real time. You can trigger a pull request to actually just get all of the information you need before you start working on design. And even if it was just updated an hour ago, then you get everything you're supposed to have.
Speaker A: Could you explain real quick what a pull request is? I assume that most designers don't work with GitHub, but I think it's important to know
Speaker B: GitHub is basically an engineering platform for having all of these development files stored in one place that can be accessible from the terminal as you're working. Like traditionally as you're working and you're coding things, you could do a pull request that gets the latest versions of files that you need to keep working, and you do that all from within the terminal. Now what we can actually do as insights teams or design teams is have a repository, basically like a collection of files that is for Example, um, within a project like my Whoop Screen project, Latest Insights could be a folder and the Insights team could be pushing the latest analysis and findings and results into that folder, and the design team or the product team could pull that straight into their terminal where they're working. It's like another more advanced version of, if you're using Dropbox and you're sharing folders and everything, that usually works too. But pull requests tend to be a little bit more of a stable way of doing it and controlling and making sure that you have the latest version of everything. Whereas if you're using Dropbox, that tends to be stable too. But depending on how you have your sync settings set up, you might be a little bit delayed. So that's just more kind of, you have more control there. But the point is, we want to have a system where everything is file based, but whatever is delivered by Insights, as soon as they save that, we can kind of keep moving because we can immediately just pull those results into continued product and design work without someone having to send it to us in an email or a Slack message. So that gets pulled straight in, synced in read only, ideally. Right. So no one should be kind of modifying those files, but we just get kind of like extract what we need from the originals. And then as design or products start working, those teams have their own versions of the Claude MD file and other context and rules files in their folders. So they can kind of take the Insights but run their own workflows and in a more controlled environment where they've already hard coded their standards into documentation that Claude picks up when it needs it, and so on.
Speaker A: And, um, just a quick question between, um, you had your Woob app locally on your computer, right? Um, some things you, you probably also share with your team. How, how do you make sure that the things that you have locally on your computer are still up to date with the things that is on GitHub? And maybe, you know, my, uh, colleague has updated versions on these MD files on their computer and then they're working on different. How do you make sure that everything's aligned?
Speaker B: Yeah, so ideally this would be something that I can actually pull from GitHub. Um, so the whole folder might be like a shared team folder or even just something like you can choose whether it's the whole folder or it's data or it's the context files. Um, but that I would be able to pull the latest versions directly from a GitHub repository that the whole team is using. So in theory, I'm Working solo here. But if this is a team folder, then everyone needs to be able to access it from the same place. And so it would be kind of like the master version is the one that's updated in GitHub or in your Dropbox. Right. And you just need to make sure that either you have pulled it from GitHub, you've done a pull request, or you are syncing this in real time from Dropbox so you always have the latest version.
Speaker A: Okay, great. Yeah, that makes a lot of sense. Cool. Thanks for the tip. I assume that a lot of teams are struggling with the consistency and um, manually changing things. So I think makes total sense to pull this from GitHub, have there this shared, um, space and then pull all the content from GitHub and use it from there.
Speaker B: Yeah. One thing I highlighted here because I think it's really important to talk about an additional benefit that I see of using CLAUDE code or sort of like a file based system. The ability to automatically keep a record of decisions along the way. Because this has been a big topic for me the whole time that I have worked in research and product teams especially that the decisions made are not always based on insights. I think that's a frustration. A lot of insights people have, but also even when they aren't made on insights for good reason, because there are other kind of facts or forces that just override that for the moment, that's not documented very well or it's not shared with everyone very well. And so when we're working with this kind of file system where CLAUDE can also be set up to automatically take in the information from different parts of this workflow, also pull in from connected transcription services like if you're using granola or Fireflies or whatever to record your stakeholder meetings, pulling that into and keeping a log of the discussions and the decisions that are made throughout this process that can be so valuable both for just making sure everyone's on the same page and understands what decisions were made and where did they come from, and also keeping us accountable so that when we kind of redo some part of the product and we launch it and we either did or didn't base those changes on insights that were documented. That's here. Um, and so we launch something, we see the results, CLAUDE can automatically kind of pluck the information from the launch results, if that's in a database somewhere, take the decisions that were made, access the original insights and so on, and actually do a comparison on its own to help us understand how well are we making decisions what result did our decision have on the results that we had and so on? The documentation takes on a kind of life of its own here because we have the documentation like this file system throughout, that we're using as the backbone of any workflow we do. But we can also set up Claude to kind of document really interesting things along the way without us doing it at all.
Speaker A: Yeah, I think it's fascinating to see how you set up this whole folder structure, how well this is all documented. I think for me, super cool to see. So first of all, thanks for sharing. I think this is super, super interesting. Um, do you usually prepare these files for clients? As a researcher, is this something as like, who does that, you know, for inside a team? I'm curious,
Speaker B: Let me just come back to you here. Yeah, sometimes I do, but really it's done with Claude. I mean, so whether I do it or somebody else does it, it's really again, like back to my first principle of working with AI. It's somebody does a brain dump, right? Me, the team that I'm working with, somebody has to get their knowledge down on paper in one way or another. And then you feed that to Claude and Claude creates everything for you. So if you want to create some sort of decision making workflow or document decisions in a certain way, the best way is always to say what are the pieces that would be important for me? What's the outcome that we want to be able to get from this? Or what is the value that I would want from this? Is it about accountability or is it about improving our decision making capabilities over time and how well we make decisions as a team? Telling that sort of thing to Claude and then letting it put together all of the details and pieces into, uh, in that case it would probably be a hook or something that runs automatically to be capturing specific pieces of information from certain workflows that are running at certain times and then document it in the right way? Most things are that sort of template, like your brain, write down some notes, give it to Claude and then iterate from there.
Speaker A: And where would you say things break in this process? Are there some limitations as well that we should be aware of?
Speaker B: Yeah, for sure. I think the biggest places that things break are still where we put together workflows or documents that are really just not solid enough. All of this is not, uh, a 100% certain solution to avoiding hallucinations. Hallucinations or things being pieced together from different places in the wrong way by AI is still a fact of life using AI. So that's still absolutely. Potentially, uh, a breakpoint. I think there are other more technical problems. Like we were talking about certain people not using the same tools or functionality as others and therefore having outdated files, outdated folders. Um, that's a big barrier to making sure that every process that I've kind of talked about on a high level can work properly. Because if someone has updated all of those context files or really important Claude MD or rules files for doing certain types of work, and someone's using version two, two versions ago, they're going to be thinking that they're running a workflow and a process with the same standards baked in as you are, but they're not. And I think those are the biggest points where we can design really good workflows, we can design really good file structures, but we do need to still keep our eye on it and know that the technology or our use of it can still break down a little bit.
Speaker A: Cool. Yeah, makes so much sense. Right. Um, uh, thanks for sharing. Also the tips about the limitations. I think this definitely is a problem. Also like what you mentioned earlier with also fact checking, um, if the content is right, setting up your own MD M files to measure, uh, the results and see if everything's accurate. I think this makes a lot of sense. So, um, yeah, such a fascinating process how you set everything up with the different files and for all the listeners and viewers, if you would like to dive deeper into everything that Kaitlyn, just share it. If you're really into AI customer research, Caitlin is also running a course, especially on that, where you will go much deeper and really dive into all of that in a live cohort. Right. Like when is the next one actually?
Speaker B: June. So beginning of June. It's two weeks. We do two hands on live sessions which are like working sessions. So everyone leaves those having actually been build skills and complex workflows in skills and yeah, we talk about all of these things. Um, because there are a lot of complex things here to unpack. So I do my best to answer the kind of beginner blocker questions that come up as someone's like, I just opened the terminal for the first time in my life and what am I doing here?
Speaker A: Yeah, a lot of questions. Cool, that's wonderful. So I will definitely link it in the show notes. Um, there's still some spots left, right?
Speaker B: Yeah, yeah.
Speaker A: Cool. Okay, perfect. So if you're interested, you can find the link in the show notes. Definitely check it out. I think super interesting to dive deeper into that. Um, so really, really cool. Especially if it's live. I Think this is an extra bonus as right now if you don't need to watch any like, I don't know, YouTube tutorials, but if you can actually ask questions, I think this is so valuable, uh, right now, really be part of this community. So really, really cool that you're doing that. I love that. So what would you say are your last tips for people who would like to get started with Claude, with AI in research and customer research? Um, what are maybe like your three tips for people to get started?
Speaker B: Oh, honestly, the first one is what I've said if you few times, but I'll say it in clearer terms. Spend some time outside of a large language model by yourself. I know, reflecting on how you do things and how you do them. Well, I think this is the biggest lever for people that no one is spending enough time on. So if you can, before you go, I'm going to have Claude do this workflow for me. Think really critically about how you do it yourself, if you do it manually and what makes the outputs from your own work good, reliable, trustworthy and so on. Try to get really explicit about that. Very, very clear on that. It will make everything else easier because then you can hand those criteria and those checkpoints and so on off. The second is start thinking of agents really as just a series of files. I think I spent last year, uh, dabbling with agents here and there. And despite being an AI customer research person, feeling like behind on a lot of agentic things. But now I realize agents that work really well are essentially a sequence of very good files. So if you were to break down your own workflow into a series of files, what might that look like?
Speaker A: I love that. I love that because last year I also had this agent rabbit hole where, I don't know, I created my own agent. It was also complex. I spent so much time on it and what you just said, I think this is pure gold. This is exactly it. Right. It's all about the documentation, having really good documents of how things like basically the system prompt, right. Like foreign agent, a bit longer and more structured. This is what we need. Agents are not that complex. You can create your own agents today. It's very easy. You just need to document it. Well, thank, uh, you for pointing that out. Yeah.
Speaker B: And, oh, if I have to come up with a third, I feel like the third one is the hardest. I think we still should not hand off our judgment. So to be honest, I have decision making workflows that I have built in Claude code, uh, as skills where I get Claude to Walk through decision making processes with me. But I'm very careful not to let Claude guide me in a decision as far as I can control that or make the decision for me. I even tell it not to make a recommendation at the end of those. So I think we need to continue thinking of large language models as the execution layer for us. But uh, going back to the first tip, use your own brain at the very beginning. Download your own way of thinking about things and the criteria that you would have and the way that you would judge things in the first place. And then at the very end, make sure you're still making decisions on your own.
Speaker A: Cool, thank you. I think super important to not give the AI too much power, but that we still have our own judgment and taste and making our own decisions. Sometimes it's exhausting to make decisions, so we use AI, but in the long term, this is not the right way. We need to do it if we want it or not.
Speaker B: We need to keep the exhausting part a little bit. We need to prevent our brains from decaying by doing some of the more difficult thinking work. I think of it that way. If I feel like I haven't thought hard during the week, then I realize I have to cut out an AI workflow. Mhm.
Speaker A: That makes sense. Be critical about yourself and your own workflows. I think this is also a good reminder to, um. Yeah. Uh, and end this wonderful conversation with you. Caitlyn, thank you so much for being in the future of your ex. I loved this episode, so hands on. I also learned a ton, which I always love.
Speaker B: Cool.
Speaker A: And for everyone who would like to join the cohort, I will link it in the show notes so you can check it out, um, and learn more. And if people would like to ask questions, reach out to you. I think the. The best is probably LinkedIn, right?
Speaker B: Yeah, definitely. Great.
Speaker A: Cool then. Caitlyn, thank you so much for coming to the future of your ex and being my guest. I really loved it. Thank you.
Speaker B: Thank you for having me.
Speaker A: Mhm. Sam,
Other episodes covering the same guests and topics, from across The B2B Podcast Index.