The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Possible
Possible artwork

Inside an AI-powered personal dashboard | Matthew Tiemann

Possible · 2026-08-26 · 32 min

0:00--:--

Key moments - from our scoring

Substance score

63 / 100

Five dimensions, 20 points each

Insight Density13 / 20
Originality11 / 20
Guest Caliber15 / 20
Specificity & Evidence12 / 20
Conversational Craft12 / 20

Matthew Tiemann, who works at Foundry Logic on AI-driven energy grid efficiency, shares his journey from data analytics through robotics and logistics into building personal AI agent systems. The conversation explores how agents compress workflow timelines - turning week-long analysis tasks into minutes - and enable personalized, real-time dashboards that adapt to individual preferences and biometric data like sleep patterns from Oura rings. Tiemann's practical experiments reveal the gap between objective, code-based tasks (where agents excel deterministically) and subjective work like game asset design (where alignment and adversarial review agents become critical). His personal dashboard system, powered by agents that proactively draft meeting notes, suggest activities, and make multi-threaded decisions, demonstrates a shift from tool use to decision amplification. The discussion addresses the jagged edges: how agents drift off-course without frequent realignment, why 12-hour check-in windows optimize productivity without context-switching overhead, and how defining success criteria for subjective tasks requires more human oversight. This resonates with B2B operators managing knowledge work automation, particularly in forecasting, data curation, and complex decision-making where speed must be balanced against alignment and oversight.

Key takeaways

  • →Agents reduce analysis workflows from a week to under 30 minutes, enabling 1,000x more analysis depth and fidelity, transforming how data dashboards and business intelligence are built.
  • →Personalized, agent-driven dashboards - integrated with biometric data like sleep and goal tracking - can proactively surface decisions and reduce decision-making friction without removing human control.
  • →Subjective tasks (art, design, game development) require tighter agent alignment and adversarial sub-agents that harshly review against success criteria, while objective tasks (code) are more deterministic and need less oversight.
  • →Agents drift and make poor decisions on long-running tasks unless realigned frequently; 12-hour check-in windows with minimal friction provide the optimal balance between autonomy and course correction.
  • →Fast iteration powered by agents - building prototypes in hours instead of weeks - compresses the feedback loop and surfaces design flaws early, accelerating product discovery.

Guests

Matthew Tiemann

Topics in this episode

AI agentsCodex (agent platform)Personal dashboardsOura ring (biometric data)Demand forecasting systemsSub-agents and agent orchestrationAdversarial review agentsiOS game developmentNotion (task management)Foundry Logic (energy grid optimization)

Questions this episode answers

How can agents reduce time spent on repetitive analysis and forecasting tasks?

Agents can handle manual data entry, spreadsheet updates, and multi-step analysis across dozens of forecasts in minutes rather than the 30+ minutes of clicking required before. For example, updating 10 different forecasts with 5 variations can be done via voice instruction instead of manual spreadsheet manipulation, collapsing week-long analysis into 15-30 minute iteration cycles.

What's the difference between using one shared dashboard versus personalized agent-curated dashboards?

Shared dashboards require a template approach that fits general audiences but doesn't adapt to individual preferences. Agent-curated dashboards can build hundreds of personalized versions - each aligned to how a specific client wants to see their data - enabling better business question answering without the manual design overhead.

How do you prevent AI agents from making poor decisions on long-running tasks?

Define success criteria upfront, use sub-agents to review work against those criteria (especially for subjective tasks), and check in frequently - Matthew found a 12-hour window optimal. Without frequent realignment, agents drift off-topic after the first bad decision and are hard to correct.

What biometric and contextual data should agents have access to for daily productivity?

Access to sleep data via devices like Oura rings allows agents to flag when you're sleep-deprived before you make important decisions, provide context on mood and energy levels, and time dashboard reviews for when you're most cognitively sharp - usually an hour after waking.

Why is adversarial review by a second agent effective for improving quality?

A skeptical sub-agent critiques the main agent's work, preventing overconfidence and catching errors that a single agent might miss. This works especially well for subjective tasks like art or design, where you need multiple perspectives to validate alignment with a vision.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

13 / 20

The episode contains genuine insights about agent-powered workflows, personal dashboards, and task acceleration (collapsing week-long tasks to minutes), but much of the conversation cycles through the same ideas repeatedly rather than introducing new conceptual frameworks. The discussion of context layer challenges and adversarial agent review offers real substance, but overall density is diluted by meandering anecdotes and restating the same core themes.

I can do a thousand times the analysis at a greater depth and fidelity than ever before, which, I mean it's like superpowers of cognition.
you want to have a fine balance of where you actually need the agent to have the latest context because you are almost going to have to maintain and input that information into your shared context layer

Originality

11 / 20

The core framing - agents as cognitive multipliers, personal dashboards, and fail-fast iteration - closely mirrors existing AI thought leadership discourse. While the personal dashboard with Oura ring integration and the multi-agent review structure show some novelty, the contrarian insights are sparse. The episode largely reinforces familiar narratives about automation and productivity rather than challenging assumptions about when agents actually fail or their systemic limitations.

agents have unlocked tremendous agency in our day to day lives
if you're not trying to do things with AI that are failing, you're not exploring the edge enough

Guest Caliber

15 / 20

Matthew Tiemann demonstrates real practitioner credibility: hands-on experience across analytics, logistics, robotics, and energy systems; currently deployed at a portfolio company building AI agents in power grid optimization; and actively shipping experiments (game prototypes, personal dashboards). The hosts (Reid Hoffman, Parth Patil) also bring genuine operational experience. However, Tiemann is primarily an operator exploring AI applications rather than a senior executive at scale driving measurable business outcomes, slightly limiting the caliber.

Matthew and I go way back actually when uh, I think it was maybe like 10 years ago...he was in high school and he was looking for a job
I'm at one of our portfolio companies, um, starting to learn and get ready to deploy, figure out where we can deploy AI agents to get efficiency gains

Specificity & Evidence

12 / 20

The episode includes concrete examples (personal dashboard with Oura ring data, iOS game with 20-30 decision trees, Codex Chief of Staff thread, demand forecasting system updates) and real timelines (12-hour agent check-in cycles, 15-minute dashboard builds versus week-long prior process). However, evidence remains anecdotal rather than quantified; no hard metrics on productivity gains, financial impact, or deployment scale at Foundry Logic. Numbers mentioned (1000x analysis, week-to-minute compression) are illustrative, not substantiated.

collapsing that down to a minute and being able to then do the next minute and the next minute
that monitor is my personal dashboard that my agents will interact with...when I wake up in the morning, we'll walk into the living room

Conversational Craft

12 / 20

The hosts ask follow-up questions and occasionally probe deeper (e.g., 'Have you thought about how well it can approximate your level of energy?', 'what's another example where you failed?'), showing genuine curiosity. However, questioning often remains surface-level; hosts rarely challenge claims about productivity gains, request evidence, or probe contradictions. The conversation is conversational but lacks the adversarial sharpness or willingness to pressure that would expose gaps in thinking. Reid's closing comment about needing human engagement somewhat contradicts earlier claims about agent capability.

You were talking about collapsing a task that used to take one week...collapsing that down to a minute
where in your personal agent amplification have you seen the potential or facts for bad decisions?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B52%
  • Speaker A29%
  • Speaker C19%

Most-used words

agent35game33agents27context18dashboard16different16matthew15back15data13decision13almost11decisions11world10check10analysis9personal9

Episode notes

AI strategist Matthew Tiemann is building toward a world where AI agents work around the clock - and humans step in for the decisions that matter. Across his work at energy services company Foundry-Logic and his personal projects, he uses Claude Code, Codex, and Fable to automate knowledge work and accelerate software development. Those projects range from an AI-assisted iOS game to a personal dashboard informed by his goals, to-do list, and Oura Ring sleep data. Drawing on experience across data analytics, supply chain, robotics, logistics, and energy, Matthew joins Reid Hoffman and Parth Patil to explore proactive “heartbeat” agents and shared context. They also discuss adversarial review, 12-hour human check-ins, and AI’s limits in subjective judgment - revealing why human vision and alignment remain essential.

Full transcript

32 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: You were talking about collapsing a task that used to take one week, collapsing that down to a minute. And so you can imagine the amount of analysis you can do in one day. I can do a thousand times the analysis. Greater depth and fidelity than ever before, which, I mean it's like superpowers of cognition.

Speaker B: I actually bought a very small monitor for my, my place. That monitor is my personal dashboard that my agents will interact with. When I wake up in the morning, we'll walk into the living room. Instead of pulling up my phone or pulling up a newspaper, I have a dashboard that is personally curated to me.

Speaker A: The token grantee program gives $1,000 a week in tokens to high potential creators already deep in AI across film, gaming, comics, print and digital art. There are no tool restrictions. The freedom to choose is the point. Grantees also get access to my own custom fleet of agents and an ongoing collaboration with me. The goal, to close the gap between an idea and the world.

Speaker C: Matthew, welcome to Read Rifts. This is awesome. Um, you know, part of the kind of creativity as we explore kind of what can be done with AI and so Parth, uh, over to you to introduce Matthew and kick us off.

Speaker A: Oh man, super excited for this one. Um, welcome Matthew, welcome to the show. Uh, so good to have you man. Matthew and I go way back actually when uh, I think it was maybe like. Was it like 10 years ago? It might have been. I was at my first.

Speaker B: 2017-2018-2017-2017.

Speaker A: Okay, that's right, 2017. I was, I was working at data on data analytics at my, for, at a startup and Matthew was in high school and he was uh, looking, I guess his like he was looking for a job coming in out of high school and I was like, oh, I could totally like take a high school intern and like turn him into a data analyst. I saw that as a challenge. So I was like, oh, let's, let's uh, let's interview him. Let's take, bring him on the team. So we brought Matthew on the team and he immediately just like picked up data analytics on the fly as like on the job. And um, and I see like how quickly he learns and then over the last uh, I guess almost, yeah, nine years we've been, you know, I've just been throwing interesting hard problems at him and he's just been like, you know, pushing on them, you know, like trying new things. Going into his own career is like he's, he's deep into his own career, robotics, logistics experience. So he kind of expanded from the data world. To the world of, uh, the physical world and the world of uh, atoms. And when language models came out, um, we were talking, we were just bouncing ideas. And I went very deep on language models and code generation. And I basically told Matthew, I was like, I have a feeling that it is really important for us to get good at this before it becomes the next big thing. And he kind of trusted me. Back when I think the GPT3 era, where it was kind of like a fringe, the stuff barely worked but you know, we could connect the dots. It's like if this gets to a certain level, you know it's going to change the way we do analysis, the way we do almost all of our knowledge work. So it'd be cool to like get ahead of that curve and watch it. And so he kind of jumped joined on that journey. And we've been unpacking the themes of generative AI, the implications of code generation in an increasingly automated world and applying that into the real world. And so really excited to have you, Matthew. And now you work at Foundry Logic, uh, company which Reid invested in, uh, working on, um, building uh, a more efficient electrical uh, grid in the United States. So super exciting to have you on the show and can't wait to dive in.

Speaker B: Thank you. I'm actually forward deployed today, so to speak, traveling.

Speaker A: Oh, really?

Speaker B: Yes.

Speaker A: What does that mean?

Speaker B: Uh, I'm at one of our portfolio companies, um, starting to learn and get ready to deploy, figure out where we can deploy AI agents to get efficiency gains.

Speaker A: Awesome. So you've. Yeah, so I guess let's start with that. Like you've worked across a pretty wide variety of fields, right? So, um, analytics. We met, we were working in analytics. Then you got into logistics and robotics and hardware, software, you know, your side projects, you do VR, game design, um, all kinds of stuff. So I guess what drives you in those spaces and what common problems come up across those fields?

Speaker B: Yeah, so I guess my career so far has really been just exploring everything. So supply chain, robotics, logistics, uh, now in energy. And I think the piece that combines across all of those, uh, that brought me to agents is two things. I think. The first is I have a very high itch to uh, automate. So for example, if I'm in uh, the cloud code terminal and I take a screenshot, it really bothers me if I have to go drag that image into the cli. Um, so I actually set up a macro for myself to both take the image and paste it into my cloud code cli. Even that small, just little friction, uh, bothered Me, so I have this itch to automate. And I think the other piece of it is, um, as you said, when LLMs came out, it was the GPT4 era. You had. I had that feel, the AGI m moment quite early. And once you get that, you understand what agents can do. Um, and then across all of those different fields that I've been in, really the common theme is every business runs on spreadsheets, right? So when, when you see somebody using a spreadsheet, maintaining that spreadsheet, and you realize an agent could be helping you do this, it's hard to not try to bring agents into that. Um, it's like if I'm doing manual entry, I can't stand it because of that ish to automate and because I know what agents can do.

Speaker A: Yeah, I was the same. You know, I was one of the best at spreadsheets and Excel. And then one day I found Python and I was like, wait a minute, we should be doing this faster and more automated. And then when agents came out and they started writing Python, GPT4 started writing really good Python, I was like, oh, wow, I don't even have to do the spreadsheet, I don't even have to write, write the Python. I just have to describe the business problem and have the agents kind of do the work. And so it's like we get to do it again. You know, what we did 10 years ago with data analytics, we get to do again with this time. It's like the actual, like the agent layer, where it's like taking actions across the intelligence of the business.

Speaker C: And by the way, both of you, just so the people don't misunderstand that it's only the rote, like, kind of like whether it's Python or anything else, go to the amplification of it. Like, what new capabilities both in, like, in ability to do things in speed, et cetera, that kind of become game changers here. Because I think this is a good, uh, lens in everything else.

Speaker B: Yeah, I'll give an example of one. Um, uh, one of my other jobs was in, uh, demand, uh, forecasting. And you have a system that you update to keep that forecast. Right. But typically it would be very much clicking and doing a lot of analysis on that. Ideally you can just talk to the agent and say, hey, go do this analysis on these 10 different forecasts and um, update them in these five different ways. And the agent can just go do that. I don't have to sit there for 30 minutes clicking through each screen, clicking to update Filling in the new one and saving that, the agent should just be able to go in and actually do that itself. So that's just one place in my world that um, at least voice unlocks that significantly. And then there's also the intent to which uh, what you intend gets actually executed a lot faster.

Speaker A: Yeah. And I think you can also be greedy with the level of the amount of analysis that we can ask for and that can get done um, in one iteration and then also the depth. Right. So I feel like when just three years ago, four years ago, when I was doing analysis, uh, through code in Python and SQL, it would take me maybe like, it would take me like a week to build a dashboard into one slice of the business. And now when I use um, tools like ChatGPT, tools like Claude, I can build that first dashboard in almost like 15 minutes. And then I have so many follow up questions and I'm like, oh, let's look at this from this different lens. Let's uh, let's dive deeper into the like, give me drill down capability. I want to go down to the transaction level. I want to get a high level view so you can change the zoom, the granularity at which you're looking at the data. And because all of that happens in basically 15 minutes to like 30 minute iteration cycles. And pretty soon with the uh, high speed inference, once the Cerebras chips come online, it's going to feel almost instant. You were talking about collapsing a task that used to take one week and a full data analyst entire day, um, for a whole week, collapsing that down to a minute and being able to then do the next minute and the next minute. And so you can imagine the amount of analysis you can do in one day totally dwarfs uh, the analysts that we were just three years ago before these tools. So even though I'm no longer just a data analyst, I can do a thousand times the analysis at a, uh, greater depth and fidelity than ever before. Which, I mean it's like superpowers of cognition.

Speaker B: Right. And I'll just add one thing to that. Um, you mentioned one thing we did back when we were both in that analytics role was curating that dashboard. And that took time for us to design that dashboard into a way that it was generally applicable for, let's say a whole account management team to go interface with client. Now what I'm very interested in and going and building is very curated and personalized dashboards per client preference. Um, every client wants to see their data a little bit differently. They have slightly different Businesses. Having your agent understand their preferences and then go build that dashboard and curate it to each person, uh, unlocks a whole new world. As opposed to having one dashboard, you could have hundreds that are of the same quality and answer your business questions even better.

Speaker A: Right. We don't have to do the template approach just to save time.

Speaker B: Exactly. Yeah, yeah, yeah.

Speaker C: Look, one of the things I think is particularly interesting in your journey, Matthew, is how this leads you into building agents that can coordinate information and help run parts of your daily life. Because that's uh, something I think could be illustrative to everybody. But all of this kind of intensity and data science moving into these agents, running your daily life, say a little bit about like what you've been doing there, the path that you were accelerated in terms of getting there, and then kind of like what you see in the future as part of that.

Speaker B: Yeah, in my daily life I am obviously interfacing with agents everywhere. So using dictation, um, deploying multiple agents at once. I think the most important unlock for me has been the heartbeat agents that can wake up and proactively go and do something for you. Uh, so I have Codex doing that now for me. I have my main chief, uh, of staff thread in Codex that can then go talk to my other Codex threads. Uh, so I'll have different threads based on different projects I'm doing and I'll interface with kind of all of them. But the Chief of Staff 1 has been a very big unlock. Uh, so even just yesterday I get out of a meeting and I will see in my notion environment that my Codex agent has proactively taken action to draft a follow up to that meeting. Uh, like that proactivity from agents is I think an incredible unlock. One other really cool thing I'm doing is it's kind of an experiment, but I have, I actually bought a very small monitor for my, my place back home. That monitor is my personal dashboard that my agents will interact with. So when I wake up in the morning, I will walk into the living room, get some morning coffee or whatever it is, and uh, instead of pulling up my phone or pulling up a newspaper, I have a dashboard that is personally curated to me. That agent that's powering that dashboard knows a lot about me. So it knows my goals, it knows what my general to do list kind of looks like. Uh, it should know kind of where I'm at in that to do list as well. Uh, what's really cool is since it knows all that about me, it knows, for example, I like the Lakers Um, I'm a fan of the Lakers basketball team, so it'll know that, hey, there's a Lakers game on tonight. And it'll put that on my dashboard and say, hey, do you want to, uh, add this to your calendar to watch the Lakers game tonight? And so I'm still in control of my life, right? I'm still driving it, but now I just had that extra layer of push personally curated news or information insights that I can interface with and just get more value in my day to day because otherwise I could have totally missed that. And that's just something it knows that I like little things like that. And I think the really big highlight I want to make there is having that personal dashboard that you can also make decisions on. So I still, I'm the one making the decision. Do I want to actually add this to my calendar, yes or no? I think that that's kind of where things are going. Um, another thing I want to get onto my daily dashboard is the idea of a software factory being you have an agent that's kind of just running 247 against an app or a goal. Uh, I should just be able to wake up, say yes or no to an agent working on, let's say a video game and come back 12 hours later, see the progress of that. And on my dashboard it should just show me, hey, I need you to approve these three big decisions and then I'll go work for another 12 hours. Uh, so that's where I'd like to get that too is a point where it's just a decision machine for me and it's some way that I can have some control over my to do list calendar, Just get good insights but also make decisions out of.

Speaker A: Yeah. So when you talk about your personal dashboard, have you thought about how well it can approximate maybe your like, level of energy and your product or you know, productive momentum in your work? Like, like what keeps us in, in the flow, keeps us in motion. Have you thought how, have you thought about, like how it thinks about that bigger picture?

Speaker B: Yeah, it has access to my Oura ring data, so that's kind of useful. Just knows how much sleep I got. And maybe it should try to personally curate it to wake up at a certain time. Um, so I'm only reading it maybe an hour after I wake up or something. If it knows I'm maximally, uh, productive at that time. I think the other piece is if it were to bring decisions to you, it should know generally your goals and how to align that with your goals. Um, so like in the example of a game, it, it should try to, if it's developing a game asset for you, it should try to bring you three different versions of that. And maybe version one is kind of um, a certain art style, version two is another art style and version three is another one. So they're all uniquely different on purpose because then I can give it that direction and push it off in an alignment way that it's going to go further towards that direction. So instead of three assets that are super duper close together, maybe I want it to be uniquely different in a way that when I say go off and do this, I will come back 12 hours later and actually get something decently good.

Speaker A: We're, we're, we're like constantly. We're currently looking at re having an agent rebuild Reed's website and I asked for like 25 variations and today I'm like, my job is to review them. It's so interesting. I was like, I need 25 design variations. You know, go in a slightly different direction and the pro, the process becomes like this review thing. Um, um, but it's very, very exciting.

Speaker C: Well, and one of the reasons why that kind of review thing is, you know, we, we have this amazing amplification with these agents and it can, you know, cross, uh, check like. And by the way, Parth also has an aura ring so to stop the AI agents from waking them up all the time. Uh, you know, I think the actual

Speaker A: value of giving it access to my sleep data, uh, which I've noticed just in the last week was if I'm making an important decision, it reminding me that like, like if I haven't slept well and I have to make an important decision, just getting that reflected back to you before you make the decision. So you know, like, yeah, you might be cranky, but actually that's like more your sleep. You're not actually frustrated with the person you're talking to, but having that extra level of like context the agent makes that brings it, you know, brings you back to a level point before you make a really important decision. I think it's very interesting when you tie the, that biolog, like that energy level context back into um, your decision making process.

Speaker C: And by the way, all kinds of super interesting things about helping shape and guide and amplify the cognitive decisioning process. Now one of the things that we still do see in addition to all of the amazing applications is some places where it can uh, facilitate bad decisions or, or do something that's bad. I mean obviously, you know, kind of earlier this month, as the time of this recording, there was the Codex agent escaping to hack. Hugging face, bad decision in various ways. Um, but where in your kind of personal agent amplification have you seen the potential or facts for bad decisions? When, where are you shaping it, uh, to make it uh, kind of better. But where is the jagged edge of decision making for you, Matthew? Uh, in this, for the personal agents,

Speaker B: I think the most recent example that comes to mind is developing a video game. So actually the art style example I just gave has been really hard to get alignment on. You really have to align your agent before it's going to go off and do a long running task. Otherwise it's going to get off topic, it's going to start making bad decisions and the second it makes that first bad decision, it goes in another direction and it's really hard to kind of correct that. You almost have to revert back to that checkpoint, continue going towards the actual vision that you want it to get to. Spending the time up front to give it that alignment of whether it's context dumping, building the skills or memory that the agent needs to get that job done. I think that's a big piece of it. And then also having whatever context it needs to be able to fetch, um, having that context available but also up to date with your latest goals is a very interesting challenge as well that I'm interfacing with day to day now. Um, but I think the other piece of that is also just being able to check in frequently, uh, to realign that agent. It's really hard for an agent to go for two days straight and get it perfectly on the first try unless you really give it that upfront alignment. Uh, I think the, the 12 hour check in is almost my sweet spot of if I can check in and realign it before it's going to go off and make those mistakes, that's really good. And if you can make that stepping in point as frictionless as possible, uh, that is also of the most, most important to me because if I, if I'm checking in every hour then I'm doing a task and I have to context switch over to that. So spacing out the check in points and also making them as frictionless as possible has been really important to, for me to keep that alignment.

Speaker C: Go ahead and do you have um, you know, any like in terms of like a panoply in terms of how you use Hermes or things like multiple agents that are kind of, as it were, like a workflow agent to say, oh, we should get you to check back in before, like, context goes awry or check back in before a major port or portion of work happens. Like, how do you orchestrate that for your own kind of personal productivity on this? And what does that mean for future agent team orchestration?

Speaker B: Yeah, I think what's cool about the latest agent harnesses is you can build custom sub agents. So if I'm using cloud code, for instance, I will create a custom subagent that is specifically designed to review that task that I'm doing. So if it's generating an art asset for a video game, that subagent needs to understand, here's the art style we have built. Here are the 20 other art assets that we've built for the game. And you need to review this one that we're making that the main agent is making against all that and really having it define that success criteria well is the most important part. Um, but then when I'm interacting with my main agent, I will always say to go use that sub agent to check your work. Uh, and that sub agent needs to check it very harshly against the success criteria. So I found that to be a very good strategy. Uh, there's definitely ways to improve it, but I think it's a lot easier when you can define the success criteria for things that are more subjective, like building a game asset. Again, uh, it's hard for the agent to really determine that the asset it has just produced really matches that art style because it's still somewhat subjective, even if you have 20 examples. And for something like code, it's a lot easier because that's more deterministic and much less subjective. I mean, there's subjectivity in the architecture, but, uh, for things that are more subjective, it's a lot harder to get that right. And I think that's where you want to actually step in more, is for subjective tasks like that.

Speaker A: Right, right. That's a, that's a good way to frame it. The subjective versus objective deterministic. Um, and also, I agree, adversarial review is probably one of the most powerful. Creating a second agent, just telling it to be skeptical and to critique the work of the first agent. It helps, um, deal with the like, oh yeah, I'm so confident in this answer. And then it's like, well, how about, you know, you prove it to this other guy who's very skeptical and you have that second agent come in and be like, I m don't know about that. I think it's a very effective way to get a better final output. And you get more confidence when you have multiple perspectives attacking the problem. All right, so what is an experiment you've run recently related to your projects? One that worked and then maybe one that failed in an interesting way.

Speaker B: Over the weekend, I built an iOS game. And that iOS game, it kind of started in a place of failure, but that fast iteration is what unlocked the one I'm actually kind of happy with. Uh, so originally I had an idea and I had the agent just go in and scope out that idea. The idea was, uh, you tell a story and you swipe to make decisions on an iOS app and it actually, uh, edits a tabletop tile based world. Um, and on top of those decision, um, cards that you swipe on. And the problem is for that the game loop there is, you're kind of paying attention to too many things. But the reason I was able to figure that out so quickly is because agents helped me iterate and build an actual pretty good prototype of what I originally thought very fast. So I got to that point where I prototyped and I had that almost working first vision and I said, okay, let me play it. I played it. The game loop just did not click to me. So then I stepped back and you're

Speaker A: playing it on your phone, you have it.

Speaker B: Uh, I was just doing on the simulator on the Mac simulator for an iPhone. Uh, but even that was well enough to get a feel for how the game loop worked. Uh, so then I was able to quickly pivot into something that did work. And that's something that did work is me taking a step back and saying, okay, this format might work for a computer game with a longer game loop, but for an iOS game with a shorter game loop, I'm going to really hone this in on one feature of my original idea that I wanted. That original idea was the swiping to make decisions. The game that I actually ended up with is a full story game with 20, 30 different decision trees, uh, or different branches of the story that you can go down, um, that you just swipe left and right to tell the story. And what's cool is Fable actually helped me build a whole storyboard against every single decision that you can make in the game. So it's kind of like those storytelling games, but even more amplified because there's so many different decisions. And you can have Fable help you write that story, um, very effectively. So that's like a choose your own,

Speaker A: choose your own adventure.

Speaker B: Still things to refine there, but I was a lot happier. Yeah, choose your own adventure. IOS swipe left and, right. Uh, so excited to get that over the line. But that's where it really helped me find a game that I was happy with.

Speaker A: I guess like we, I find that like, you know, fail fast, learn quickly and try to iterate is a huge theme of the AI era. You can very quickly get more, more uh, rapid iteration. So curious, uh, another example of something where like you failed, you initially failed, but like you, you kind of like what you learned from that and where it took you and the unexpected path that you kind of get to your quality outputs from. Talk, uh, about another example maybe.

Speaker B: Yeah, so I'm very interested in the idea of, especially when it comes to personal agents, having a shared context layer of your latest context that I can use and make good decisions on. The problem that I've really hit with that is building that out where it's really hard to embed your latest context into, especially in an enterprise sense where it's almost a multiplayer game of you have many different people working towards the same goal. So if I'm working on a project and I need my agent to know the latest on the project, to draft an email to somebody, well, the latest on that project might have been a conversation I had with my coworker in the hall five minutes ago. So if I wanted my shared context later to have that information, I would almost have to constantly be updating it or wearing some sort of pin that, uh, is recording my full day and staying up to date with every little detail of my life that's going on. Uh, so the big, that was a challenge was I wanted to get it to the point where it could have that front edge context. Uh, but it's really hard to keep it at that front edge for everything. I think it's possible for some things it's almost, you want to pick and choose. And that's kind of what my learning was, is you want to have a fine balance of where you actually need the agent to have the latest context because you are almost going to have to maintain and input that information into your shared context layer. So that was a big learning. And I think that the outcome of it is that my new shared context layer is a lot more balanced in what I'm collecting and how I'm maintaining it.

Speaker C: We've been talking a lot about the kind of, the various kind of workflows and kind of efficiency and kind of amplification. This also kind of brings me to, you know, kind of what happens as people panic when they think about AI and creativity. So Matthew, what do you, what do you wish they understood about this and what's kind of like as it were, a lens or a uh, kind uh, of like hope and solidity for them as they think about, you know, AI and creativity in the future?

Speaker B: We all know that agents have unlocked tremendous agency in our day to day lives and I think that agency will. I'll go back to the game example. There should be more games. There just should be a lot more games. And a uh, lot of the pushback in the gaming community is they don't want to see AI in game, in game development. But for me I would rather have more great games to go experiment with and enjoy. And if there's a single developer out there who is building a game and can't afford to go hire a full art studio to help them build assets, but they can get that leverage out of using an agent, I think that by all means they should get that to go build a great game. Um, because it just democratizes the game dev cycle. It unlocks more great games and I'm really excited for having more games overall and more great games.

Speaker A: Yeah, I definitely look forward to the Cambrian explosion of the games. I think we're at the beginning. Yeah, it's like we're in the GPT2 era of game game generation. So this could get very interesting very quickly. Well, thanks for, thanks for joining us today, Matthew. Um, really learned a lot from, from this conversation, as I always do. Um, but before, before we leave, I would love to open it up, open up the field and if you have a question for me and Reid, um, that you'd like to uh, pose.

Speaker B: Yeah. So I'm really curious how you think about the gap between the incredible capability of the latest mod and the overall productivity gains in the economy. So what's the next big step for the industry to unlock those significant, uh, productivity gains? Uh, do you need a shared context layer that's way up to date on everything or do you need a really way better agent harness to wrap agents against problems? Even better. I'm uh, just curious your thoughts overall on where things are going.

Speaker C: So by the way, you definitely need those things. Uh, those are key. That's one of the reasons we have you and the other reason we have you chatting with us here on Redress is a lot of it is human engagement with this. Like, is even if you build, you know, much better context, you know, frameworks much better, you know, um, not just individual agent harnesses, but harnesses for a collection of agents in terms of, of how you work better. You need to have people pushing in it. And part of the reason why we were talking about like what are the know kind of like where are you leading the edge and learning where the edge works and where it doesn't work, like what are the learnings from failure is like one of the pieces of advice that I give individuals and organizations is if you're not trying to do things with AI that are failing, you're not, you're not exploring the edge enough because of how it's developing. And so I think a lot of the places where we see that we're going to see the productivity games is not just the frame further evolution of the tools. The tools are already pretty magical as per, you know what you know you're doing Matthew and Parth's doing and, and what a whole bunch of different things you can do here is but you need to be engaging them and what uh, in terms of how you work now that being said, part of what you're gesturing at which is hey you know, how do we you know get HERMES set up the right way? How do we get a context layer, how do we get enough memory and knowledge about what it is I'm doing and what I'm trying to do, how do we get you know, red team agents to get the quality the right level, how do we get check ins for you know, how do we operate? All of those things need to happen, all of that can be productized better. But like there's a lot of raw capability there right now. And so precisely the reason why we're doing these uh, redrifts with the pioneers and the scouts and the creators and the token fund grantees like you is because it's there now. Start deploying, start using it, you know, you'll get some ah, uh, that didn't work great. That's part of means you're learning fast enough. Uh, and you're, you're part of what this accelerating universe is. So um, I think that's, it's the human amplification, not the human engagement, not just the pure AI frameworks that are going to be key to seeing the realization in like the industries, in the economy and in terms of what's happening including some of the work that I'm most excited about what you're doing. The um, like the whole range is you know, kind of personal enhancement and I do think we'll see gaming too. Very cool. So Matthew, thank you Parth, always, always a pleasure.

Speaker B: Thank you so much.

Speaker A: Thank you Matthew. Great to have you.

Speaker C: Possible is produced by Palette Media. It's hosted by Ari Finger and me, Reid Hoffman. Our ah, showrunner is Sean Young. Possible is produced by Tenasi Delos, Katie Sanders, Spencer Strassmore, Emo Zhu, Aman Suri, Danny Garrison, Trent Barboza and Tafadzwa Nimurundwe.

Speaker B: Special thanks to Surya Yalamanchili, Sayda Sepieva, Ian Alice, Greg Beato, Parth Patil and Ben Rellis.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Why AI Pilots Fail: How to Escape AI Pilot Purgatory and Scale Enterprise AI with Ronnie Kwesi ColemanUsing AI at Work · on AI agents92 / 100
  • More Agents Than Employees: How Zapier Disrupted Itself Before AI CouldTalking AI · on AI agents90 / 100
  • PMs Are Fintech: How Column is Rethinking Property Management BankingThe Profitable Property Management Podcast · on AI agents89 / 100
  • 90% of AI prototypes never reach production (w/ Temporal's Samar Abbas) | AI BasicsThis Week in Startups · on AI agents81 / 100
  • Fighting Fire with Fire: How CyberProof Is Automating Cyber Defense with Edy AlmerCyber Sentries: AI Insight to Cloud Security · on AI agents80 / 100
  • From Billable Hours to Business Outcomes: How AI Is Rewriting the Rules of IT ConsultingEvolving the Enterprise · on AI agents78 / 100

More from Possible

All episodes →
  • America's aerospace rebirth97 / 100
  • The whole studio is one guy | Jonathan Brazeau
  • Making taxes fun with Pikachu and AI | Cadi Zhang
  • How Sougwen Chung teaches robots to pause
  • The artist using AI to sell brands | Joe Salvatore
Explore the best B2B AI & Data podcasts →
All Possible episodes →