The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Talking AI
Talking AI artwork

More Agents Than Employees: How Zapier Disrupted Itself Before AI Could

Talking AI · 2026-08-18 · 40 min

0:00--:--

Key moments - from our scoring

Substance score

70 / 100

Five dimensions, 20 points each

Insight Density14 / 20
Originality12 / 20
Guest Caliber18 / 20
Specificity & Evidence15 / 20
Conversational Craft11 / 20

When GPT-4 launched in March 2023, Wade Foster recognized the pace of AI advancement - faster iteration cycles, dramatic capability improvements, and rapidly declining costs - posed an existential threat to Zapier's core automation business. Rather than wait, he issued a company-wide Code Red, effectively shutting down for a week to run a hackathon that shifted AI adoption from 11% to over 50% of employees. Foster explains Zapier's differentiation strategy in the AI era: while frontier labs like OpenAI and Anthropic optimize for coding tasks, most enterprise workflows require hybrid architectures combining deterministic, rule-based processes (reliable but rigid) with agentic systems (flexible but costly and slower). Zapier's Automation Bench - a benchmark testing models on real workflows across sales, marketing, operations, support, finance, and HR - reveals even the best models (GPT-5.6 Solon Max) achieve only 18.1% success, underscoring why pure agent delegation fails. Foster discusses token optimization, the importance of rigorous private evals, and Zapier's positioning as the orchestration layer managing the interplay between workflows, models, and deterministic processes - the unsexy but critical infrastructure that frontier model companies won't build.

Key takeaways

  • →The top frontier models achieve only 18.1% success on end-to-end workflow automation tasks, making hybrid deterministic-plus-agent architectures more reliable and cost-efficient than pure agentic systems for most enterprise work.
  • →Model routing and evaluation frameworks matter more than raw model capability; many organizations unnecessarily use expensive frontier models for tasks adequately solved by cheaper open-source alternatives, mirroring Brian Armstrong's 50% token cost reduction at Coinbase.
  • →Deterministic workflows win on reliability and cost but fail on edge cases; agents handle randomness and unstructured data better but are slower and more expensive - the real value is blending both approaches rather than choosing one.
  • →Private, difficult evaluation datasets prevent benchmark gaming and ensure evals predict real-world performance; leaked benchmarks end up in training data, causing models to optimize for the benchmark rather than actual business outcomes.
  • →Habit formation and ongoing exposure (show-and-tell sessions, periodic hackathons, continuous tool releases) drive organizational AI adoption far more effectively than a single announcement, converting skeptics by demonstrating tangible use cases across functions.

Guests

Wade Foster

Topics in this episode

AI agentsZapierGPT-4Model routingToken optimizationDeterministic workflowsFrontier models (OpenAI, Anthropic)Code Red hackathonAutomation BenchHybrid workflow automation

Questions this episode answers

Why did Zapier issue a company-wide Code Red in March 2023?

Wade Foster recognized that GPT-4's combination of rapid iteration (released just six months after ChatGPT), dramatic capability jumps, and sharply declining costs signaled an imminent threat to Zapier's business model, necessitating an immediate strategic reset and organization-wide experimentation.

What is Automation Bench and what does the 18.1% score tell us?

Automation Bench tests frontier models on end-to-end workflow execution across six business functions (sales, marketing, operations, support, finance, HR); the top model's 18.1% success rate indicates that for most structured business tasks, pure agent delegation is still unreliable and that hybrid deterministic-plus-agent workflows are necessary.

What is the difference between deterministic workflows and agents?

Deterministic workflows (like traditional Zapier automations) execute the same way every time, providing reliability and cost efficiency but failing on edge cases; agents receive goals and decide their own execution path, handling randomness and unstructured data better but at higher cost and with lower reliability.

How should organizations choose between different AI models for a task?

Rather than defaulting to the most powerful frontier model, organizations should evaluate which model is actually sufficient for the task; Brian Armstrong's Coinbase optimization demonstrates that many routine workflows run adequately on cheaper open-source models, potentially cutting token spend by 50% without losing quality.

What makes an effective evaluation framework for AI models?

Good evals should be difficult for models but easy for humans (exposing true capability gaps), use entirely private test data to prevent benchmark gaming and training data leakage, and recognize that benchmarks are proxies for real-world work and can mask poor performance on tasks outside their narrow scope.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

14 / 20

The episode contains substantive discussion of agent vs. deterministic workflows, model evaluation methodology, token optimization, and organizational AI implementation strategy. However, there is notable filler including a mid-roll ad break, softball questions about daily recaps, and repetitive framings that could have been more concise. The best insights cluster around Automation Bench findings, hybrid agent-workflow architecture, and institutional vs. individual AI adoption - genuine operational concepts - but these are interspersed with generic advice.

So you have to think about what are the things that these frontier labs can't do or they want to. And you really just try and understand what are the things that are really valuable to customers that those companies can't fulfill.
The best evals are typically evals that are actually kind of easy for humans to do, but really hard for these models to do because now you're testing something really interesting.

Originality

12 / 20

Wade offers some genuinely fresh framings - the deterministic-vs-probabilistic workflow distinction is well-articulated and the Automation Bench benchmark result (18.1% success on real workflow tasks) is concrete and counterintuitive. His distinction between 'floor raisers' (broad organizational fluency via hackathons) and 'ceiling raisers' (institutional AI bottleneck-breaking) is useful. However, much of the narrative follows familiar patterns: the Code Red story is now well-known, and broader points about model selection, hybrid systems, and organizational transformation are circulating widely in enterprise AI discourse. Limited true contrarianism.

The best model right now is performing these tasks only 18.1% of the time. So the best model right now is performing these tasks only 18.1% of the time.
You have this like, hybrid agent workflow setup and you want something like Zapier to sort of manage those processes end to end.

Guest Caliber

18 / 20

Wade Foster is exceptionally well-calibrated for this topic: founder and 15-year CEO of a $5B SaaS infrastructure company directly affected by AI commoditization, early mover on organizational AI integration, and someone making real-time product decisions in the space. He speaks from operational experience, not theory. His credibility is reinforced by concrete examples (Zapier's employee count vs. AI agent count, Automation Bench data, actual product pivots). This is a practitioner at the right level and scale for B2B AI strategy discussion.

Zapier has more AI agents than it has employees.
Wade Foster co founded Zapier in 2011 and he has led it ever since, turning a scrap UI Combinator startup into the $5 billion plumbing of the SaaS era.

Specificity & Evidence

15 / 20

The episode contains strong specific data points: Automation Bench's 18.1% success rate, the Code Red hackathon moving employee AI adoption from 11% to 50% in one week, 600M automated tasks, and references to Coinbase's 50% token optimization. Concrete workflow examples (daily recap journal, email drafting) are included. However, many claims lack granular support: no details on the actual 'rethinking' that occurred post-Code Red, vague references to 'most sophisticated customers,' and limited financials on how Zapier itself is monetizing AI. The institutional AI section relies heavily on conceptual framing without case-level specificity.

We saw our daily usage of AI from our employee base go from about 11% of folks using it to over 50% in one week.
And as you can see if you look at the leaderboard, the top of the leaderboard right now is GPT 5.6 SOL max and it's scoring 18.1%.

Conversational Craft

11 / 20

The host asks reasonable setup questions and does pursue some follow-ups (e.g., on Automation Bench, on model efficiency), but overall the interview lacks edge and genuine productive friction. The host is respectful and knowledgeable but rarely presses back or challenges claims. Questions tend to invite Wade to elaborate on pre-formed talking points rather than exploring tensions or testing assumptions. The daily recap discussion, while interesting, goes largely unchallenged despite being somewhat generic productivity theater. The conversation meanders and doesn't drill into the hardest questions around organizational change resistance, financial impact of AI transformation, or Zapier's actual business model shifts.

What was going through your head at that moment in time? I think it was a seminal moment for everybody, but I think few leaders took it to that extreme and really saw what was coming.
Yeah, and there's several rabbit holes I want to go down here.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker A74%
  • Speaker B26%

Most-used words

models28model25zapier24inside21best18build17folks14workflow14better13important13deterministic13agent13task13product12daily12gets12

Episode notes

The best AI model in the world just scored 18.1%. On Zapier's own benchmark for real business work - the cross-app tasks any white-collar worker does every day - even the top frontier model completes them barely one time in five. That's the number Wade Foster keeps pointing at, and he runs an automation company that stands to gain from the hype. Instead, he makes the case for what actually works right now: not turning a model loose, but blending deterministic workflows with agents where each is strong. In this episode of Talking AI, Matt Paige sits down with Wade Foster, co-founder and CEO of Zapier, who built a scrappy Y Combinator startup into the $5 billion plumbing of the SaaS era on barely a million dollars raised. Foster called a company-wide “code red” the week GPT-4 launched, and he's spent the years since rewiring how Zapier - and its customers - actually use AI.

Full transcript

40 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: We looked at those three things and said, holy cow. If this continues with any sort of pace at all, this changes the entire industry. And so we felt like we needed to take a quick breather and say, okay, we gotta reset roadmaps, we gotta reset how we think about our vision, we gotta rethink, we think about automation. Hence the Code Red.

Speaker B: Welcome to the Talking AI podcast where we talk AI with both experts in the field and early adopters. I'm your host Matt Page and we're here to demystify AI for you so you can get some value from it. Let's talk some A Zapier has more AI agents than it has employees. And across its platform, nearly 600 million tasks have already been automated with the number climbing every second, literally. There's a running counter on their website and Wade Foster co founded Zapier in 2011 and he has led it ever since, turning a scrap UI Combinator startup into the $5 billion plumbing of the SaaS era on barely a million dollars raised. Now comes what may be his most exciting chapter yet. Deliberately disrupting his own company before AI does it for. Welcome to the show, Wade.

Speaker A: Yeah, thanks for having me.

Speaker B: I'm excited for this discussion. Zapier is one of those iconic names in the space and you've run it for 15 years now. But I want to go back to March of 2023 when GPT4 first dropped and for the first time in company history, you issued a company wide code red and effectively shut down the company for a week. What was going through your head at that moment in time? I think it was a seminal moment for everybody, but I think few leaders took it to that extreme and really saw what was coming.

Speaker A: Yeah, well, the GPT4 launch for us was pretty eye opening. We had seen obviously ChatGPT launch, we played around with it, uh, product was fantastic, really enjoyed using it, but it didn't really create this crazy sense of urgency inside the company, at least not yet. It was, hey, this would be cool. How can we operate better with this? How can we make better products with this? It was, it caught our Curiosity. Yeah, the GPT4 launch, which was about six months later, had a couple characteristics. One, six months later, it was pretty quick. Two, the difference between 3.5 and 4 was pretty big. Uh, the capabilities had just advanced a lot. And three, the cost curve was meaningfully coming down. And so we looked at those three things and said, holy cow, if this continues with any sort of pace at all, this changes the entire industry. And so we felt like we Needed to take a quick breather and say, okay, we got to reset roadmaps. We got to reset how we think about our vision. We got to rethink how we think about automation. Hence the Code Red. And a big part of what we did around the Code Red was, you know, we did a lot of stuff, but probably the most impactful was the hackathon where we paused the company for a week and we said, hey, everybody, just go build. If you're an engineer, play with the APIs. If you're not in engineering, go mess around with ChatGPT. That was really the main tool at the time, and just get a sense of, like, what is possible, what is coming with these things. And, you know, we saw our daily usage of AI from our employee base go from about 11% of folks using it to over 50% in one week. And so that really was kind of like that jumpstart that we needed to say, okay, something important is happening here.

Speaker B: I was not received. You literally essentially stalled the company for a week. I got to imagine some people were like, yeah, okay, weren't taking it as seriously. How did people react to that?

Speaker A: Yeah, this is 2023, right? The AI frenzy is just starting. It's not in peak fervor as it is now. There's a mixed reaction. Honestly, I think there were some folks who had already been using the technology who felt like I was behind. They were like, come on, Wade, we should go faster. We should go faster. We should go faster. There was definitely folks in the company that felt like this was unnecessary. It was. I was creating chaos where there was. Didn't need to be any. I was called sensational and things like that. And it felt like the way I think about all, uh, that was just. It was a moment where I had to be really, like, really help educate the company. Really had to be clear about what I believed and what I felt was coming to help people understand why this was such a critical moment. And it wasn't just sounding the alarm bells unnecessarily. And so it. It definitely was like a. A moment of just, whoa, what's going on?

Speaker B: So. So you got people excited. You got up to 50% usage. But I feel like a lot after a lot of those, you have the excitement, and then human nature kicks in. You go back to how you did things. How did you actually get that to stick? In terms of habit formation, which I think is one of the biggest undervalued things of this entire transformation, is for creatures of habit, at the end of the day, totally we like the way things are.

Speaker A: Yeah. I think there's a couple moves that help it stick. One, you gotta just keep doing show and tell. So on our own hands we would have show and tell, where we'd show off. What are you building with AI? And it's not just engineers. We just have everybody from without, from across the company, myself included, that would just show off things that we're doing. And people learn by seeing. You'd see someone do a cool thing and you go oh, I should try that out or oh, wow, I have that exact same problem. So just the act of just seeing people and getting exposure to it, you start to feel the art of the possible. You periodically step back and run more of those hackathons. So you give people more space to come back in and pick up where they left off, see the new capabilities. That's one of the fun and exciting things is that these models are, we're constantly getting released new models that have new capabilities, new powers. The application layer is figuring out how to do more things with them. The harnesses are getting better. So there's this just pace of improvement that is pretty invigorating because the things you can't do today, you are very likely to be able to do not that far in the future. And so you just gotta get people in the habit of keep trying. It really isn't. The answer isn't, hey, no, it's not possible. It's, it's not yet. It's not possible yet.

Speaker B: Uh, quick break in the pod. I keep hearing the same pattern with companies. I talk to quads helping employees move faster. But in many companies the business itself hasn't changed the value. Uh, still trapped in isolated chats and experiments. And that execution gap is why Ford deployed engineers have become one of AI's most talked about deployment models they embed with your team instead of advising from the outside. It's also why the FDE model is now central to every client engagement we lead at Hatchworks AI. As an official Anthropic partner, we embed Anthropic certified fdes to identify high value business problems, build and deploy the solution and put governance and security around it. Then transfer the capability back to your team. If your cloud rollout's still mostly individual usage. Check out how Hatrick's AI FDES work@hatchworks.com Claude FD. You can also find it in the show notes.

Speaker A: Now back to the show in most of these areas. And you just keep the exposure up, keep giving people places to do it, keep sharing and along the way you just kind of, the momentum starts to build. It's like a snowball.

Speaker B: Yeah. And you mentioned the application layer. There's different layers of the stack. You got the frontier models, you have the application layer, obviously chips and things like that. How do you think of the level of disruption, modes, differentiation, all of these things? Because literally, Zapier is a workflow automation tool. Right. And that's one of the things that AI is best at. And people can go into Quad or Codex and have it do things. How do you think about differentiation in this new era? You're obviously not going out and building a model, but you're leveraging models in a major way throughout your entire platform.

Speaker A: Sure, yeah. You know, I think a big part of it is you have to think about, you know, what are the things that these companies, uh, the frontier, the models, um, what are the things that they can't do or they want to. Like that's a really important question. And, uh, Sam and Dario and the leadership teams at those companies have been pretty forthcoming about, you know, the places that they intend to invest and the places that, you know, they sort of want the ecosystem to invest in. And so you'd like to pay attention to that. And then you really just try and understand what are the things that are really valuable to customers that those companies can't fulfill. So, for example, you know, most of these folks are pretty nervous about vendor lock in and they want to have the best capabilities no matter what model company it is. And so, you know, a lot of these application layers have model routing baked into them or have exposure to different tools. Zapier does as well, where you can pick and choose from the best. You can allow Zapier to sort of figure out what the right tool for the job is. You know, you're trying to figure out what are the capabilities that is not best served by AI. So right now there's a lot of token maxing going on. And it's, uh, in the best interests of these labs to have people continue to do token maxing. But it turns out there's a lot of workflows that you benefit from having deterministic workflows that run on code that are backed in a reliable way. Those are things that are not possible. Or you have this like, hybrid agent workflow setup and you want something like Zapier to sort of manage those processes end to end. And so, you know, a lot of this is just trying to think through, like, where are the places that customers really value and you're just not Going to get it from one of those labs. And how do you build, you know, a great product for those types of use cases that are going to be really important to a segment of the M market?

Speaker B: Yeah, and there's several rabbit holes I want to go down here. But you triggered one thing, the deterministic versus probabilistic nature. And Zapier's put out this benchmark called Automation Bench, which is effectively this leaderboard. And I uh, actually like you to talk through what it is and what it's scoring. But I think it hits at that point at uh, where, at uh, which when a model is executing on a workflow it's not perfect all the time. Even when you're using some of the most, the best frontier models out there Even I see GPT 5, 6, which the day of this recording just launched what about an hour ago or so. It's uh, listed on there. So what's the purpose of Automation Bench? How do you think about that and what's leaving this benchmark?

Speaker A: So you bet. So Automation Bench, it's a benchmark that is used to test these various models on end to end workflow execution. And we have a whole set of tests that we test against and it's across six business functions. So sales, marketing, operations support, finance and hr. These are the types of workflows that you know, you or I, any white collar worker will do every day working inside these tools. And it's meant to test how effective these models are at completing those tasks and at uh, what price points. And as you can see if you look at the leaderboard, the top of the leaderboard right now is GPT 5.6 SOL max and it's scoring 18.1%. So the best model right now is performing These tasks only 18.1% of the time. And I think what that screams to me is that hey, for a lot of things, these workflow related tasks, you probably don't yet want to delegate these to a model to go do them. And furthermore they're all expensive as well. So you know, if you can, if there are parts of your workflow that you actually can run deterministically, you're going to have higher reliability and lower costs. But there are definitely workflows you want an agent to do it working with unstructured task, writing code, generating emails, like there's a whole bunch of things that only a model can do. And so that's where these like hybrid agent workflow setups I think are so powerful and underrated right now where you can get the best of both worlds. And that's where Zapier just invests a ton of time where you can, in natural language describe what is you want to go build. And we're going to help optimize what gets built on the other side. So it's not just purely a model running amok, token maxing and you spending a lot of money on these things and getting out the other side, you're getting something really purpose built for the task you want done. And we're trying to be smart about how that gets built so that it isn't just willy nilly spending tokens for you left and right.

Speaker B: Yeah, and I think that it's an interesting pattern because you're allowing AI agents models to leverage technology in a deterministic way. It's that marrying it's very similar to how a human would go and leverage technology in a sense. What, like are there other bottlenecks or constraints as to why the top tier model is only scoring 18%? Are there other culprits at play other than the fact that probabilistic versus deterministic, are there avenues to actually increase that capability?

Speaker A: Yeah, you know, look, I can sort of only speculate because I'm not building these models. But you know, from what I understand is a lot of how these frontier labs are optimizing is they're optimizing for coding models. And one of the great things about building these models to focus on coding is that coding is easily verifiable. You know, you can have tests that run against code and you can see did the code work or did it not, did it perform the task or did it not. And so it's a really great place for AI to work because you can know that there is a factually correct way to do this or not. Now there is some knowledge work that has that same characteristic. Uh, you know, maybe there's accounting workflows, you know, you either balance the budget or you did not. Right. Those types of things. I suspect the models will continue to get a lot better at. But there's also a lot of white collar work where it's a lot more subjective. Is this a great email or not? Maybe we might have different styles of communication and uh, we might say, well, I really like this type of email. It was short and tight. You might say, well, I like this other email that had a lot of detail in it. So there's these other types of tasks that happen inside of knowledge work that have a little more variability to them. And so I think this is where when we talk about like white collar work and AI being able to automate a lot of this stuff, I do think it will be able to automate it. But the quality aspect, as long as there's sort of a, ah, human evaluating and judging it, those tasks I think are going to be trickier for these AIs to objectively say, hey, it is absolutely better. I think you could have some humans that say yes, it is better. And I think you could have some that say I don't think so. I think I'd rather have it this way. Which is also why I think you also want to have multi models and you're seeing this now bear out where these models have different capabilities. And so you might say, well, I really like having you uh, know the OPUS model be the one that like talks to me because I really like how it talks but I really want to have you know, these GPT 5, 6 model like write all the code because it's like a really good engineer on this stuff. So like you're starting to see. It's weird to call it a personality, but it is almost like a personality that sort of emerges from these different

Speaker B: models that gets into what I think anthropic. Just put out something around consciousness and the J space, which is a whole interesting separate topic. And then you have people, what is it when you. They're identifying patterns of AI, uh, in terms of writing and whatnot and they're purposely trying to do things to look not like it's AI, which is just productive in, in some ways. But on um, the token maxing front, I think you're in an interesting position in Zapier, you're both using AI internally obviously. Right. So there's a cost component there. But Zapier, literally the product itself, people are consuming tokens using the product, so you're getting hit from both sides. So how do you think about model efficiency? We had Coinbase, uh, CEO recently came out and optimizing which models they were using, saving 50% roughly on tokens. And tokens usage is still growing. How do you think about that in the nature of your business where it's on both sides of the equation?

Speaker A: Yeah, well I think, you know, this is a big opportunity for us to help our customers navigate a lot of this stuff. You know, when you go into, I think many organizations they don't really know like what is the best model to choose for a job. Like most people don't sit down and think about that kind of stuff and so they just say ah, you know I just want to use the best model for the job. And you know, they might be doing pretty mundane tasks where it's like, ah, you know, build my daily brief, but you're sicking Fable 5 on it. And it's like, well, that's just way overkill. You don't need that type of model for the job. And you know, furthermore, I think these models are starting to get where they are very smart and the incremental advancements that we're seeing in them are now past what is required for many types of tasks inside of the workplace. And so that's why you can see, you know, somebody like Brian Armstrong at ah, Coinbase say, hey, we're going to do a lot of optimization around our spin here because some of these open source models are plenty capable for a wide variety of tasks inside of Coinbase. And so we no longer need to be on the frontier for certain types of tasks. And that is where I think we're going to start to see a lot more optimization into the future. And so it's going to be really interesting to see how this plays out. My guess is that we're going to live in a world where most of the tokens are consumed by open source models, but most of the spend still goes to the frontier.

Speaker B: Yeah, no, that's a great point. It is funny, there is this element of model FOMO where it's, yeah, maybe the lower tier model would be okay, but maybe I may get that extra bit of something from use. I'm guilty of that all the time. But it comes down to, especially when you're doing things in a more structured way, the evals and evaluation. So how do you think about that? And is there a method for actually creating evals to determine, okay, this is, this model is perfectly fine for XYZ task.

Speaker A: Yeah. What I've learned about what makes a good eval is a couple things I think. One, you really do want to have your eval be tough. It should be really hard for these models to achieve these tasks. And the best evals are typically evals that are actually kind of easy for humans to do, but really hard for these models to do because now you're testing something really interesting where you're like, huh, huh. There is a true capability gap there. So I think that's one that's really important. The second thing I think that's really important is it's really valuable to have the test data entirely private. This really, this prevents basically the models ending up like benchmark maxing on these by accident. I think one of the things we've often seen is that as these benchmarks get leaked, all of that stuff gets leaked into the training data. And then all of a sudden the model is like hyper focus on fixing it. And so the model gets good at the benchmark, but it doesn't necessarily get better at, uh, just the things you might actually be exposed to in real life. Because ultimately a benchmark is still just a proxy for the type of work you do in real life. It's not actually real life. And I think this is what can happen. A lot of the times when folks get frustrated with AI in the real world, you'll see these benchmarks where it's 99% or whatever. This is a crazy model. And then you go use it yourself on a task and you're like, dang, it kind of sucked at that. Like, how is it so good at this? But it kind of stinks at this other thing. And it's because the benchmark is still just a small representation of all, like the knowledge that sort of humans might ask the model to go achieve.

Speaker B: Yeah, it totally makes a lot of sense. I want to get into the topic of agents. Right? You obviously are leveraging agents within the business, within zapier, but the term gets thrown around like crazy. How do you define what an agent is for our audience? Because I'm sure there's a lot of folks that have differing opinions or have no clue how to actually define it.

Speaker A: Sure, I think locally agents have just become basically like things that automatically do work. But I do think it's important for folks to understand the differences between a deterministic workflow and a purely agentic system. A deterministic workflow is something that's been around for ages. Pre AI, it's a program. It's zapier in some ways is build agents before agents exist, where it's like, hey, a trigger happens over here and we're going to automate. We're going to automate adding this customer to a CRM and then we're going to alert somebody inside of Slack and all that sort of stuff. These are just the workflows that sort of exist inside of a company. And these things have only grown in popularity in the age of AI. Now, an agent is slightly different. Way an agent actually performs is you give it a goal. You say, hey, I want you to go complete this task. And then the agent gets to go decide how it wants to go execute on that thing. So based on its training data, based on how the model works, it will Say, maybe I'll do this task and then I'll do this next. And then with this combination of data, I can go complete my goal. And every time it gets a similar task, it may not execute it the same way every time. It might choose to follow a different path and it still might get to the same outcome. It might get to a slightly different outcome, but it behaves more like a human does where you give it a task and it just goes about its day. Might mostly do it the same way, but it doesn't necessarily follow strict reels. Now, there's pros and cons to both of these models. A deterministic workflow. What you get is reliability, consistency, and cost advantages. It does the exact task the same way every single time. So that's the real benefit that you get from these deterministic workflows. The failure case is it's rigid. What happens if something goes wrong? What happens if something unexpected comes up? What if you have to work with, like, messy sets of data, like big unstructured, um, tests or images or things like that. All of a sudden it gets really, it gets a lot harder to do the task inside of these determinix nistic workflows. So you run into those problems. Agents are almost the exact opposite because they have the agency to go complete the task. Reliability isn't quite so good. Cost is a lot higher. As a result, it can be a little bit slower to complete these things. However, they can be a lot better at handling edge cases, because an edge case pops up and it starts to reason through it. What should I do about this? How could I go solve this one? And it might find its way through the problem. It can handle a lot more randomness inside the situation. And so that is like the blurring of these two systems. Now, oftentimes the magic that I find is inside of a business setting, you often don't have just purely deterministic setups and purely agentic ones. So the magic is like when you can actually blend these worlds together, where you might say the first couple steps of these always just work the same way every single time. And so we're going to just have that be purely deterministic. But right here we have one of these ambiguous situations that right now a human is making all these judgment calls on, because we can't really do anything with it. Okay, let's put an agent right in the middle of that. But then once it spits out an output, we're going to feed it back into a deterministic system to complete the Loop. And I think the most sophisticated customers we see, uh, the folks on the frontier are often getting really smart about how they do these things. They're doing the stuff like Brian and Coinbase is doing where they're mixing models or thinking about what is a workflow versus what is an agent. They're really starting to optimize these steps. Most people aren't there yet, but I do think that is the world we will find ourselves in as more and more tokens get spent in these organizations.

Speaker B: Yeah, and that's the beautiful thing you mentioned is the marrying of both, because to your point, when it was purely deterministic, you had to figure out every edge case or you had failure patterns all over the place. But now with probabilistic, it's a whole new world. But how do you think of iterating on a workflow and optimizing it over time? Because I feel like there is that element of being able to improve something. Because back to the deterministic nature is it did the thing or it didn't, but now that you have AI in there, there's this using the same term again, probabilistic nature of the workflow where it can improve over time. Any thoughts on, like, people that are actually doing these things in their business? How do you, like, methodically think about starting small and improving a, uh, workflow that you're trying to automate over time?

Speaker A: Yeah, yeah. I think what the best folks are doing is they have these, they're building their own mini evals for these setups. And as more and more scenarios come m through, anytime the agent fails to complete it, they're analyzing why did it seem to go wrong? And based on how it went wrong. Let me see if I can give it another test case or another scenario that it can compare against. And then they'll add that to the eval suite and then they'll run more through and then each time that happens, they'll just keep adding more examples to it. And now look, not always more examples doesn't always make it better. So sometimes they are having to subtract from it. But by and large they're, they're like. The best way I can describe it is these little test scenarios act as like a guide to the model where you're trying to steer it in the direction that you want more over time. And the more you do this, you can usually eke out higher and higher percentage of accuracy with these models.

Speaker B: And so, uh, zapier, you were effectively built on the API explosion as the basis of Being able to provide this product and MCP is the new thing. Connectors, being able to connect all your tools together. What similarities do you see between Those 2? And APIs are still a massively important thing in this new age too. So how do you think of MCP versus AI and this ability to connect things together?

Speaker A: Well, they're really just the extension of the same thing. At the end of the day, it really is just about connecting to the tools that you use. And so APIs and for most of our history have just been running a call in a very specific way to sort of extract data from a tool or perform an action in another tool, uh, so on and so forth. MCP just allows that for an agent to do it. And the nice thing about what MCP gives agents is the ability to kind of search and discover and kind of work in a sort of fuzzier way. Uh, and so you don't have to be as like, specific about like the way to perform the task. The mcp, Gordon, sort of gives the agent the ability to kind of go figure it out on its own, but under the hood, it's still effectively doing the same thing.

Speaker B: So in terms of the nature of jobs and work, there's a lot of debate on is AI taking jobs, is it creating new jobs? What is your take there? And I guess what are you seeing in terms of your own people? How would you describe the ideal type of worker right now? What type of qualities do they have? So you're hiring somebody new at Xavier. What are you looking for?

Speaker A: Yeah, the most general thing I think I can give is that the people who are like, very curious, who are doers, who have a high level of caring about their work, are just absolutely thriving right now. AI is giving them a jetpack to go figure out all sorts of things. And so whether you're just entering the workforce or you're on the cusp of retirement, these characteristics seem to accelerate people who really orient themselves that way.

Speaker B: Yeah, totally. And so let's shift to this. I'm curious, how are you using AI in your day to day life? Or what is the most unique, interesting use case of AI that may be not a common one that everybody's playing with today? Anything weird off the wall, weird off the wall stuff?

Speaker A: I do a lot of.

Speaker B: We're just super productive. It's something that just.

Speaker A: Yeah, I mean, my favorite thing is. Yeah, I mean, I. One of my favorite workflows is a little bit mundane. So a lot of folks like to talk about their daily brief. Candidly, I Think the daily brief is like just okay, it's like fine for most people. But the flip of that is the daily recap. I think the daily recap is one a lot of people are sleeping on. So what does the daily recap do? It basically reviews everything that happened for you in that day. So you can have it go loop over all of the digital exhaust that you have from the day. So this might be emails, Slack, meeting recordings, everything, your calendar. And it can help summarize everything that's happened. It can help track down the key action items. Even better than that, it can actually start to take action items. So if you were in a meeting and you said, hey, I'll make sure to make an introduction to this person or I'll make sure I'll follow up with this. Great, it'll already have those emails drafted for you. It already have like sample comms ready to go for you. And so you can do a lot of this stuff. The other nice thing I have it do is I often have it prompt me to say, how'd the day go? And so it's a very lightweight journal for me. And so usually I'll just uh, do voice to text. Yeah, I'll just do voice to text real quick and I'll say, these three things were awesome about my day. These three things sucked about my day. And in the moment, that's not hugely valuable to me. But now I've got, I've been doing this for probably about six months now. And so I've got six months of data about what gets me pumped up during the day and what sort of makes me frustrated during the day. And so now I can do a whole bunch of interesting analysis on that and say there's some commonalities here that makes zapier more successful, makes me more successful. And then there's some anti patterns here where I find myself not being at my best or find the company not being at my best. And so how do I try and architect my day and architect my work to avoid some of those scenarios? And so I actually find that the daily recap to be just under discussed. And I think it's way more valuable than the daily prep.

Speaker B: No, I love that there's this memory component as well. Like when you're doing that reflection, it's learning over time. That's one thing for the, for lack of a better word, daily brief that I do is it's actually learning about the previous day. If I've slipped on something three times in a row, it knows. Right. So there's this element of not repeating the same thing again essentially over time. And so in terms of work. So, um, maybe wrap it here. But the companies are obviously, they know the impact of AI. They know it's important. It's the top of everybody's strategy. But from somebody that literally called the code red at the beginning of GPT4, what recommendations would you give to companies to actually drive this adoption internally? Because it's not just about rolling out a bunch of licenses saying, hey, go do this thing. Yeah. What advice would you give to companies other than start using Zapier to automate some things throughout your.

Speaker A: Of course. Use Zapier for everything. Yeah, no, I think I, uh, think there's. I would probably boil it down to a handful of things and I would bucket it into two categories. I would say one I would call floor razors and the other category I'd call ceiling razors. Now the floor razors, this one's really important. This is about building widespread AI and fluency inside your whole organization. It's building a comfort level and excitement and energy for what is possible here. Now, the floor. The best floor razor activities are generally things like hack and hackathons, workshops, lunch and learns. Any place where you're asking people to put their hands on the keyboard and actually learn and play around with this stuff. What can they do that practical in their day to day? We just talked about the daily recap and the daily brief. Great, go build one. Everyone should have something like that they can go feel, figure out. And whether you're an engineer or a marketer or a sales rep or an accountant or an hr, there's probably something in your day that you can use to build something that would just make your life better. And then you can start to build stuff that helps your team and all that sort of stuff. So these types of hackathons are just so powerful for doing this. Now the other thing you get from this is you build a culture of an experimentation and you build a culture of excitement around what's possible. With AI, this becomes really important because you're going to have to go do some more interesting and harder things down the line. You might want to rethink like how the company works top to bottom. You might want to rethink job families and job structures. You might want to rethink just a lot of core concepts. And uh, it's really important if you built some of these, this for raising culture internally because now people are going to be a little more supportive of that because they're going to have A much more practical understanding of what's possible with these tools, what's not possible with these tools. And so any other changes you're making in the organization will be less scary because it comes grounded in actually information and knowledge versus just people responding to the external headlines. I find the external headlines and what is actually happening in companies to be, uh, honestly pretty different in a lot of cases. So that to me is like the first category, then the second category is the ceiling raisers here. This one's important. It's how do you go from individual AI to institutional AI? And this one I think a lot of companies are struggling with right now. They have figured out how to accelerate a lot of individuals inside their company. We can all point to that engineer, that marketer, whoever who's doing like crazy stuff with AI and it's massively more productive. But it's a lot harder to go find companies where you can say, wow, this company is growing so much faster, or they've reduced their cost basis or all this, the company is totally transformed with AI. Those are just a lot harder to find those examples. And I think a lot of this boils down to just because you accelerate one individual inside of a company, it doesn't mean that you've actually accelerated the whole company as a whole. You've just moved the bottleneck around. And so these institutional AI use cases are a lot more challenging. You have to look around and try and break down what are the bottlenecks that prevent your company from growing? What is the thing that prevents you from getting another customer or making that customer happier or uh, making that customer grow faster. And oftentimes those are things like how do we actually truly accelerate time to market for a key product? How do we actually improve conversion rates in our sales and marketing funnel? How do we actually take waste and cost out of this business? That's unnecessarily there. And so those you really do have to rethink some of the core ways in which you go to market or you build your product or you operate internally. And that requires a lot more tops down leadership and identification of the big levers inside of an organization. Much harder to do, much harder to do than the individual AI setup. But the rewards are much, much larger. Uh, if you do accomplish this, and I would say there's very few companies who have really pushed the envelope on this. Even the companies you read about, I would say are still not as good as the things you've heard.

Speaker B: Yeah, ah, it's so spot on the individual productivity, but it's not spanning across the org. How do you think of the org of the future? You have Jack Dorsey talking about this hierarchy to intelligence. There's a lot of hypotheses of how the nature of an organization is going to evolve. Any thoughts or perspectives?

Speaker A: Yeah, I think Jack has got a lot of interesting ideas. The shift from individual to institutional AI is one of the things I'm really excited about. And I think the big bottleneck is you want to figure out how to increase the amount of time that your organization spends on those activities that are more likely to create a customer. It's how do you ship a, uh, product faster? It's how do you do better sales and marketing? And so you want to increase the time energy, people that are spit on that. The challenge is as companies grow, you often have a lot of work that goes into stuff that's not that you have all the. You have a bunch of managers inside of a company. You have a bunch of telephone games that sort of have to transport information from one team to another. If you've ever seen that movie, Office Space, like that's kind of what happens to companies as they grow. It's the, the TPS report that gets handed from one department to another. And a lot of that stuff doesn't really actually add value to the customer. It doesn't make them happier. At the end of the day, it's just something that kind of has to get done because that was the best way to solve the problem inside that company today. So what I get really excited is how do you actually put AI at the center of how your company operates? How do you build a brain so that the company actually has all of its knowledge institutionalized around it? How do you put automation and tools on top of that brain so now the AI can take action on those efforts? How do you put a governance layer around that so that it's really clear what humans can and cannot do in that system? And also it's clear what agents can and cannot do inside of that system. And then how do you put a harness around that so that there is an interface between the human layer and the AI layer inside the company to take action? All stuff. I think if you do those things really well, you can start to build a system inside your organization that does take out a lot of that operational overhead where you can start to run closed loop meetings where you can start to have the AI acting on your behalf in certain places. And if you do that well, you can push more of your energy and effort to things that humans are Uniquely good at, which is this product development, which is the sales and marketing side of the house. Those things are stuff that distinctly bring value to customers. And that I think is when organizations get a lot more fun, but I think it's also when they get a lot more leverage from AI.

Speaker B: Yeah, it's so akin to system design, system architecture in a way. And um, last question for you. How do you think about your product roadmap? Things are evolving so quickly and do you leverage AI on the strategy front, thinking through the product roadmap, what you're going to build next and how Zapier evolves? Has that changed at all for you over the years?

Speaker A: Well, yeah, you, uh, know, I think the thing that we have noticed is that the like six month roadmap is basically dead. What I see internally, what I hear from other companies is that you're having to be a lot more nimble on your toes about where you're heading and what is possible with these models. And so, you know, you're often thinking two, three weeks ahead, you know, at most. Now that said, I do think there are certain things that you're always trying to figure out which are like, what are the foundational things that we do care about and how do we sort of keep those at the center of where we're heading? We know that this is the mission, we know that these capabilities are where we're heading. And so we're going to go hill climb against this and we'll hill climb against this for 6 months, 12 months, for however long as we think is important. But the way in which we get there, this sort of scoped out epics and stories and things like that, that stuff is kind of falling to the wayside, uh, in favor of these much more agile two week roadmaps, or it's weird to even call them a roadmap. It's sort of just like we just know we need to go achieve these things. Let's go figure it out.

Speaker B: Yeah, it's like there's tent poles in your strategy to wrap us up. Who are you following? Other CEOs, other leaders, other people in the space? Are there a few folks that you consistently look to for leaders in the world of either your world, AI? Uh, in general, yeah.

Speaker A: There's a lot of companies I think right now that are trying to be on the frontier. Obviously you pay attention to the labs themselves. Those folks are really interesting to pay attention to on the application layer. People who are maybe more closer to my seat. You look at the Brian at Coinbase and the ramp crew, the folks over at DoorDash are doing really interesting things. So there's a whole bunch of folks that are trying to figure out, like, what do we, what, how do we operate differently with these things? What do we build that's different with these things that I think are really useful to pay attention to these companies and CEOs and whatnot.

Speaker B: Awesome, Wade. Well, thank you for being on. Thank you for Talking AI. Where can people go learn more about Zapier? And honestly, if. If you've used Zapier in the past and haven't used it in while, like, the product has evolved so much, I was just playing around with it recently. But where can people, you know, find you and learn more?

Speaker A: Yes, uh, I mean, you can follow me on, uh, LinkedIn or X, I'm Wade Foster there. But you know, if there's one thing I'd recommend that's different about Zapier is I would try and install it into your agent harness of choice. So you might be using Claude or ChatGPT, you might be using Codex or Claude code or whatnot. To me, this is the best new, interesting capabilities that we've added to Zapier. And so I think you'll have a really fun time playing around with all of the connections that Zapier has inside of these new tools. I think that is the one really exciting thing that you should check out with Zapier.

Speaker B: No. Such a great call. It's like a new modality and way of working, in a sense. Wade, thank you for being on.

Speaker A: Thanks for having me.

Speaker B: Matt, thanks for listening to the Talking AI Podcast. If you enjoy the show, give us a follow or subscribe on your favorite podcast podcast. And don't forget, forget to leave us a review. We love those. For more info on talking AI, visit talking aipodcast ah.com Quick break in the pod if you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high impact tailored AI use cases based on your business, your goals, your pain points and your industry. No fluff, no generic use cases, just real ideas that fit your business and the rank by ROI potential. It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com AI opportunity finder.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Why Companies Force AI: Zapier’s CEO Wade Foster Explains | Belkins Podcast Episode #19Belkins Podcast · features Wade Foster72 / 100
  • 5 Rules for Building AI Agents That Work in Production | Nan Yu & Jacob ShumwayBehind the Craft · on AI agents91 / 100
  • Where's the Smart Money Going in AI? Rob May, Co-Founder & CEO of NeuroMetric AI, on Inference, ROI, and the Bets That MatterMaking Data Simple · on Frontier models (OpenAI, Anthropic)91 / 100
  • Building A Company of AI Super Workers - Tracy St. DicFNDN Series · on Zapier87 / 100
  • Inside the AI Hiring Pipeline: Interns, Apprentices, and Full-Time Coworkers | Vinay Gidwaney & Mike Sullivan, OneDigitalThe AI Why with Liam Lawson · on AI agents87 / 100
  • Stop Asking What AI Can Do. Ask What Your Staff Hates to Do.Small Business Big AI · on AI agents84 / 100

More from Talking AI

All episodes →
  • 99% Correct Is Still Failure: The Last Mile for Mission-Critical AI62 / 100
  • The State of AI 2026 Mid-Year Reality Check
  • Context, Control, Collaboration: Why Capability Was Never the Bottleneck
  • Past the Productivity Ceiling: Rebuilding the Enterprise from First Principles
  • The VC's Lens: How AI Is Rewriting the Rules of Defensibility
Explore the best B2B AI & Data podcasts →
All Talking AI episodes →