The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/The Road to Accountable AI
The Road to Accountable AI artwork

Logan Kelly (Waxell): The Accidental Agent Governance Company

The Road to Accountable AI · 2026-06-18 · 33 min

0:00--:--

Key moments - from our scoring

Substance score

62 / 100

Five dimensions, 20 points each

Insight Density13 / 20
Originality12 / 20
Guest Caliber15 / 20
Specificity & Evidence11 / 20
Conversational Craft11 / 20

Waxell started as a sales engagement platform but pivoted into agent governance after realizing early deployments exposed dangerous attack surfaces - agents spending tokens uncontrollably, exfiltrating data through tools like Claude Code and Codeium, and spawning cascading child agents outside governance perimeters. Logan Kelly describes Waxell's three-layer solution: Waxell Connect (monitoring on-machine AI tools), an SDK for instrumenting custom agents like LangChain implementations, and a governed runtime that encrypts entire agent environments. The control plane applies dynamic policy injection - drawing from frameworks like NIST and OWASP LLM - with 45+ policy categories that executives can set intent-first while engineers scope implementations. Unlike point solutions or vendor-locked walled gardens (ServiceNow, Salesforce Agent Force, Claude), Waxell positions itself as platform-agnostic overlay governance, handling the non-deterministic nature of agent behavior by identifying deterministic invariants (tool calls, function inputs) where policies can be enforced before, during, and after execution. The company emphasizes making governance accessible to non-engineers (finance, legal, compliance) through custom dashboards and blast radius visualization rather than requiring engineering expertise.

Key takeaways

  • →Agent governance differs fundamentally from traditional MLOps or data governance because LLMs instruct agents to take actions on machines with escalated privileges, creating data exfiltration and supply chain attack vectors that existing security tools don't intercept.
  • →Policy enforcement in agentic systems requires finding deterministic invariants (like tool calls and function parameters) within probabilistic agent behavior, then injecting governance rules at runtime with sub-millisecond latency to prevent user-experience degradation and bypass attempts.
  • →Organizations should implement platform-agnostic governance overlays rather than vendor-locked solutions, because the AI agent ecosystem evolves daily with new frameworks and tools that need to remain pluggable within a single pane of glass.
  • →Agent lineage tracking with cryptographic runtime stamping is essential for managing parent-child-sibling agent relationships and enabling kill-switch functionality when agents iteratively spawn sub-agents consuming cloud resources uncontrollably.
  • →Current risk tolerance for agents is zero - there are no trade-offs; companies should implement cost, data access, and compliance policies simultaneously because they create only millisecond performance overhead while dramatically reducing blast radius exposure.

In this episode

  1. 1From Sales Engagement Platform to Agent Governance
  2. 2Understanding the Agent Governance Problem Space
  3. 3The Control Plane Solution: Three Sets of Agents
  4. 4Making Governance Accessible to Non-Engineers
  5. 5Key Customer Challenges: Cost, Data Exfiltration, and Security
  6. 6Handling Non-Deterministic AI: Rules and Deterministic Invariants
  7. 7Agent Lineage and Kill Switches for Spawned Agents
  8. 8Avoiding Vendor Lock-in and Building Open Governance Overlays

Mentioned

WaxellLogan KellyKevin WerbeckOpenAIClaudeClaude CodeClaude ArtifactsCloud CoworkLangChainServiceNowNVIDIADatadog

Guests

Logan Kelly

Topics in this episode

Claude CodeLangChainMCPs (Model Context Protocol)Agent lineageCodeiumWaxell ConnectDynamic policy injectionNIST compliance frameworkOWASP LLM frameworkCloud resources kill switch

Questions this episode answers

What makes agent governance different from traditional API gateway or MLOps governance?

Agent governance must intercept at the runtime level where LLMs instruct applications to take actions on machines (like bash commands, database queries, file access), creating attack surfaces that don't exist in traditional prompt-to-output flows and requiring protection before the agent action executes, not just monitoring afterward.

How do you enforce policies when agents behave non-deterministically?

Waxell identifies deterministic invariants within probabilistic behavior - like the specific tools an agent will call and their input parameters - and injects governance rules at those decision points in sub-millisecond timeframes, enforcing yes/no binary rules before execution occurs.

What happens when agents spawn multiple child agents automatically?

Waxell uses agent lineage tracking with cryptographic runtime stamping to maintain parent-child-sibling identity relationships, allowing operators to fire kill switches that shut down entire agent families in one action, preventing resource exhaustion from iterative agent spawning.

Should companies choose a vendor-locked governance solution or platform-agnostic overlay?

Platform-agnostic overlays are preferable because the agent ecosystem changes daily with new frameworks; vendor-locked solutions like ServiceNow, Salesforce Agent Force, or Claude-specific governance become obsolete quickly unless you can remain plugged into a neutral control plane that handles multiple agent frameworks simultaneously.

Who should have visibility and control over agent governance policies in an organization?

Non-engineers including finance, legal, and compliance teams should have intent-first visibility and policy-setting capability through custom dashboards showing blast radius and risk, while engineers maintain deep audit trail and tool-level scoping access in the same control plane interface.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

13 / 20

The episode contains moderately useful architectural insights about agent governance - particularly the control plane concept, policy injection mechanisms, and the distinction between agent governance vs. traditional AI governance. However, there's substantial filler (throat-clearing explanations, repeated points about non-engineer accessibility) and several abstract passages that lack concrete substance. The technical details about lineage, kill switches, and deterministic invariants are valuable but underexplained.

The problem though with that agentic governance looks to solve is basically that model is telling an application what to do on that machine. That is a whole level of of you know kind of complexity, right.
our goal is to find in every policy category that we work with, you set rules. And then our goal and and how we've structured our systems is to find the deterministic invariance of that particular thing that's happening.

Originality

12 / 20

The framing of agents as requiring an observability + governance 'control plane' layer is moderately fresh, and the insight about policy parity driving multi-vendor strategies is useful. However, most of the underlying concepts (cost control, data access restrictions, kill switches, audit trails) are borrowed from traditional security/compliance playbooks. The 'deterministic invariants' framing is somewhat original but not deeply explored. Few genuinely counterintuitive claims.

Everything that could possibly happen around agents is connected. You have the governance or the observability, and then you have the governance that can then, you know, is what um, you know, kind of the rest of the org, not just the engineers, can can kind of log into and and use to, you know, control things.
I think there's so much model parity, right? I think there's so much provider parity that a a well constructed organization is going to have the kind of workflows internally to say this model is right for this or this provider is right for this and this provider's right for that.

Guest Caliber

15 / 20

Logan Kelly is CEO of an early-stage governance startup with direct operational experience building agent control systems and first-hand exposure to real customer problems. He demonstrates hands-on technical knowledge (virtual machines, function interception, cryptographic stamping) and has clearly iterated through product design with multiple customer cohorts. However, he hasn't managed agents at enterprise scale or previously built systems in adjacent domains, limiting the depth of battle-tested perspective.

So we started as a sales engagement platform kind of built from the ground up with AI. Uh and and quickly realized that people didn't really want to learn another software
We just saw a supply chain attack in NPM where like people are just have clawed code, you know, downloading stuff onto their machine

Specificity & Evidence

11 / 20

The episode offers moderate specificity: named tools (Claude Code, Copilot, LangChain, Datadog), framework categories, policy counts (45 policy categories, 15,000 API URLs cataloged), and runnable examples (bash scripts, file system access, database deletion). However, there are few concrete customer outcomes, no named client examples, no specific cost figures or metrics showing the governance impact, and vague claims about 'thousands of agents.' Much of the risk discussion remains abstract.

we have 45 different policy categories
I think of like fifteen thousand different uh API URLs and we're constantly building that

Conversational Craft

11 / 20

Kevin asks solid foundational questions (how Waxell pivoted, how the control plane works, how policies scale) and probes on concrete challenges (child agents, kill switches, shadow agents). However, he rarely pushes back on vague claims, doesn't ask for specific customer wins or financial impact, and misses opportunities to challenge Logan's assertions about 'no trade-offs' or the superiority of the general-purpose model. Follow-ups are polite rather than rigorous.

Walk us through then what the solution you built looks like.
And how do you solve it? Given that these are non deterministic tools, there is this agent interaction problem that's out there.

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

governance32agents30agent30different21data15start15machine15code13cloud13call12problem12cool11policy10control9built9policies9

Episode notes

Logan Kelly never set out to build an AI governance solution. Waxell spun out of CallSine, an AI-native sales engagement platform, when the team realized that agents that could act on their own produced a cascade of problems: burning through tokens, accessing databases, creating data-quality issues, and generally doing things no one had explicitly approved. Unable to find existing tooling that addressed the problems effectively, the team built a control plane for agents, which became the foundation of Waxell. In this episode in our series on governing AI agents, CEO Logan Kelly emphasizes that governance should be legible to finance, legal, and compliance teams, not just developers. As he explains, agent governance is less about exotic AI risks than about visibility and control over things companies already care about, such as cost, data access, and who's allowed to do what. Kelly makes the case that the worst outcome isn't an agent misbehaving but companies losing trust in agents altogether and missing their value - arguing that every major technology, from cloud computing onward, arrived with new risks that good governance ultimately made manageable.

Full transcript

33 min

Transcribed and scored by The B2B Podcast Index.

Hi, I'm Kevin Werbeck, Professor of Legal Studies and Business Ethics at the Wardens School of the University of Pennsylvania. For decades, I studied emerging technologies from broadband to blockchain. Today, AI is promising to transform our world. But AI needs accountability, mechanisms to ensure it's developed and deployed in responsible, safe, and trustworthy ways.

On this podcast, I speak with the experts leading the charge for accountable AI. Waxhell started out creating a sales engagement platform and ended up building a control plane for autonomous AI agents to address the challenges of agents spending wildly, leaking data, exposing vulnerabilities, and spotting more agents outside the governance perimeter. In this conversation, CEO Logan Kelly walks me through what agent governance looks like in practice and what it means to make it accessible to business people in addition to entrepreneurs.

Logan, great to have you. Welcome to The Road to Accountable AI. It's Kevin. I just bear.

As I understand it, you didn't start out trying to build an agent governance company. Uh so tell us the the story of how you wound up doing Waxella. Yeah. So we started as a sales engagement platform kind of built from the ground up with AI.

Uh and and quickly realized that people didn't really want to learn another software, especially we were focused at that time on kind of the SMB, right? So they had Hopspot and some of them had the sales lofts and outreach.ios of the world. There's a million tools.

So we said, what if we could just do everything for the for the user and that, you know, kind of has become what we call agents. Um, and then once we launched the agents, it was like, wow, we're not, we're not uh kind of at the mercy of them clicking a button to to be, you know, spending money on the tokens. That agent starts working, it's gonna burn all the tokens. Then it becomes a data, you know, quality per issue.

And then it becomes uh, did it say the right thing? And so As we launched these, like early last year, we started to to look around for platforms that could do this and and we couldn't really find anything because it was, you know, we were pretty early and it's just such a complicated world where if somebody solves this problem, there's a lot of open source code and and this kind of stuff. And that kind of led us to we have a small team and we, you know, we've got smart people who are not engineers.

And so we needed to give them a surface that they could really. be a part of that and that, you know, as we, you know, have progressed in the company, that's become a a big value prop because agentic governance shouldn't just be a bunch of engineers running around with their hair on fire. It should be, you know, finance, legal, compliance. Everybody needs to have the visibility and the the ability to do that.

So uh that we kind of innovated through you know, the the problem and and ended up getting a lot more excited about, you know, the the controlling agents and and building a really sweet governance surface as opposed to, you know, just being another sales engagement tool. So you define the the problem set. How do you understand what the solution is that's needed to address the agent governance challenge? Yeah.

Uh it in the early days, it was problems that came to us, right? So you see the the behavior of an agent, and these are probabilistic and very, very complex, right? So, you know, there's this you can start to try to govern the prompt, right? And you can start to try to inject, you know, different instructions into a prompt, but you start to realize very quickly as you deploy agents or you should that.

You start to need to the the AI is telling the agent to take a certain action. So now it's running a function and now it's accessing a database and now you have human identity, right? So the early days of understanding kind of this like primordial soup of all the different problems that could be contained in in, you know, a gentic governance were kind of experiential. And then we took a step back and said, we need a lab that we can start to you know, really experiment and and test with different use cases and and different, you know, agentic frameworks and different run times and all of this.

So we've gone from experiential to very much a uh, you know, day-to-day new new system comes in into play and and now you put it into the lab and you try to break it, you try to see what the behavior is and, you know, who who who in a company is going to to be, you know, or what element of a company is it a cybersecurity issue or uh finance issue, these kinds of Walk us through then what the solution you built looks like. Yeah, so we call it a we call it the control plane. So uh in an organization, you have what what we would classify as three sets of agents.

So a lot of companies who are just getting started with agents, the agents that they know best are things like cloud code, cloud cowork, right? Uh co-pilot, the these kinds of things that are interfaces with a uh you know, with a user, but they're not built. by the company. And that kind of creates this whole surface of okay, you have an agent, right?

Where a user is saying to do something and now it's going to go access the file system or it's going to go, you know, potentially like in cloud code, like could delete a production database. Stop it, it's happened. And so so we have that. And so that's called waxel connect where people we we have on machine uh apps that that get downloaded and then we have the kind of control plane, our interface where the organization can start to put policies in place using what we call dynamic policy injection that says, you know, if this happens, and we have 45 different policy categories.

And then that extends to okay, I've got a you know lane chain agent that is doing, you know, actuarial, you know, work automated for us. You have to instrument that. So we have an SDK to instrument that that that data is also then, you know, the observability and the governance is done through the control plane. And then the third piece is we have a governed runtime system that uh allows agents to run in a uh, you know, kind of the entire envelope, the entire runtime is encrypted, has, you know, all these different, you know, access points for governance, where like human identity or non human identity is managed and all of that.

So You put that all together, we have what we call the control plane. Everything that could possibly happen around agents is connected. You have the governance or the observability, and then you have the governance that can then, you know, is what um, you know, kind of the rest of the org, not just the engineers, can can kind of log into and and use to, you know, control things. Yeah, so how unique is it to each organization?

You talk about lots of different policies and so forth. Does does a company need to understand what it needs for agent governance in order to configure it from it the the cool thing about how we did it is like an executive, right? So a business owner, right? Not an engineer, comes in and they look at their they look at their their, you know, the line item on their PL and it's like they go to their CFO, like, whoa, where what was that?

Right. And then they're like, it's like, we spent a lot of money on on, you know, open AI last month, right? It's like, okay, cool. I want to control cost.

We have, you know, in in Waxle, we have this thing called a uh policy pack where it's like, Somebody who is a non-business owner can say, hey, we need to implement these policies. And then an engineer can then scope them down. So maybe there's a tool or a workflow or a a certain class of agents or a user set, these kinds of things. So they can scope it down and make it a little bit less, you know, every single agent shouldn't be able to spend more than this amount of money, right?

Which I think becomes problematic when you have millions of different agent executions. So, so we've built it so that there's kind of objective overlays and then there's you know, security framework and and compliance framework overlays, right? We've got like a a NIST pack and an OWASP LLM pack, right? So I think the way we've tried to architect this is the non engineers can be intent first in the platform.

The engineers can get, you know, all the kind of, you know, deep, deep, deep, you know, tool call agent, user group, you know, audit trails, all that kind of stuff if they need. What is it you mentioned a couple of times the engagement with non engineers, both the in your own team and and in the customers. What what does it take to to build something that addresses the needs of those groups? I mean, anybody who's built software products for people, right?

It's a lot of iteration and a lot of like, you know, conversation and and what what's important and that kind of thing. And and one of the cool things now, and this isn't the first product I've ever built. One of the cool things now with AI is it's such an important piece of the puzzle for companies, right? Like there's board pressure, there's market pressure, right?

It is it's something that has to get done. And I think that When we when we're in these conversations, we're getting a lot of feedback about, you know, this is how I look at the world, right? And it's like, okay, cool. We can translate the this data and and what we've seen is is these really cool conversations with customers that say, This is this is kind of what I'm facing.

Awesome. Well, here's a custom dashboard for that. Right. And and that allows people to kind of short circuit needing to learn an entire concept of agenda governance, which by the way, changes every day, right?

Like we have a a lab that's running, you know, a thousand agents a day on all these different frameworks. So like if you think you know what it is, wait a day, right? W what are the things that the the customers, the perspective customers come to you with most? You you mentioned that the cost issue is one thing, but what are the other challenges?

Yeah, I think cost is cost is where people start. I think it's the the easy place to start. Um, but then kind of uh with these on machine, you know, cloud co work, cloud code, codex, these kinds of things, there's a a big concern about, you know, are we exfiltrating RIP? Are we exfiltrating PII?

Uh does cloud co work? Or Claude Coates somehow get around the the kind of you know role-based access that I have implemented in my company. Um and so uh I was talking to a chief security officer the other day, and and there's this like concept of like, I almost don't want to know because once I know, right, there's gonna be a lot of work involved. And and so I think it where a lot of the work right now is starting to happen and needs to happen is understanding that Claude Code and cowork and all these tools are super powerful, but there needs to be the layer that we already have when you think about like email and and just connecting so connecting an app to your Microsoft O365 account gets admin permission.

So you have MCPs that people are just adding into their, you know, agents. And it's like, are do we know that those tools are are actually trustworthy? So are you fingerprinting them and and these kinds of things? So so I think there's this external agent.

piece that is where most companies are starting. Uh I think the the further ahead companies or the kind of more like data science driven companies are are, you know, okay, I've got this, you know, eight, eight agent kind of fleet that works together to solve one problem. How do I, you know, use something that's better than like a a, you know, the normal APM, like a datadog or something, to really see how those agents are, you know, we call it lineage, how they're all communicating with each other.

But I think the real thing that we need that organizations are worried about and and, you know, are focused on with us is that, you know, what's on my employees machine and what, you know, what kind of attack surfaces is that, you know, laying out for people. Yeah, and you alluded to some of this already, but can you say a little bit more about how that agent governance process is different from what they might already be doing, either for for data governance or or ML ops or observability or things like that?

Yeah, um it's The the funny thing is like there's been a lot with like AI governance and prompt level governance and and all of this and AI gateways and all this, but but where things start to explode in people's faces is that say you you know we've got an enterprise open AI contract, we know that the data that, you know, uh is being sent to open AI is private, at at least that's what they attest. Um but when that comes back with a with a Like chat GPT or something like this, that that comes back as text or something that that that the user interacts with.

That's fine. That's like, okay, cool. We can, we can, we know that that flow is is probably safe, right? Now, did the person put, you know, my customer's username and password and and you know, address and social security number in it?

That could be problematic. But if we've got some level of, you know, okay, are we using a private LLM? That that kind of AI governance is good. The problem.

though with that agentic governance looks to solve is basically that model is telling an application what to do on that machine. That is a whole level of of you know kind of complexity, right? That is very difficult for any of the tools that are currently sitting on on people's, you know, machines to to actually intercept before something bad happens, right? And uh that's that's where we're seeing You know, data exfiltration.

We just saw a supply chain attack in NPM where like people are just have clawed code, you know, downloading stuff onto their machine that they've downloaded 9,000 times and they didn't check that, whoa, you shouldn't download that today. And how what tool in in anybody's tool shed right now solves that if you're not specifically, you know, focused on agenda governance? Mm-hmm. And how do you solve it?

Given that these are non deterministic tools, there is this agent interaction problem that's out there. So we believe that when it comes down to it, a policy, whether it's a HR policy or a spending policy, whatever it is, there becomes a a level of binariness to it, right? So can you or can't you do that? And that's how we see policies need to work.

And so abstractly, our goal is to find in every policy category that we work with, you set rules. And then our goal and and how we've structured our systems is to find the deterministic invariance of that particular thing that's happening. So uh Cloud Code has tools, for example, right? So if it's gonna do a, you know, if it's gonna run a a bash script, you know, and make an API call or something, it's gonna call that particular tool.

And then that tool is going to have inputs. And so you can intercept all of that, right? And so now you can evaluate those. And you can say yes or no.

And you have to do it fast, or the user experience sucks. And then, you know, people start to um, you know, try to bypass governance, which is the age-old, you know, that's what everybody tries to do in every other type of governance. So we need to make sure that that doesn't happen. And so, so it's rules and then find the invariants of of something that happens.

These are they feel like magic and they feel like black boxes, but they're not, right? There, there are things that that you know that you can. And then once you find those seams and those surfaces, now you can start to inject governance at different parts, you know, before, during, and after, right? You can, you know, uh inject kind of long-term, you know, budget or long-term bias or or these kinds of you know, governance because you have all of that data.

Um, and so yeah, it's it that's to me, that's the most fun part of of looking at any of this is it's like finding how do you how do you, you know, get your Get your finger in the crack as you're as you're, you know, uh trying to scale the wall of governance here. Yeah, what was the most challenging aspect for you to build? Uh Claude Cowork runs on virtual machines and so so do a lot of these. So that's really difficult because they've kind of made it into a you know, this like uh right self contained thing.

makes it more secure at some level, at least more limited. Well, you you think, right? But the problem is it still can, you know, with a virtual machine, you can still mount a a file starter, right? So it still can go get files from a machine, right?

So it still can do that, and then it can grab that and it can send it to to something. So it's like, you know, it in a lot of ways it's it's more dangerous than something like Claude Code, that clawed code can still, you know, can run bash scripts and stuff, but you you have hooks, you have all this stuff that you can so So with clo cowork, we've actually built tooling for machines on Mac and Windows that that can kind of interact with the different, you know, components of co work on the machine to to you know do that.

But yeah, cowork is the coolest. And what the the funniest part of that is like cowork is used by the least technical people in the company. So cowork could say, I'm gonna do this thing, and it's like all this, you know, gobbly gook, you know, command stuff. And somebody's like, I guess so.

Accept. Right. And then it does it. That should be intercepted.

You've talked a few times about about policies. How much do organizations even understand what their policies are or or more broadly what their risk tolerance is? Because obviously there's some level of trade off here. Yeah.

Uh what's funny, like risk tolerance, uh I think we're you know, there's like the tur the the concept of like marginal gains, right? We're we're like in the time of like maximal gains, right? There there are no trade-offs right now, right? It's like get get policies in place, right?

Get cost policy in place, data access policies in place. You're not gonna make any trade offs. They're gonna run milliseconds slower, but you're gonna feel a lot better, right? I think When we look at risk profiles, that's something that Waxel does, is is each policy category shows we call it blast radius, right?

So if this happens, here's what's going to happen in your company. And we we rate the kind of, you know, here's the things that are generally should be looked at. So as you come in, you you kind of had this like boilerplate, right? Here's here's what good governance looks like in in an agentic fleet.

But then as you have your agents running. Waxel basically says, Hey, this is an ungoverned database. You know, you just made a an ungoverned, uh, raw Postgres query. Don't do that anymore, right?

Here's the policy to to to you know create and you know put it in place. Um, one of the challenges is agents can spawn sub agents. Yeah, and then you get this this fan out. How do you deal with that challenge?

it it is irresponsible to deploy something like that without without having the understanding of how to control that, right? And the way we look at this is um we we we call it agent lineage, right? So you have parent, children, siblings, the you know, generally speaking. Um but then you need a way to ensure that that that kind of identity, you know, lineage Is not lost across that.

And so that's that's really where we use a lot of you know runtime kind of um there's like a cryptographic element to it, right? Where things are stamped, where they're running. So, you know, if it's not in the actual code, it's it's from a runtime environment, right? So you can start to you there needs to be defense and depth around that because that if you don't have that, then you can't fire a kill switch.

Kill switches are the you know the thing that you know, say something just spawned 20 agents and is iteratively spawning agents, you need to be able to kill that whole lineage in one one click of a button, or you're gonna have a big problem. You need to start shutting down cloud resources, all that kind of stuff. So uh lineage and defense in depth on on that, uh, I can say that's one of the unsung heroes of what we've done because that is not a fun thing for any developer to try to, you know, build that.

You know, themselves. Yeah, but as as you say, it's i i it's a really critical thing in this agency world. Um well, so broadly, you know, there there are people that focus in more on that identity piece, that that fundamentally the key challenge is identifying and observing ages, the people who focus more on the security piece uh and other aspects. how how do you look at what's what's really the core of agent governance?

So The problem is like, you know, you've got you've got some walled gardens springing up, right? So you got ServiceNow, you know, and I think they just launched something cool with NVIDIA. I haven't looked into it too much. I think they're calling it control tower, which is awesome.

But ServiceNow is this walled garden. Then you have Agent Force, walled gardens. Then you have Claude putting out, you know, this different stuff. And then you've got 9,000 uh agent frameworks and new ones popping up every day.

And you've got different, you know, vector databases that you can use in all of this. So the problem is um if you if you try to you know go vendor locked, right? Then you you lose all this beauty of what happens in AI, right? Because there's always something that's cool that's happening in AI that you should be, you should have somebody on your team looking at that.

You know, new blog article comes out, you see it on LinkedIn, Twitter, whatever, like go install it. Let's go, let's build it. So, so then if you've if you Focus on uh, you know, okay, I'm gonna secure things at, you know, for what's in this moment of time. And I I built this like security for this moment of time.

And that's only for this set or that or that. I think you're not you're not future-proofing the infrastructure. So the way we see it is like have the overlay, have everything plugged in, make it easy to plug in. And then over time, what we're what we're doing is partnering with the security companies, partnering with these, you know, different companies.

Because you have to have the overlay that everything can can bubble up into so that you have that single pane of glass that, okay, I have these agents over here, these agents over here. And then, you know, my dev team is about to push a whole new class of agents that I gotta like see what was happening and then build the policies before they, you know, go out the door. So I think if you're building a a a kind of point solution right now, it's kind of that whole like you can see a uh what's it like you can see an electron either where it is or how fast it's moving, but you can't see both at the same time.

I think that's what we've got going on in AI, right? So we're trying to say, Go build the entire, you know, um the capsule that all my stuff is gonna run in. And then you can start doing cool stuff with it, you know? What about the shadow agent problem where organizations don't even realize what someone's deploying?

Yeah. Uh so I I think that's that's like that that is the uh craziest problem right now because it's like on one hand, somebody finds this cool thing that the company doesn't know about. That should be like you should they should get a cake, right? And we should all clap because that's awesome.

Thank you for trying to make my company better, right? And then on the other side it's like, whoa, whoa, whoa, but what is happening with this, right? Um And so what what we've built is uh on machine apps that that you know are are looking at the the kind of what's going out from the the network and are they a are you know we have a catalog I think of like fifteen thousand different uh API URLs and we're constantly building that. So it's like, okay, is there API are are there calls coming from this machine that are not, you know, part of the governed flows and and part of the governed agents?

And then that gets brought to an admin's attention and then they can they can, you know, basically turn those turn those off. So so kind of classic API gateway, but at the machine, um at the machine level, not at like the firewall or or network level. Mm-hmm. And you mentioned kill switches.

how is it possible to shut something down once it's in deployment, especially if, you know, there's there's child agents and so forth? Yeah. So the uh the old way of doing this was you just turn off the cloud resource, right? Just turn off the machine.

It's like so so I but but that's problematic, right? Because you you you know, sig kill, everything goes away. So your audit trail is gone, all the cost data is gone, everything is gone, right? ah and so the way we see it is uh, you know, this is where like Protecting state, like agent state at all times, right, is the most important thing.

So if the agent's progressing, it's creating, you know, it might be accessing different data, you know, uh resources. It might be, you know, what we talked about generating cost, all of this. There might be intermediate work products that are still salvageable, right? So it might have written a blog article but didn't get to the uh, you know, place that it, you know.

checked all the sources or something, right? Okay, we shouldn't throw that away, right? So what you have to do is you have to to mi uh protect the state of the agent at all costs, which means you have to have you what we call a uh agent envelope, right? That you know everything is kind of hitting that envelope.

It's not hitting the actual functions that are running. Like if you, you know, put the the electron microscope on it. And so so The uh kill switch comes into the envelope, it's then going to the envelope is then going to manage the shutdown of the of the actual agent. Um and the nice thing is uh it once you have the infrastructure and the runtime sorted, this becomes a much less challenging issue, especially if you're you're managing the kind of agent lineage and all of that.

But if you're just running like raw lane chain, you know, or or these different, you know, agent frameworks on like a a uh celery task or something, that kill switch is gonna you're gonna just lose all the data. That's problematic when you think about like, well, I've got a HIPAA auditor. I've got a, you know, all these different things, right? You've got to have that data preserved.

So that's kind of the the kill switch in theory, turn off the machine. In practice, that doesn't fly. You need to have something very graceful. Um That's very protected and and slow.

And going forward, what do you think are the the next challenges, um risks, dangers we're going to see uh with agents? Biggest risk is that people don't trust them and the whole concept slows down greatly. Um I think the the risk of agents, we know what we know what they are, right? Theoretically, right?

Cost and data access, and it could do something. You know, you had mythos breaking out of its harness, right? That would be a good idea to just uh send a sick kill to the process on the machine, right? They like it, there's not that's not real, right?

They just didn't. you know, they just didn't build their harness, right? Um, so it's like I think the the biggest pro the biggest risk with agents is that companies don't get to see the value of the agent because they don't have the the governance uh infrastructure. Every single technology that we've ever put in place, right, has always brought inherent risks.

cloud infrastructure created a a whole attack surface, right? But we adopted it like crazy and and now, you know, it's a pretty it's a pretty secure world. Obviously there's always people trying to break into it. So yeah, I I think that that's a bigger risk than, you know, any one thing that, you know, an agent might be able to do.

Mm-hmm. And then finally, how do you see this landscape of governance providers around agents evolving? You you talked about some the value of a a a general purpose layer versus the walled gardens, but what do you see as the future of the space? I think that there's going to be so much parody.

Like this week I you know, I use Cloud Code a lot. And then uh Cloud Code started, you know, kind of being weird. So I I spun up Codecs for the first time. And so now I'm using Codecs.

So I think there's so much model parity, right? I think there's so much provider parity that a a well constructed organization is going to have the kind of workflows internally to say this model is right for this or this provider is right for this and this provider's right for that. And so that's why I believe that over time, uh it it's kind of crazy to think that a company would commit to a walled garden agentic infrastructure. And so I think that's where and you already see the explosion of observability companies right now, which observability you can't govern what you can't observe.

So I think there's going to be movement. of those types of companies into governance, I hope, right? Because, you know, it'd be nice to have have some, you know, analogs here. And and um so I think the the the general purpose governance layer, I think is what's going to win because that's that's what that's what the the big money in the space is going to kind of ask for, right?

Let me have an opportunity to compete. You see that with OpenAI puts out five point five. Now they're saying Let's get all the cloud code users, right? So they're not gonna want walled gardens all over the place.

Otherwise, you know, they're they're, you know, we go back into this like year-long lockdown and all that kind of stuff. So yeah. General general purpose, I think, is gonna be what wins and then who can plug the most kind of accoutrements into the different, you know, pieces. Great.

Well Logan, thank you so much. We gotta wrap up, but really appreciate your time in the conversation. Thanks, thanks Kevin. Appreciate it.

This has been the Road to Accountable AI. If you like what you're hearing, please give us a good review and check out my Substack for more insights on AI accountability. Thank you for listening. If you want to go deeper on AI governance, trust, and responsibility with me and other distinguished faculty of the world's business school, sign up for the next cohort of Wharton's Strategies for Accountable AI Online Executive Education Program.

Featuring live interaction with faculty. Expert interviews and custom designed asynchronous content. Join fellow business leaders to learn valuable skills you can put to work in your organization. Visit execed.

warden.upen.edu slash acai for full details. I hope to see you there.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • AI Was A Waste of Time, Until It Wasn't with Megan BoshuyzenMaking Sense of Martech · on Claude Code91 / 100
  • How Organizations Can Thrive in the Human + AI Era with David ChestnutThe Edge of Work · on Claude Code85 / 100
  • Small Models, Massive Wins: The New Shopify AI FormulaBeyond The Pilot: Enterprise AI in Action · on Claude Code85 / 100
  • Episode 018: Season 2, the $75 Consult and the Frankenstein StackAI Tools for Practicing Lawyers · on Claude Code84 / 100
  • AI for Engineering Is Leaving the Demo PhaseAI Across The Product Lifecycle Podcast · on Claude Code83 / 100
  • What happens after coding is solved? | Fiona Fung (Manager of the Claude Code and Cowork Teams)Lenny's Podcast · on Claude Code82 / 100

More from The Road to Accountable AI

All episodes →
  • Harish Peri (Okta): When the Thing Accessing Your Systems Has a Brain77 / 100
  • Nadav Cornberg (Eve Security): Interrogating Agents Before They Act83 / 100
  • Venkat Siva (Compfly): Governing Agents at the Execution Boundary95 / 100
  • Munmun De Choudhury (Georgia Tech): Conversational AI and Mental Health83 / 100
  • Emre Kazim (Holistic AI): Why AI Governance is Life Cybersecurity90 / 100
Explore the best B2B AI & Data podcasts →
All The Road to Accountable AI episodes →