The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Engineering & DevTools/M365.FM
M365.FM artwork

The Death of the UI: Why CUA is the End of SaaS as We Know It

M365.FM · 2026-07-01 · 1h 8m

0:00--:--

Key moments - from our scoring

Substance score

38 / 100

Five dimensions, 20 points each

Insight Density13 / 20
Originality14 / 20
Guest Caliber0 / 20
Specificity & Evidence11 / 20
Conversational Craft0 / 20

This episode presents a provocative thesis: the entire SaaS industry's valuation model is collapsing because agents don't need UIs, dashboards, or navigation menus - they need stable APIs, explicit logic, and governance. The speaker traces how the graphical interface, revolutionary in the 1980s, was always a workaround for human cognitive limitations. Now that agents can reason about data without visual representation, the scaffolding that powered $300 billion in software equity becomes dead weight. The core economic crisis: when one agent replaces 50 human seats, traditional seat-based pricing generates an 80% revenue loss per customer. Mid-market SaaS is getting hit hardest because it lacks the margin cushion of enterprise software to absorb the transition to consumption-based or outcome-based pricing models. The speaker advocates for API-first architecture, treating agents as first-class identities with governance and audit trails, and crucially, protecting workflow capital - the accumulated operational logic that differentiates companies. The real competitive advantage isn't the software; it's how you use it. Organizations that encode their specific risk models, approval hierarchies, and decision heuristics into private agents will compete on defensible moats. Those sharing workflow data with vendors to improve generic agents are commoditizing their competitive advantage.

Key takeaways

  • →Seat-based SaaS pricing is structurally incompatible with agent workflows - one agent replacing 50 employees means an 80% revenue loss per customer, causing a market repricing from 12x to 4x revenue multiples.
  • →Agents require API-first architecture with atomized, callable operations, not UIs; 82% of enterprises claim API-first strategy but only 25% actually implement it, leaving them unable to deploy agents at scale.
  • →Workflow capital - the accumulated operational logic, decision criteria, and institutional knowledge embedded in how organizations work - is the real competitive moat, not the software itself.
  • →Computer using agents (CUA) that read screens succeed only 67-85% on simple tasks and 9-19% on complex workflows, making API-first design a necessity rather than an option.
  • →Organizations must treat agents as first-class identities requiring provisioning, permission management, governance, audit trails, and approval workflows, similar to but distinct from human users and service accounts.

Topics in this episode

Consumption-based pricingSeat-based pricing modelsAPI first architectureFoundry Agent ServiceService accountsComputer Using Agents (CUA)Workflow capitalSaaS valuation compressionEnterprise identity managementNon-human identity management

Questions this episode answers

Why is the graphical user interface becoming irrelevant for enterprise software?

Agents process information through direct data structures and logic without needing visual representation, making the UI scaffolding that exists solely to bridge the gap between human cognition and machine logic unnecessary and counterproductive.

How does agent adoption destroy SaaS revenue models?

When one agent handles the work of 50 employees, seat-based pricing generates an 80% revenue reduction per customer because the unit of value shifts from headcount to work completed, not the same metric.

What is workflow capital and why does it matter?

Workflow capital is the accumulated operational logic - approval chains, risk thresholds, decision criteria, and institutional knowledge - that differentiates how organizations actually work; agents trained on your specific workflow capital are defensible competitive advantages that competitors cannot replicate.

Why do most enterprises claim API-first strategy but fail to implement it?

Enterprise software was built around human interaction for decades, making UI-first design the embedded practice; most companies bolt APIs onto UI-first systems rather than rebuilding with API-first architecture as the primary contract.

What governance do agents require that differs from user identity management?

Agents need provisioning, permission management, audit trails, and approval workflows like users, but also require governance frameworks that didn't exist for traditional service accounts, including clear boundaries on allowed actions and accountability for agent-made decisions.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

13 / 20

The episode contains substantive claims about architectural shifts (API-first vs UI-first, seat-based pricing collapse, workflow capital as defensibility), but relies heavily on abstract frameworks without sufficient concrete operational specifics. Many assertions lack supporting numbers or real-world examples that would help an operator immediately apply the lessons.

An agent that replaces 50 human seats, the vendor's revenue collapses
Workflow Capital is the accumulated operational logic that makes your company work

Originality

14 / 20

The episode reframes enterprise software economics around agent adoption and workflow capital as a moat, which is relatively fresh. However, the core arguments (UI as legacy, API-first as necessity, pricing model compression) are increasingly common takes in 2025-2026. The 'workflow capital' framing is novel but underdeveloped; much of the architectural discussion follows predictable patterns.

The dashboard was scaffolding. So is the form... the moment you remove that constraint, the scaffolding becomes dead weight
Workflow Capital is not software. It is not a feature set and it is not something you can buy

Guest Caliber

0 / 20

This is a solo monologue, not a podcast interview. There is no guest. The speaker delivers prepared remarks without any interlocutor, which fundamentally disqualifies it from guest-based evaluation.

For two decades, we built enterprise software around a single assumption

Specificity & Evidence

11 / 20

The episode provides some concrete numbers (82% vs 25% API adoption gap, CUA success rates of 67-85% vs 9-19%, SaaS multiple compression from 12x to 4x, labor costs $120-150k per employee), but lacks named customer examples, specific vendor implementations, or real deployment timelines. Many high-stakes claims about agent-driven ROI remain illustrative rather than evidence-based.

CUA succeeds on simple UI tasks, about 67% to 85% of the time. On complex multi-step workflows, it drops to 9 to 19%
SaaS multiples have compressed from 12 times revenue down to four times revenue

Conversational Craft

0 / 20

There is no conversation. This is a continuous solo presentation with no host questions, follow-ups, challenges, or dialogue. Without a conversational partner, this dimension is not applicable.

This is the operational model entirely

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

agent229agents114data64model52software45build44human35specific34built31workflow31logic31aren30real29first29platform29team28

Episode notes

For more than forty years, enterprise software has been built around one fundamental assumption: humans need graphical interfaces to interact with machines. Dashboards, forms, navigation menus, search boxes, workflow builders, and endless clicks became the foundation of the software industry. But what happens when the user is no longer human? In this episode, we explore one of the most disruptive shifts in technology since the rise of cloud computing: the transition from human-driven software to agent-driven systems. As Computer-Using Agents (CUA), autonomous AI agents, and API-first architectures become mainstream, the traditional SaaS model faces an existential challenge. We examine why user interfaces were always a workaround for human limitations, how agents interact with software differently, and why the economics of seat-based software licensing are beginning to break down. More importantly, we explore what replaces the UI and how organizations must rethink architecture, governance, security, identity, workflows, and business value in a world where agents increasingly perform the work once done by people. This conversation goes far beyond AI hype.

Full transcript

1h 8m

Transcribed and scored by The B2B Podcast Index.

For two decades, we built enterprise software around a single assumption. Humans need graphical interfaces to navigate complexity. Buttons, forms, dashboards, navigation trees. That entire apparatus existed because the human brain cannot reason about data at scale without visual representation.

It was necessary. It was real. That assumption is collapsing. Computer, using agents don't need dashboards.

They don't need navigation menus or search filters or workflow builders. They need something different, logic, context, and stable APIs. They need to understand what they're looking at, decide what to do, and execute. No clicks required.

And right now, your entire SAS valuation model is compressing because of it. The seat-based pricing model, where you multiply headcount by a monthly fee to calculate recurring revenue, is structurally incompatible with agent workflows. When one agent replaces 50 human seats, the vendor's revenue collapses. When that happens across the entire software industry, $300 billion in equity value evaporates, not because software is broken.

But because the operating model underneath has shifted, this isn't about software disappearing. This is about the assumption that broke the moment agents became economically viable. By the end of this conversation, you'll understand why your current software investments are becoming bottlenecks instead of assets. You'll see where the real mode lives now, and you'll understand exactly what has to change if you want to compete in the next decade.

The UI was always a workaround. The graphical interface solved a real problem in the 1980s. Before that, you interfaced with computers through command lines and punch cards. You had to translate your intent into machine syntax and memorize commands, which meant the cognitive load was immense when the GUI arrived.

The desktop, the window, the mouse pointer. It was revolutionary because it compressed the gap between what you wanted to do and what you had to tell the computer to do. But it was always a workaround. It was a workaround for the fact that humans and machines reason about information differently.

Humans think visually. We see a button and know it's clickable. We see a form and nowhere to type. We scan a dashboard and extract the pattern.

Machines don't think that way. They process symbols and logic. Every SAS product built since then has been constructed on the same architecture. Database plus business logic plus a UI layer designed specifically for human interaction.

Salesforce didn't invent a new paradigm. It took the three-tier architecture and made it cloud native. Same pattern. Humans click, machines respond.

Humans interpret the response. That cycle repeats. The problem is that cycle has become the constraint, not the solution. When a human uses software, they're translating intent into clicks.

You type in the search box, click the filter button, and wait for the page to load so you can scan the results. If you don't find what you need, you try again. That translation loop intent to interface, interface to action, action to result, result to interpretation is where latency lives. It's where errors accumulate.

It's where cost hides. An agent doesn't operate that way. An agent reads data directly. It understands the structure without visual representation.

It operates on the logic, not the interface. It doesn't wait. It doesn't misinterpret what it's looking at because the layout confused it. There's no translation loop.

The dashboard was scaffolding. So is the form. The workflow builder, the drag and drop interface, and the admin panel were all tools designed to make software usable for humans who couldn't code. They worked because they mapped the underlying logic into visual metaphors that the human brain could grasp.

But that scaffolding only exists because of a constraint, the human user. The moment you remove that constraint, the moment you have a system that can reason about data without visual representation. The scaffolding becomes dead weight. It slows things down.

It adds latency, where there should be none. It introduces opacity, where there should be clarity. This is why the entire software industry is being forced to rebuild, not because software is broken, but because for the first time, there's a viable alternative to the human-driven UI model. And that alternative doesn't need any of the visual apparatus we've spent decades perfecting.

Understanding why UI's exist is the first step to understanding why they're about to become irrelevant. The seat-based model is mathematically broken. The math has been simple for 20 years. You count the people using your software.

You multiply that by a monthly fee. You add it all up for the year. That is your annual recurring revenue. For two decades, we build business models around predictable growth in that one number.

You hire more sales reps. You build features that force people to add more users. The more headcount your customer has, the more they pay you. It is clean.

It is predictable. Wall Street understands it. That formula worked because of one stable assumption. We assumed headcount growth was tied to business expansion.

When a company hires new employees, they buy more licenses. When they open a department, they add more seats. The variable you price against, the number of humans, moves in the same direction as the organization. You can forecast it.

You can model it. You can build a repeatable sales machine around it. Then agents show up. One agent handling invoice reconciliation can process what used to take a team of five people.

That is not a projection. That is what we are seeing in pilot programs right now. One system. Five seats of value destroyed.

The company that bought your software no longer needs five people for that function. They need one person to oversee the agent and one person to handle exceptions. Now calculate what happens to your revenue. The customer still needs your software, but they need four fewer licenses.

That is an 80% reduction in what they pay you. If you have 100 customers and this pattern hits even half of them, your revenue collapses by 40%. Your retention metrics crater, your expansion revenue evaporates, your growth curve inverts. This is why traditional SaaS companies are trapped.

If they do not adopt agents and integrate them into their platform, they lose market share to competitors who do. A vendor that builds agente capabilities into their product becomes more valuable than one that does not. Customers migrate. Market consolidation happens fast.

The laggards get acquired or they simply disappear. But if they do adopt agents, they cannibalize their own revenue model. You cannot price a system the same way when the unit of value has shifted from the number of humans to the amount of work completed. Those are not the same thing.

A company with 50 employees might process 10,000 invoices a month. With agents, they might still process 10,000 invoices, but with 30 employees. Same output, 30% of the license revenue is gone. The market understands this dynamic and it is responding with brutal efficiency.

SaaS multiples have compressed from 12 times revenue down to four times revenue across the board. This is not because individual products are performing worse or because customer satisfaction has declined. It is happening because investors are reprising the entire category based on the structural incompatibility between seat-based pricing and agente workflows. If you are a SaaS vendor and your revenue model is still based on headcount, your valuation is deteriorating in real time.

Mid-market SaaS is getting hit hardest. Enterprise vendors have enough margin to absorb transition costs. They can experiment with usage-based pricing. They can layer on consumption models.

They have enough cash flow to survive the reprising. But mid-market companies selling into mid-size organizations at 10,000 to $100,000 annual contract values do not have that cushion. They need growth to be clean and predictable. They cannot afford the margin compression that comes with agente integration.

The vendors moving fastest are scrambling toward new models. They are looking at consumption-based pricing, outcome-based pricing and hybrid models that blend seats with usage. But these require completely different operational infrastructure. You need metering systems.

You need billing that can handle variable charges. You need finance teams that can forecast revenue when it is no longer tied to a simple headcount multiple. Most of these systems do not exist yet. The reprising is not a temporary correction.

It is a structural shift. The market has decided that the old model is obsolete. What is being priced in right now is the assumption that SaaS vendors will successfully transition to agente architectures and new pricing models. But that transition is harder than anyone anticipated and the cost of failure is extinction.

This is why the economic pressure is forcing a fundamental shift in how software gets built in price. That shift starts with understanding what agents actually require. What agents actually require? So what does an agent actually need to do its job?

It is not what you think. Your instinct is probably to say it needs a good UI or a clever interface. That is wrong. An agent does not need a dashboard.

It does not care what color the buttons are. It needs something completely different. Stable, predictable APIs. It needs explicit logic.

It needs clear boundaries on what it is allowed to do and what it is forbidden from doing. A dashboard is a visualization layer. It exists because humans need to see information to process it. An agent processes information without visualization.

It can read data structures directly. It understands what a JSON object represents without needing a table or a chart or a graph. Give it a field called invoice amount and it knows it is dealing with a number. Give it a field called approval status and it understands the possible states without a legend.

This distinction matters because it changes what usability means. For a human usability means the interface is intuitive. The controls are discoverable. The information architecture makes sense.

For an agent usability means the API is stable. The scheme is consistent and the business logic is exposed as calbable operations. When we talk about computer using agents, we are talking about systems that can read a screen, interpret the visual layout and execute actions the way a human would. That is, CUA, it is powerful because it works with legacy systems that do not have APIs.

But it is also a workaround. CUA succeeds on simple UI tasks, about 67% to 85% of the time. On complex multi-step workflows, it drops to 9 to 19%. That is not because the agent is not intelligent enough.

It is because the agent is trying to infer intent from visual design instead of operating on explicit logic. CUA is friction masquerading as a solution. The real requirement is API first architecture. Every piece of business logic needs to be exposed as a calbable service.

It cannot be hidden behind the UI. It cannot be locked in a workflow builder. It cannot be embedded in a proprietary data format. It must be exposed as an API that an agent can invoke, understand and chain together with other operations.

But here is where it gets harder. An agent does not just need access to operations. It needs context. It needs to understand not just the data, but the rules, the constraints, the organizational knowledge that determines what actions are valid.

When a financial agent is processing an expense report, it cannot just check if the amount is under $10,000. It needs to know that expenses over $5,000 in the EMEA region require a different approval chain than expenses under $5,000. It needs to know that travel expenses get flagged differently than software subscriptions. It needs to know that the Requestors Department has a specific budget that is already 60% consumed.

That is context. That is institutional logic. That is the difference between a generic agent and one that actually understands your business. An agent also needs governance.

It needs clear boundaries on what a canon cannot do. It needs an audit trail for every decision. If an agent approves an expense, you need to know which agent made that decision. What data it was operating on, what policy it applied, and when.

This is not for data analytics. It is for accountability. Because when something goes wrong and it will, you need to understand what happened and why the agent did it. Most critically, an agent needs to be treated as a first class identity in your system.

It is not a script. It is not a background process. It is an entity with its own identity, its own permissions, and its own life cycle. It needs to be provisioned like a user.

It needs to have its access managed. It needs to be deprovisioned when it is no longer needed. It needs audit logs that show what it did and when. It needs approval workflows that confirm it is authorized to perform sensitive operations.

This is why Foundry Agent Service exists. It is not just a runtime. It is an orchestration layer that enforces these requirements at the platform level. Context management, governance, identity, observability.

These are not nice to have features. They are foundational requirements. Once you understand what agents actually need, you start seeing how the current enterprise software stack is fundamentally misaligned with that reality. The API first imperative.

Here is the uncomfortable truth about enterprise architecture today. 82% of organizations claim they have an API first strategy. They say it in their CTO presentations. They put it in their modernization roadmaps.

And they have likely spent millions on consultants who told them this approach is non-negotiable. But only 25% actually operate that way. The gap between what companies claim and what they actually do is where the real bottleneck lives. The reason is historical.

We built software around human interaction for so long that the practice became embedded in how we think about architecture. When you design a system with a UI first mindset, the interface becomes the primary contract. The database lives at the bottom. The business logic lives in the middle.

And then you eventually realize you need an API. So you build one, you expose whatever the UI needs, and make sure the API can support the same operations, the dashboard supports. You call it done. You think you are API first.

But in reality, you are UI first with an API bolted on. The difference matters because it changes everything about how an agent works with your system. When the UI is primary and the API is secondary, the API is shaped by what humans find convenient rather than what machines need. You get thick endpoints that combine multiple operations and you get response formats that include visual hints and metadata that make sense in a table, but add noise in a service call.

You get rate limits and pagination designed for human browsing instead of machine automation. But when you build API first, the UI becomes the afterthought. It is just another consumer. The API is designed around discreet, calibable operations where the business logic is exposed as atomic services.

A tool is a tool. An operation is an operation. The UI consumes those operations and arranges them visually for humans, while another system, an agent or a workflow engine, consumes the same operations and chains them together logically. For agents to work at scale, this is the requirement.

Every business capability needs to be atomized into discreet, calibable services. They cannot be hidden behind a dashboard or embedded in a workflow builder. They must be exposed as operations that an agent can invoke, understand, and chain. This is not just a technical change.

It is an organizational change and it is harder than most enterprises realize. Your CRM is no longer just a system for humans to manage customer relationships. It is a service that agents can query for customer data, update with interaction history, and orchestrate across your entire workflow. Your ERP is not just a place where financial data lives, but an API that agents can call to make decisions, validate transactions, and execute them with full audit trails.

The enterprises moving fastest are not the ones with the most advanced AI. They are the ones that already invested in deep API architecture. They did not do it because they were planning for agents, but because they needed APIs for mobile apps or integrations or external partners. Now those APIs exist and they can plug agents into existing services without rebuilding the entire stack, companies that are still UI first are facing a choice.

Modernize your architecture now or watch your software become a constraint on what agents can do. The cost of waiting is accelerating. Every month without API first architecture is a month your competitors are gaining density in their agent deployments. They can build new agents faster because the underlying services are already exposed and they can iterate more quickly because they are not fighting the friction of UI first design.

They can scale automation to more use cases because the architecture supports it. This is why API first has become a competitive necessity rather than a preference. It is not about being trendy. It is because agents require it and agents are becoming the operating model of enterprise software.

The next question is, once you have built the architecture who actually owns the advantage, the answer is simpler than you would think, but it requires a new way of thinking about what makes a company defensible. Workflow Capital is your real mode. There is a concept that has become the defining strategic insight for enterprises thinking about what actually matters anymore. I've only introduced it and once you understand it, you cannot unsee how everything in enterprise software is being reorganized around this single idea.

It is called Workflow Capital. Workflow Capital is not software. It is not a feature set and it is not something you can buy. It is the accumulated operational logic that makes your company work.

It is the unwritten rules and the edge cases everyone knows about but nobody documented. It includes the decision criteria that separates a good approval from a bad one, the approval chains and the sequences of who talks to whom and in what order. It is the informal workarounds that actually make the formal process function and the judgment calls that experienced employees make without even thinking about it. It is how you use the software, not the software itself.

Here is where it gets concrete, take two banks. Both run the same CRM with the same vendor, the same features and the same interface, but they have completely different workflows because they have different risk tolerances, different customer segments and different regulatory constraints. One bank might approve a commercial credit line at the branch level if the amount is under $100,000 while the other requires regional approval for anything over $50,000. One bank flags international transactions for enhanced due diligence immediately but the other has a more granular risk model that factors in the customer history and counterparty reputation.

One bank routes small business lending through a different approval chain than commercial lending while the other uses the same process for everything. Both banks are running the same software but their workflow capital is completely different and that difference is what actually drives their competitive advantage. It is not the CRM, it is the way they use it. In a WorldWare agent's execute workflows, that shifts everything.

An agent that understands your specific risk model, your approval hierarchy, your exception handling and your decision heuristics is worth infinitely more than a generic agent because the agent is not just executing, it is embodying your logic, it is making decisions the way your organization makes decisions, handling exceptions the way you handle them and escalating the right things to the right people. It is following your playbook instead of some vendor template but here is where it gets dangerous.

If you let vendors train on your workflow data to make their agents better, you are teaching them how to replicate your competitive advantage, you are handing them the playbook. The vendor sees your approval patterns, your risk thresholds and your decision criteria. They use that to train their models, they make their agents smarter and then they sell that intelligence to your competitors. You have just commoditized your secret source.

This is happening right now. Vendors are asking customers for access to their real work artifacts like email, documents, meeting transcripts and transaction logs. They want this data to train agents that work better in real office environments. What they are actually doing is learning how your organization works so they can productize that knowledge and sell it to everyone else.

This is why organizations need to be deliberate about what workflow data they share with external vendors. The enterprises that will win in the agentic era are not the ones buying generic SaaS tools. They are the ones encoding their workflow capital into private or controlled agents. These agents reflect how they actually do business.

They embody institutional knowledge and competitors cannot replicate them because the logic is proprietary, the context is internal and the training data never left the company. This requires a fundamental shift. We have to move from buying software to building agents that embody our logic. We have to move from configuring the vendor workflow to coding our workflow into our agents.

We have to stop treating workflow as configuration and start treating it as intellectual property that needs to be protected and compounded over time. It is a different operating model entirely. Most software companies today optimize for feature parity. They ask what their competitors offer and they build that plus a little more.

It is an arms race of functionality. But in an agentic world, that race becomes irrelevant. The features are fast to copy, but the workflows are not. An agent trained on generic best practices performs worse than an agent trained on your specific logic.

An agent that embodies your risk model, your customer segmentation and your operational heuristics is defensible. That is durable. That is your mode. The organizations that understand this are already moving.

They are inventorying their workflow capital and asking what makes them unique. They are looking for the operational logic they have that competitors do not, where their decisions are different and what their experience people know that they could encode. They are treating that as their most valuable asset. Protecting and compounding workflow capital requires a governance model that most enterprises do not have yet.

The non-human identity problem. For decades, enterprise identity management was a solved problem. You have employees, you provision them when they joined, you manage their permissions as they move around. You deprovision them when they leave platforms like Entry, D and Octa were built on this one assumption.

Everything flows from the fact that you are managing humans. Humans have names and email addresses. They belong to departments and report to managers. They have tenure.

Because humans are predictable, you can write policies around their behavior. But agents break this model completely, an agent isn't a user. It doesn't have an email address or a performance review. But it's also not a service account.

In the old model, a service account was the answer for non-human access. You'd create one account and give it a password. It lived forever. It had broad permissions.

So it could do everything it might ever need. Multiple applications or scripts would share that same account. It was convenient and simple, but at scale, it's a security nightmare. The problem is that service accounts are too rigid and too privileged.

A service account created in 2015 might still have access to systems that don't even exist anymore. It might hold permissions for operations it hasn't performed in years. If that account is compromised, an attacker has access to everything. And they have all the time in the world to figure out what that is.

There is no audit trail of which application used it or what specific operation happened. It's just a blob of standing privilege sitting in your infrastructure. An agent is different. An agent needs to perform a specific operation for a specific period of time.

It doesn't need broad permissions. It needs minimal permissions just enough to do the job and only while it's doing it. This is the non-human identity problem. How do you govern an entity that acts on its own and makes its own decisions?

The traditional answer of giving it a service account and hoping for the best doesn't work. You end up with hundreds of accounts with no clear owner and no audit trail. Microsoft's answer is "entraagent id". Each agent gets a unique identity just like a human user.

That identity can be provisioned and governed using the same frameworks you already use. You can assign it to groups or apply conditional access policies. You can monitor it for suspicious activity and revoke access the moment the agent is done. But here's the critical difference.

Agents should operate on the principle of zero standing privilege. This means an agent only has access to the specific resources it needs for the task it's doing right now. Not everything it might ever need. Not everything that would be convenient.

Just what it's doing in this moment. After the task is complete, the access should be revoked. This requires just-in-time access and continuous monitoring. You need the ability to say that an agent can read customer data from a specific database, but only if a specific business process invokes it.

And only for the next 30 minutes. Most enterprises don't have this level of granularity in their systems. They were built to assign permissions to a role that stays active all day. You need granularity measured in minutes and tied to specific data contexts.

This is not how traditional platforms work. It's not how your current infrastructure was designed. They're going to have to build it. And that's where agent 365 comes in.

Where agent 365 is the control plane. Agent 365 reached general availability in May of 2026. It's the first enterprise product designed to solve one problem. How do you govern autonomous agents at an organizational scale?

It's not a tool for building agents. It's a tool for governing them. Think of it this way. Agents are emerging everywhere right now.

You have copilot studio agents, foundry agents, and custom ones built by your engineering teams. Some are authorized, some are experiments, and some were built six months ago and forgotten. Without visibility into what's running, you have a governance crisis. In agent 365, every agent in your tenant gets registered.

They are classified and monitored. You can see who owns them, what systems they touch, and how often they run. This matters because without that visibility, you end up with shadow agents across your infrastructure. An engineer might spin up a custom agent to automate their workflow without telling anyone.

Suddenly, an unmanaged entity is querying customer data in your CRM. You have no idea it exists because nobody reported it. Agent 365 forces that into the light. Visibility is just the foundation though.

The real power is in the governance layer because each agent has an intra-ID, you can apply conditional access policies like you would for a human. But these policies are measured against what the agent is actually doing. You can set a capability boundary, stating an agent can read customer data but not write it. You can set a temporal boundary, meaning the agent only has access from 9am to 5pm.

You can even set a data classification boundary. Per view, sensitivity labels are enforced at the platform level, not in the agent code. If the agent tries to touch something, it shouldn't. The platform blocks it before the request even executes.

You can also set a control boundary. This means an agent might require human approval before it can execute a financial transaction. It can draft the transaction and validate the data, but then it stops and waits for a person to review it. These policies are enforced as hard stops at the platform level.

They aren't just guardrails or recommendations. The agent cannot do the thing because the platform won't let it. That's the difference between hoping agents behave and actually ensuring they do. Agent 365 integrates with the rest of your security stack.

Per view handles the data governance while defender watches for threat signals. If an agent starts accessing resources, it doesn't normally touch, you get an alert. Then there's the observability piece. You can audit every decision, every tool call, and every interaction.

You can replay a conversation to understand exactly what the agent did and why. When something goes wrong, you can see the data it was using and the policy it applied. This helps you figure out if there's a genuine problem or just a false positive. The licensing model is per user, not per agent.

You aren't paying $100 for every agent you create. Instead, you pay for the governance capability itself. Everyone in your organization gets access to manage agents under a single license pool. This aligns the incentives because you aren't penalized for having hundreds of agents.

You're simply paying for the ability to govern them effectively. The more agents you have, the more critical this infrastructure becomes. The pricing reflects that reality. This is how you move from shadow agents everywhere to governed agents everywhere.

The platform becomes the final word on what's allowed. Foundry agent service as the runtime. Agent 365 handles the governance. It defines the rules and keeps your agents within the lines.

But governance without a place to run is just a set of rules with no game. You need a home for these agents to actually do the work. That home is the Foundry agent service. Foundry is the managed platform from Microsoft, built for running agents at a massive scale.

It is not co-pilot studio. That's a tool for building chatbots that answer questions in teams. It is also not Azure Functions. That's just a place to run snippets of code when something triggers them.

Foundry is different because it was built for one specific purpose. It runs agents that need to manage complex workflows. Remember what happened in previous steps and talked to multiple systems at once. The way it works is actually very simple.

You start by defining your agent. You write the prompt that gives it instructions. You give it tools so it can interact with the world. You connect it to knowledge sources.

So it has the right context. Once that's done, you have an agent. Foundry takes care of the rest of the infrastructure. The tools are the most important part of the setup.

A tool is how your agent reaches out and touches another system. It might be an API call to your CRM, a quick database query, or even a command to move a file. You define what these tools are and what they can do. The agent learns what is available in its toolbox.

When it needs to get a job done, it picks the right tool and uses it. Foundry manages the actual connection, handles the errors if something breaks, and keeps a log of every single move. One of the biggest shifts here is how Foundry handles multiple agents working together. You aren't stuck with one giant agent trying to do everything.

Instead, you can build an orchestrator. This is a lead agent that takes a request and figures out where it needs to go. If the task needs a deep dive into data, it sends it to the retrieval agent. If it needs to check a legal rule, it hands it off to the policy agent.

If it needs to actually move money or update a record, it calls the action agent. Foundry manages the handoffs and makes sure the context stays consistent as the task moves between them. This is how you actually scale an AI strategy. You don't try to make one agent smarter.

You build a team of specialists that are great at one narrow thing and you let an orchestrator coordinate the workflow. The pricing model also shows how much the value of software has changed. You aren't paying for seats anymore. You are paying for what the agent actually does.

Foundry charges for compute by the hour based on the resources your agent containers use. You pay for tokens based on the model you choose. If you use GPT-4, you pay those rates. If you use a smaller, cheaper model, your bill goes down.

Even the tools are meted. If your agent calls an API a thousand times, you pay for a thousand calls. This is a true consumption model. It aligns what the vendor wants with what the customer needs.

In the old world of seat-based pricing, efficiency didn't really matter. Now, if your agent is wasteful, if it uses an expensive model for a tiny task or makes the same API call over and over, you see that cost immediately, this creates a real pressure to optimize. The agent has to be smart. It needs to cache data and pick the right model for the right job.

These aren't just technical goals. They are financial ones. That alignment changes how everyone behaves. Vendors want things to be efficient because it saves them support costs.

Customers want efficiency because it saves them hard cache. Everyone is finally on the same team. This is the future of the operating model. It isn't software that humans have to click through.

It is infrastructure that agents live in. It is managed, it is metered, and it is priced based on the outcome. The runtime is where the work happens, but agents still have to talk to systems that were never meant for them. They have to deal with old software and apps that don't have a modern interface.

That is where the bridge comes in. Computer using agents as the bridge. The reality is that not every system in your company has an API. This is the part of the job that people hate to talk about in meetings.

Your old mainframe doesn't have an API because it was built in the 1970s. Your custom desktop app doesn't have one because nobody ever got around to building it. Even your modern SaaS tools might have limited APIs that don't let you touch the specific workflow you actually need. You can see the data, but you can't pull the lever.

You're stuck. You cannot wait for every system to have a perfect API before you start using agents. If you wait for that, you will be waiting forever. Some of these systems are never going to change.

They are in maintenance mode and it's too expensive to rebuild them just so a machine can talk to them. You need a different way in. You need a way to work with systems that only have a screen meant for humans. This is why we have computer using agents or CUAs.

A CUAs is an AI model that looks at a screen exactly like you do. It sees the buttons. It recognizes the text boxes. It understands how a table is laid out.

It knows that if a message pops up in the corner, it's probably a confirmation of the button it just clicked. It can then take action. It clicks, it types, and it navigates through a process by chaining those visual steps together. It isn't the fastest way to work and it isn't the most elegant.

But it works when nothing else does. This tech actually started as a way to help people with vision loss. The goal was to have an AI describe what was happening on a screen. It turns out that if an AI is smart enough to describe a screen, it is smart enough to use it.

That mix of computer vision and understanding is what makes this automation possible. But there is a catch. Recent research shows that CUA is a bridge, not the final destination. When you give a CUA a simple task, like filling out a single form or reading a table, it does pretty well.

The success rate is usually between 67 and 85%. That is good enough to be useful. But when the task gets complex, when the agent has to remember something from three screens ago or handle an unexpected pop-up, the success rate drops to between 9 and 19%. That is nowhere near good enough for a critical business process.

That gap tells us exactly what CUA is. It's a pattern matcher. It works when things stay the same. It breaks when things get complicated.

The long term goal is still to have APIs for everything. But for right now, CUA stops you from being blocked by your oldest software. This is why Microsoft includes CUA as a tool in their responses API. Your agents in Foundry can call on a CUA whenever they hit a wall.

You don't have to build a special CUA agent as... It's just another tool in the box. When the agent can use an API, it does. When it hits a legacy system with no other way in, it switches to CUA.

Foundry tracks both parts in the same lock. You don't have to fix your entire technical debt before you start using AI. You can start with the systems that are ready right now. Use Foundry to coordinate them.

Use CUA to fill in the hole. You can build one agent that talks to a modern database through an API and then jumps into an old desktop app using CUA to finish the job. As you eventually modernize those old systems, you just swap the CUA tool for a direct API. The workflow stays the same.

You just made it faster and more reliable. This is how you move forward one step at a time instead of waiting for a perfect world. CUA is important because it kills the excuse that you can't use agents yet. You can, you just have to build the bridge.

The path is right there and you can start on it today. But as these agents start moving through your systems, you have to know what they are doing. You need to see where they are going and make sure they aren't making mistakes. That brings us to the problem of observability.

The safety and observability layer. When an agent starts acting on its own, it creates a problem that traditional software never had to deal with. You have to see what it is doing while it is happening and you have to prove exactly what it did after the fact. This is not just monitoring.

Monitoring is basic. It tells you if the power is on, if the system is slow or if the code crashed. But observability for an agent goes much deeper than that. You need to see every single decision the agent made and the specific data that led to that choice.

You need to know why it picked one tool over another. When things go wrong and they will, you have to understand the logic that led to the mistake, not just the fact that an error popped up on a dashboard. Foundry builds this tracing directly into the system. Every time a tool is called the system logs it.

Every time the agent makes a choice, the record is saved. You do not have to build your own logging system because the runtime handles it for you. You can open up a full trace of how the agent thinks and see the exact order of every action it took. You see what data came back from a database and how the agent used that information to reach a conclusion.

This creates a clear chain of cause and effect from the first question to the final result. This is the only way to handle compliance. If an agent approves a loan or moves money around, a regulator is going to ask how that happened. They want a trail of evidence, they want to see which rules were followed and which data points were checked.

You cannot just tell them that the AI made a choice. You have to explain it. Foundry gives you that explanation in a format you can actually use. But seeing what happened is only the start.

The real shift is in the guardrails. An agent should never move money without a human saying yes. It is not a technical limitation. It is because a mistake costs real dollars.

An agent should never delete a file without a confirmation because you cannot undo that action. It should never touch data. It isn't allowed to see. Because that is how secrets get out.

You cannot put these rules in the agent code. If you just put a line in a prompt saying do not delete data, the agent might find a creative way to ignore you. It might follow the words but break the intent. You need the platform to step in and stop it.

Foundry connects directly to your data rules through purview. If an agent tries to grab a file, mark the secret, the platform kills the request before it even starts. The agent never even gets a look at the data. The boundary is enforced at the gate.

Now you have a system where you can see what the agent tried to do and exactly how the platform blocked it. Then you have defender watching for patterns. If an agent suddenly starts asking for resources it never used before, you get an alert. If it starts working 10 times faster than usual or logs in at 3am from a weird location, the system flags it.

These shifts might mean the agent is compromised or they might just mean the model is drifting. Either way you see it happening in real time so you can step in. You also have to test for safety before you go live. The assert framework lets you run scenarios to see how the agent acts in a safe environment.

You can throw weird requests at it or try to trick it with bad data. This matters because agents can surprise you. A model might look at all data and decide that certain customers are always a risk. Even when that logic is outdated, it might find a loophole in your instructions that you never saw coming.

Testing catches those gaps before they turn into a headline. These layers of safety inside are what make agents ready for work. Without them, you are just making the same old mistakes. Only you are doing it faster and at a much higher scale.

The organizational shift. Running agents at scale takes a set of skills that most software teams just do not have yet. This is where most companies get stuck. You need people who treat prompt engineering like a science.

Not a trick they found on the internet. They need to know how to structure instructions so the behavior is the same every single time. They have to know how to fix a prompt when the output is slightly off. They have to tell the difference between an agent that actually understands a task and one that is just guessing based on a pattern.

You also need people who understand data at a very deep level. It is no longer enough to just say who can see a folder. You have to know what specific context an agent needs to finish a job. You have to label data so the agent knows where the boundaries are.

If you don't, the agent might accidentally leak a secret while it is explaining its reasoning to a user. Then there is the business logic. You need people who can take your company rules and turn them into something a machine can follow. This isn't about dragging boxes in a workflow tool.

It is about taking your heuristics and your exception rules and turning them into prompts and policies. Most companies do not have these people on staff. They are going to have to train them or find them and this is not a small project. This is a change to the DNA of the company.

It takes time that most leaders think they don't have. The reason this is so hard is that the old way of building software is dead. You cannot just write a list of requirements, give them to a coder and wait for the app. In the old world, the developer makes the choices and writes the code.

If it is wrong, you fix the code. Agents do not work like that. Agents are trained and guided. You write a prompt, you give it some tools and you test it with real info.

If the model does something weird, you don't rewrite a thousand lines of code. You refine the prompt and try again. You are iterating on behavior. You are watching how the model thinks and making tweaks based on what you see.

This feels more like data science than engineering. You are looking for patterns in the failures and tuning the system based on results. But it isn't pure data science either. You aren't building a model from scratch.

You are taking a massive foundation model and steering it with guardrails and fine tuning. It is a new kind of job that doesn't even have a real name yet. The way companies are organizing to handle this is through a center of excellence. One central team owns the platform.

They manage Foundry and Agent 365. They set the big rules for what is allowed and what is banned. They provide the templates so everyone isn't starting from zero. But the actual business units own the specific agents.

The people in sales or HR are the ones who define the workflows. They are the ones who understand their own processes well enough to teach them to an agent. They are the ones who keep the quality high as the business changes. This creates a balance between control and speed.

You aren't trying to build one giant agent that does everything for everyone. You are building a platform that lets you run hundreds of small specialized agents. Each one is built for one specific job in one specific department. The central team just makes sure all of them are safe and easy to watch.

This is a massive change in how big companies work. Usually the IT department controls every single thing. They make the choices and everyone else waits in line. If you want to change, you put in a ticket and hope for the best.

In the agent model, ownership is spread out. The business teams are the ones actually building. They define what the agent does and they decide if it is working well enough to use. The central IT team sets the fences, but they don't drive the car.

This takes a level of trust that a lot of companies just haven't built yet. It requires very clear rules so people know what they can and cannot do. It requires tools that are easy enough for a business user to use without breaking a policy. The way you organize your team determines what you can build, but it also determines exactly how much risk you are taking on.

Measuring agent value and ROI. Calculating ROI for traditional SaaS is almost insultingly simple. You count your users. You multiply by the license fee.

You compare that to productivity gains and divide by the cost. If you save 500 hours a year and each hour costs $40 in labor, you've generated $20,000 in benefit. If the software costs $5,000, you get a 4-to-1 return. It's a rough calculation, but it works for most tools.

With agents, that calculation explodes in complexity. The problem is the fundamental unit of measurement. It's wrong. An agent doesn't have a cost per user.

It has a cost per execution. That shift matters because the volume of work changes the entire equation. You aren't measuring productivity gains per person anymore. You're measuring productivity gains per task.

You aren't dividing benefit by headcount. You're dividing it by workload. To get a real answer, you have to ask different questions. How many tasks did the agent actually finish?

How much time did it save compared to a human doing it manually? What was the quality of that output? What was the actual cost in compute, tokens, and infrastructure? These metrics have to be tracked every single day.

You can't just check them at deployment and walk away. An agent that works perfectly on day one can drift over six months as models update or business rules evolve. If you aren't watching, you won't see the drift until it becomes a massive problem. The most mature organizations are now using a concept called "Agent work units" to standardize this.

An "Agent work unit" is just a discrete measurable piece of work. For customer service, it's a resolved ticket. For finance, it's a processed expense report. For data teams, it's a query answered.

You define what a unit looks like in your world, then you measure the cost per unit and compare it to the manual cost. If a human resolving a ticket costs $30, and the agent does it for five, you have a 6-to-1 return. But if the manual cost is 20, and the agent costs 18, the return is only 1.1 to 1.

In that case, you probably shouldn't deploy it at all. But there's a subtlety here that most people miss. Not all units are equal. An agent might resolve a ticket, but the quality might be lower than a human's work.

Maybe the customer isn't fully satisfied or the first contact resolution rate drops. You have to adjust your benefit calculation to account for that gap. If the agent hits 90% quality compared to a person, you don't count it as a full unit. You count it as 0.

9 units. Suddenly, the economics look different. It might still be worth it, but now you're making a decision based on reality instead of a guess. Some teams also track human efforts saved as a companion metric.

They ask how much time a person would have spent on this task before the agent existed. Then they measure how much time is actually being saved now. This is more subjective because you're estimating a world that doesn't exist anymore, but it captures the gap between theoretical savings and actual results. If your team is spending 40 hours a week on a task, and the agent brings that down to 20, you've saved 20 hours of labor.

That is your benefit. Your cost is simply what you pay foundry to run the agent. The critical insight is that you have to measure continuously. An agent that is 80% as capable as a human, but costs 10% as much is a great investment.

You should deploy that and use the savings to build something else, but an agent that is 60% as capable and costs 50% as much is a bad investment. Do you either kill that project or only use it in non-critical spots where good enough is actually fine? You also have to define what good looks like for different roles. A document agent needs high quality because people rely on the text.

A rooting agent can be lower quality because a human catches the mistakes later. A financial agent needs near-perfect accuracy. An infobot can be rougher because the user evaluates the answer on the fly. Your metrics framework has to account for these differences.

You aren't measuring performance in a vacuum. You're measuring it against the specific context of the job and what happens downstream when the agent fails. That context is what determines if the performance is acceptable. This is where most companies struggle.

They don't have the infrastructure to measure every day. They don't have clear definitions of success. They don't have a process to pivot when conditions change. Building that measurement capability is hard work.

But it's the only way to know if your agent investment is returning value or if you're just automating costs without seeing the benefit. Ownership and accountability. The moment an agent fails, you find out what ownership actually means. In the old software world, ownership was easy to map out.

The product owner owned the product, the engineer's owned the code, the data team owned the database. Everyone had a defined territory. When a bug appeared, you knew exactly who to call. You could trace a feature failure back through the teams to find the root cause and the person responsible for the fix.

With agents, ownership becomes ambiguous and that ambiguity is dangerous. Who actually owns an agent? Is it the business unit that asked for it? Is it the platform team that built the infrastructure?

Is it the data team that provided the training sets? Is it the specialist who tuned the prompts? All of these people played a part in how the agent behaves but none of them feels like they owned it completely. When something breaks, nobody has the clear responsibility to fix it.

When the quality starts to slip over time, nobody notices because nobody was tasked with monitoring it. This confusion creates friction that slows down your response to every problem. The enterprises moving the fastest are setting up explicit ownership models. A common pattern is emerging where the business unit owns the agent and the outcome.

They are the ones who requested the tool and defined the workflow. Since they are the ones who benefit when it works and suffer when it fails, they own the results. Meanwhile, the platform team owns the infrastructure and the safety rules. They maintain foundry, they enforce agent 365 policies, they make sure the agent stays within the guardrails.

They aren't deciding what the agent does. They are deciding the constraints it operates within. Then you have the data team who owns the training data and its quality. If the agent starts failing because the data is stale or biased, that is a data problem, not an agent problem.

This creates clear lines of accountability. The business unit knows they are responsible. The platform team knows their boundaries. The data team knows their domain, but this clarity requires a level of trust that doesn't always exist between teams.

The business unit has to trust that the platform team won't block a valuable agent just because it doesn't fit a perfect template. Real work is messy and sometimes you need an exception. On the flip side, the platform team has to trust that the business won't deploy a risky agent just to move faster. You build this trust through transparent policies.

When the platform team says no, they explain it through actual risk instead of arbitrary rules. When the business pushes back, the platform team actually listens. Some organizations are now creating a former role called the agent owner. It's similar to a data owner.

This person isn't necessarily the one who wrote the prompt or tuned the parameters. Instead, the agent owner is responsible for making sure the agent is working correctly in production. They ensure it's delivering the value at promised and staying compliant with company policy. For a small agent, this might be a part-time task.

For a critical agent processing thousands of transactions, it's a full-time job. The important thing is that a specific person is accountable for the agent's behavior. Clear ownership sets expectations and gives you a point of escalation when things go wrong. But ownership without the right tools is just a title.

You need the operational framework to make that ownership actually mean something. The operational model. Running agents at scale requires a level of discipline most companies haven't built yet, because the reality is, building an agent is easy. Running it in production is where the work starts, and that is a completely different skill set.

You need a way to deploy these agents consistently. Not a manual process where a single laptop is the source of truth. Not a mess of ad hoc deployments where every team follows their own rules. You need a repeatable, auditable process that puts an agent into production reliably.

Every single time, and once it's there, you have to monitor it. But not just system uptime. You need to know if the agent is actually doing its job. Is it handling the volume you expected?

Is the quality of the answers starting to slip? Are the error rates trending up? You navigate. You monitor.

You adjust. Because business doesn't stop just because you want to make an improvement. You have to be able to update the logic, test the changes, and roll them out without breaking the workflows that people are counting on. And if that new prompt you deployed yesterday starts making weird decisions, you need a way to get back to the old version fast.

But before you even get to production, you need evaluation. Not a gut check where you look at a few outputs and feel good about them. You need a structured systematic way to test against real criteria. When you add new training data or refine a tool, you have to know if it's actually better.

Or if it's worse in ways you haven't noticed yet. You can't guess. You have to test. Foundry handles the runtime and the tool orchestration, but the DevOps side of this is still mostly on you.

The platform doesn't have built-in AB testing. If you want to deploy two versions to different users and see which ones wins, you have to build that yourself or plug in a third-party tool. It's not a click button feature. And there is another layer here because agents aren't static software.

They are continuously learning. If you're fine tuning an agent on real-world data, you are constantly updating the model. Every week you're looking at where it failed, adding those cases to the training set and retraining the system. The agent gets better incrementally.

But now you have to manage the versions. You have to track exactly which weights and parameters are running right now. Some teams use MLOps platforms like MLflow to handle this metadata. Others build custom solutions because their requirements are specific to their context.

Then there is the part that isn't glamorous, but it's critical, human in the loop. Not every agent should be fully autonomous. Some should draft an action and then wait for a human to hit approve. The agent does 80% of the heavy lifting, gathering data and making a recommendation, and the human makes the final call.

But managing those approvals at scale is a massive operational task. If you have 100 agents generating 10,000 requests a day, you can't let those get lost in an inbox. You need a system to root them, track them, and escalate them when they stall. Tools like Power Automate or Logic Apps can sit on top of your infrastructure to handle that flow.

This is where most enterprises are going to struggle, because the model is fundamentally different. In the old model, you shipped software once a quarter and then maintained it. It was static. With agents, deployment is more like running a live service.

You are constantly watching for drift. You are evaluating improvements every day. This requires a different mindset. The companies that are ready for this are the ones already living in a cloud-native world.

If you are used to shipping code multiple times a day and watching metrics obsessively, agents aren't a big jump. But if you are used to long-testing cycles and quarterly releases, you are going to feel the friction. Operational readiness starts before the first agent ever goes live. The Incident Response Model When an agent fails in production, it happens at machine speed.

That is the reality that changes how you respond. Imagine an agent designed to archive old customer records. The logic is simple. Find anything older than 90 days and move it to storage.

But instead, the agent starts deleting them. By the time you see the spike in database activity, thousands of records are gone forever. That isn't just a bug. It's a catastrophe for compliance and customer trust.

Or imagine an expense agent that suddenly starts approving every single requested sees. Within hours, millions of dollars are leaving your accounts. The financial impact is immediate. These aren't what-if stories.

This is what happens when you scale without an Incident Response Plan. The first rule is simple. You need a kill switch. You have to be able to kill an agent in seconds.

Not after an hour of meetings. Not after waiting for a manager to sign off. You need one action that stops the execution immediately. And you need clear policies on what triggers that switch.

If a financial agent hits a 5% error rate, it should probably shut itself down automatically. For a low stakes agent, you might give it more room. But you have to define those rules before the emergency happens. You also need monitoring that catches the slow bleed.

A catastrophic crash is easy to see. A gradual degradation is much more dangerous. You need dashboards that show you what's happening in real time. Volume, error rates, latency.

If the patents deviate even a little bit, someone needs to know. And those alerts shouldn't go to a general Slack channel where they get ignored. They need to reach the people who have the authority to act. In a real Incident, the first 30 seconds are everything.

If the error rate jumps from 1% to 15, the escalation path has to be automatic. Who gets the call? What can they do? That authority has to be baked into the structure.

Then, after the fire is out, you have to figure out why it happened. Was it bad data? A model update that changed the behavior? Did a downstream API change without warning?

You have to be able to reproduce the failure so you can fix the causal chain. Maybe you add validation logic. Maybe you tighten the guardrails. But every failure has to be encoded back into the agent's logic.

So it never happens again. Some mature teams even use chaos engineering. They deliberately inject bad data or adversarial inputs just to see how the agent breaks. They want to find the weaknesses before production does.

This is the divide between the people taking agents seriously and the people who are just experimenting. The serious teams have playbooks. They have the monitoring. They have the decision authority.

They treat an agent failure like any other critical production incident because they know the damage is just as real. The experimenters are just hoping it works and hope is not an operational model. The continuous improvement cycle. Agents are not static artifacts.

You don't just build them once and let them run forever. They are living systems. And they require active refinement to stay useful. An agent that performs perfectly on day one might be a total disaster six months later.

Business conditions shift, data become stale, the rules governing your operations change. The moment you stop improving an agent, it starts becoming a liability. Because while your business moves forward, the agent stays exactly where it was and that's how your competitive advantage erodes. This is why the continuous improvement cycle is non-negotiable.

It starts with evaluation. Not a feeling that the agent is working but structured evaluation against explicit criteria. For a customer service agent, that might mean resolving 90% of tickets on the first contact while maintaining 95% satisfaction. For a financial agent, you might demand 99% accuracy on expense categorization and zero false approvals.

You have to define what good looks like before you start measuring. And you don't just measure once a deployment. You do it continuously. Weekly evaluations are the standard but some organizations are already doing this daily.

Foundry has built in evaluation capabilities but they are just the foundation. You define test cases. You run the agent through them. You check if the output matches the result you expected.

But this only works if you actually understand what to test for. Defining good test cases is much harder than it sounds because you have to anticipate the edge cases. You have to understand how the system fails. And you have to think like an adversary trying to break your own agent.

If you only test the happy path where everything goes smoothly, you will never catch the problems that emerge in the messiness of the real world. That's why many organizations supplement platform tools with human evaluation. You have real people review a sample of what the agent produces. Not every single output because that's too expensive but a representative sample.

Maybe you review 50 approvals a week for an expense agent or 100 tickets for customer service. Humans read the responses and provide the nuance that machines miss. They can see when an agent made a judgment call that is technically correct but actually violates the spirit of the company policy. Yes, human feedback is expensive.

You're paying people to spend their time grading the agent's homework. But it gives you a high quality signal that you just can't get from clean automated test cases. Once you see where the agent is failing, you can actually fix it. The fix depends on what the analysis shows.

Sometimes it's just a prompt change because the agent misunderstood an instruction that was too ambiguous. You clarify the language. You provide better examples. And the agent learns the distinction.

Other times it's a tool change. The agent might be failing because it doesn't have the right information so you give it a new tool to provide that context or maybe it's the training data. You fine tune the agent on specific examples of yes and no until it understands the decision boundary. Sometimes you just need a policy guardrail.

A piece of logic that catches a specific mistake before the agent can commit to it. But before you deploy any of these improvements, you have to test them. This is where A/B testing is vital. You run the improved agent in parallel with the one currently in production.

You send the same requests to both and measure if the new version actually performs better. If it does, you validated the change. So you deploy it. If it doesn't, you just learned that your intuition was wrong.

And that's fine. It just means you iterate further instead of pushing a change that doesn't work. The continuous improvement cycle is where agents become genuinely valuable. Think about it this way.

An agent that starts at 70% of human capability but improves by 5% every month will eventually outperform a person. After 12 months, that 5% monthly improvement compounds into a 60% cumulative gain. Now the agent is substantially better than the human baseline. But an agent that stays at 70% and never improves is just a liability.

It's permanently mediocre. It costs money to run, but it's never good enough to fully trust. The difference between an agent that transforms your business and one that becomes dead weight is simple. It's whether or not you build the systems to keep it getting better.

The investment case. Building agent capabilities requires an upfront investment that most enterprises aren't ready to quantify. You have to build the platform infrastructure. Foundry isn't free.

And neither is agent 365. The governance systems. The monitoring and the integration points between PerView and Entra don't just appear for free. You aren't just licensing software.

You are building operational systems that didn't exist before. And you need to hire or train people with skills that most organizations simply don't have yet. Prompt engineers, data governance experts, business process designers. These people have to encode institutional logic into agent instructions and those roles don't exist in a traditional IT department.

You're either building new capability areas or recruiting from the outside. And both are expensive. You also have to modernize your software architecture to be API first. Legacy systems usually don't have the APIs that agents need to function so you have to build them.

That is engineering work. Often thousands of hours of it. Then you have to establish the processes for development, testing and safety that don't exist yet. These are processes not products.

You have to invent them based on your specific context. This is not cheap. Organizations that are serious about this are investing millions in infrastructure and people before they ever see a return. But the payoff when it finally hits is big enough to justify every cent.

An agent that replaces 30 human seeds can save a company millions of dollars every single year. Those seeds have a fully loaded cost. Salary, benefits, office space, training. A typical knowledge worker costs the company between 120 and 150,000 dollars a year.

30 of those seeds represent a cost of 3.6 to 4.5 million dollars. An agent that does the work of those 30 people running 24/7 probably costs less than 500,000 to operate.

The payoff in labor cost alone is 7 to 1. But an agent that enables a new capability doesn't just cut costs. It creates revenue. A customer service agent that handles interactions humans couldn't do profitably allows you to expand your market.

You can serve smaller customers. You can handle higher volumes. You can open new sales channels. That is growth, not just efficiency.

And when an agent reduces errors, the impact compounds. Fewer mistakes mean lower costs for rework while higher quality leads to better customer retention and premium pricing. The real question is how do you decide which agents to build first? This is where most organizations fail because they try to do too much.

They want the perfect sophisticated agent that handles everything so they spend a year planning. By the time they are ready to build, the window has closed. The answer is actually simpler. Start with high volume, well-defined workflows.

If a process has repetitive steps and clear decision criteria, it's a perfect candidate. Expense processing. Invoice categorization. Document classification.

These are workflows where the logic is explicit, where a person could sit down and write out the rules. If a human can write the rules, an agent can learn them. Workflows that rely on gut feeling or complex judgment are much harder to automate at the start. Sales negotiations or complex disputes involve nuance that is hard to codify.

You'll get to those eventually, but they shouldn't be your first project. You want the easy wins. Build an agent that handles 80% of your basic inquiries. Build an agent that processes 80% of your expense reports.

These are the projects that show a clear return on investment immediately. They build the momentum you need to keep going. Once you've proven the model works in your specific context, you can tackle the harder problems with the whole organization behind you. The investment case is also about your strategic position.

If your competitors are building agents and you aren't, you are going to fall behind. They are automating their labor costs while your stay fixed. They are building capabilities that you won't be able to match if you start three years late. But if you are the one building agents now, you pull ahead.

The fastest moving enterprises are treating agent development as a strategic imperative. They are investing heavily because they know this is how you build a competitive advantage in the next decade. The investment pays off, but only if you actually execute. The strategic choices.

Building agent capabilities requires a series of directional decisions that ripple through your entire organization. None of these choices lock you in forever, but they do shape a trajectory that becomes very difficult to reverse later on. Making these decisions deliberately is always better than making them by default. The first fork in the road is your sourcing model.

Do you build agents in-house using Foundry? Buy pre-built agents from vendors? Or blend both approaches? Building entirely in-house gives you maximum control because you own the prompts and the training process.

You understand exactly how every decision is made, and you can adapt the agent to your specific context without negotiating with a third party. But the downside is the massive investment in speed. You are starting from scratch and need a team that actually knows how to build these systems. Because building agents takes months, the business has often moved and the market has shifted by the time you actually finish.

Buying pre-built agents from vendors is much faster. Someone else has already solved the problem, so you just configure it for your context and deploy. You aren't building from first principles, but the downside is the fit. A vendor's agent is built for a generic use case, and your business probably isn't generic.

Your workflows have specific quirks, and your decision-making has nuances that a pre-built agent simply won't capture. You end up trying to configure something that was built for someone else. And you're always constrained by what the vendor decided to optimize for. Most enterprises are settling on a hybrid model.

They use Foundry as the platform and build custom agents for their specific workflows. You aren't inventing the underlying infrastructure, and you aren't arguing with a vendor about customization options. You're using a platform designed for this and building what makes sense in your own context. This requires less upfront work than pure in-house builds and produces better results than buying off the shelf.

But it does require a commitment to learning the platform and maintaining your own agents. The second choice is your governance structure. Do you centralize agent development in a single team, distributed across business units, or find the middle ground? Centralized teams make governance much easier because one team sets the standards and reviews every single deployment.

Policy is consistent and quality is predictable, but the problem is speed. Every request goes through a central queue. And every business unit has to wait for that team's capacity. The organization can't move fast because you've created a bottleneck.

A business unit, needing a specific agent, might wait months while the central team works on a different priority. Decentralized development is fast because every business unit builds its own agents. There are no bottlenecks and no waiting. Each team deploys what makes sense for their specific workflow.

But this is where governance breaks down. Different teams use different approaches. And some teams give their agents too many privileges. Some teams don't monitor their agents properly.

So quality varies dramatically. You end up with some excellent agents and some that are total operational disasters. The middle ground is a federated model. A central team owns the platform and sets the guardrails, providing templates and tools for everyone else.

Business units own their specific agents within those boundaries. The central team defines what you can and cannot do. And they provide a process for requesting exceptions. Business units build fast because they aren't waiting for central approvals on routine builds.

And quality stays consistent because the boundaries enforce the standards. Third choice. Do you use shared foundation models from vendors or deploy models in your own infrastructure? Shared models are operationally simpler because Microsoft or another provider manages the infrastructure.

You call the APIs and you don't have to run data centers or maintain hardware. The concern here is your data. You are showing the vendor your workflows. And you're indirectly training their models on your data through your usage patterns.

The vendor understands your business better with every single interaction. For some organizations, that's acceptable. For others, it's a deal breaker. Private model deployment means you own the entire infrastructure.

You deploy models in your own environment so no data ever leaves your boundary. You maintain complete control. But the cost is much higher. You're managing GPU clusters and handling updates yourself.

You are responsible for reliability. And you need infrastructure expertise that you might not currently have. That trade-off is usually worth it for sensitive workflows where your logic is your competitive advantage. Many organizations start with shared models for standard workflows and move to private deployment for their most important tasks.

This gives you control where it matters and simplicity where it doesn't. Fourth choice. Focus or breath. Do you narrow down on two or three high-value use cases or try to build agents across your entire operation?

Focus leads to deeper success. You really understand one workflow and build an agent that is genuinely excellent. You prove the model works and build organizational credibility. The risk is limited impact because you're only optimizing one workflow instead of transforming the whole company.

Breath means trying to automate everything at once. You spread your resources across too many initiatives so you get some progress everywhere but excellence nowhere. Nothing ever finishes and nothing shows clear value. The pattern that actually works is focused breath.

You pick one or two areas where agents will have the maximum impact. Build them excellently and then expand to adjacent areas once you've proven success. These four choices compound into your overall strategy. They determine your speed, your control, and your risk.

They are worth thinking through deliberately rather than letting the defaults choose for you. The execution roadmap. The transition from decision to action is where most strategic initiatives collapse. You've accepted that agents are necessary and you understand the architecture.

You've even allocated the budget. Now you need a sequence of execution that produces visible winds while building toward a sustainable transformation. The organization is moving the fastest. Follow a three phase roadmap.

Though the timeline always varies based on your starting position and how much change your organization can handle. Phase one is the foundation. This usually runs three to six months. You start with an honest assessment of where you actually are today.

What APIs exist? How mature is your data governance? Where are the biggest operational bottlenecks? You aren't trying to fix everything at once.

You're looking for the specific place where an agent delivers immediate value with a reasonable amount of complexity. Parallel to that assessment. You have to build the platform. This means implementing foundry and configuring agent 365 to provide visibility into what your agents are doing.

You need to establish governance policies before you actually need them. It's tempting to skip this and worry about governance after your first agent is built. But don't do that. An ungoverned agent in production is an operational liability.

Get the infrastructure in place first then build the agents within it. The pilot is where you prove the model actually works in your specific context. This isn't a proof of concept that everyone forgets about. It's a live pilot where the agent executes real work and produces real results.

Pick something high value but contained. Like a specific document type, the agent can categorize or a subset of customer inquiries. If a workflow currently consumes 20 hours a week, that's a perfect candidate for an agent to accelerate, deploy the agent to limited traffic and let it run for about a month or two. Measure everything systematically.

Is it faster? What is the quality? This pilot is your truth telling mechanism and it surfaces the problems before you've committed massive resources. Phase two is scale.

This runs six to 12 months. You've proven the model works and now you're building momentum. You're creating the operational muscle memory that makes building agents feel routine instead of heroic. Establish templates and best practices so building the next agent is faster and more consistent.

Don't reinvent the prompt structure or rethink the governance model every single time. Use what worked and adapted. You should be compounding your learning, not starting over from scratch. Scale the governance infrastructure to handle the increased volume.

You had one agent in the pilot but soon you'll have 50. You're monitoring dashboard and your evaluation processes need to handle that throughput. You're not hiring one person per agent. You're building systems that let one person govern many agents at once.

Train your business units to build their own agents. You aren't trying to make every employee in engineer. But the operational leaders who understand where the friction lives need to know how to collaborate on agent design. They don't need to write the prompts themselves.

They just need to translate their workflows into agent requirements. That is a learnable skill. Phase three is transform. This is the phase that never really ends.

Agents stop being a project and become how the organization actually operates. You're building agents for increasingly complex workflows because you've finally learned how. You aren't waiting for central capacity or debating whether agents are worthwhile that conversation is over. Now you're just debating which workflow to automate next.

Start thinking about how agents enable entirely new business capabilities. It's not just about faster invoice processing anymore. Can agents enable you to serve smaller customers profitably? Can they expand your support capacity for emerging markets?

The productivity gains you've made become the foundation for your future growth. Throughout all three phases. You are also modernizing your underlying architecture. You're building APIs where they don't exist and improving your data governance.

This isn't a side project. It's the structural change that makes the agent platform actually work. You cannot run sophisticated agents on infrastructure that was built for human users. You have to evolve the infrastructure.

The execution roadmap is where your strategy finally becomes a reality. He culture shift. The biggest barrier to agents isn't the tech. It's the culture.

For 30 years, we operated on one fixed assumption. Humans do the work. Software helps them do it better. We built everything around that idea.

We hired for it. We trained for it. We built entire career paths and identities around human expertise. Now that assumption is inverting.

Agents do the work. Humans oversee and guide them. For most people, this doesn't feel like an improvement. It feels like a threat.

The people most at risk aren't the ones doing high-level strategy. They're the ones executing the routine tasks. Invoice processors. See agents handling the billing.

Customer service reps see agents taking the calls. Data analysts see agents answering the questions they used to answer. These fears aren't a sign of paranoia. They're a rational response to a massive shift.

Some of these roles will disappear. Others will be completely reshaped. The people in those seats know it. And they're watching the technology very closely for that exact reason.

This is where leadership actually matters. Not with inspirational speeches, but with honest specific talk. You have to say, here is what's changing. Here is how we're planning for it.

Here is what we expect from you. And most importantly, here is how we're investing in you so you can adapt. The companies moving the fastest are transparent about displacement. They don't sugar-coded.

They admit that certain work patterns are going away. But then, they immediately show what's being built instead. New roles are appearing right now. Agent trainers.

People who ordered agent performance. People who refine instructions based on what the agent is learning. People who handle the complex workflows that agents can't touch yet. These jobs didn't exist two years ago.

They exist now. They pay well. And while they require new skills, those skills aren't impossible to learn. The investment in re-skilling is the line between being serious and just experimenting.

Serious organizations put training in the budget. They set up mentorships between the old guard and the new practitioners. They provide real career counseling. They help people see that while the market for old skills is shrinking, the market for new skills is exploding.

They fund the certifications. They create a path where an invoice processor becomes the person who validates the invoice agents. That's not a demotion. It pays just as well.

It has more growth potential. It just requires a different mindset. This shift also changes how we look at failure. In traditional software.

Failure is bad. A bug in production is a problem. Someone messed up. There's a meeting to find out who's to blame.

With agents, failures are just experiments. If the agent tries something and it doesn't work, that's just information. You learn something about how it reads instructions or what data it needs. You take that lesson and you put it into the next version.

If the agent makes a mistake, nobody saw coming. That's valuable data on an edge case. You build a guardrail to stop it from happening again. Failure becomes feedback, not shame.

This is a massive change for most companies. It requires leaders who know the difference between being sloppy and being a learner. Deploying an agent without any oversight is being sloppy. But an agent hitting an unexpected wall while operating inside its limits.

That's learning. The distinction is everything. The enterprises winning right now are the ones that already have this learning culture. They've done the agile work.

They've embraced continuous deployment. They built organizations where testing is expected. We're small mistakes teach big lessons. And where failing fast is a reality, not a poster on the wall.

For them agents are just the next step. But for companies still stuck in command and control where mistakes are punished, agents are going to feel dangerous at every level. The future operating model, we're moving from a world where humans navigate software to a world where agents navigate software. The impact is structural.

It isn't just a surface change. The user interface is becoming a legacy artifact. The seat-based pricing model is becoming a relic. Work flows through agents instead of through dashboards.

Humans stay at the top, guiding the process and making the calls where nuance actually matters. Software is being built for machines to consume first. Organizations are being structured around who owns the agent and who governs the output. Your competitive advantage won't come from your software features.

It will come from your workflow capital, the institutional logic that your agents carry. The companies that win in this model are the ones moving the fastest. It's not because they're smarter. It's because they realize this is an organizational change disguised as a tech change.

They invest in their people. They build the right architecture. They set up the governance. They accept that this is a decade of work, not a two-year project.

By 2030, the organizations that haven't made this move will be at a massive disadvantage. Their software will become a constraint. Their costs will make them uncompetitive. Their ability to innovate will be choked by infrastructure built for a generation that's already over.

But you can start today. One agent, one workflow, learn, iterate, build. The question isn't whether you should move toward agents. The question is how much time you're willing to lose before you do.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • AI Security: Shadow AI, Non Human Identities, and AI Defense (Rock Lambros)AI Security, Cyber Risk, and Cloud Strategy on ClearTech Loop · on Service accounts84 / 100
  • Zero Trust as a Mindset: Identity, Governance, and Access | Interview with Andrew GaultSecure & Simple · on Service accounts83 / 100
  • Conquering the Desperation Mindset in SaaS Pricing with Apurv Bansal, CEO of ZenskarSaaS Scaling Secrets · on Consumption-based pricing82 / 100
  • Stop Hiring SDRs and Start Building Real Partner Ecosystems Instead with Keith Bossier | Ep. #314The Modern Selling Podcast · on Consumption-based pricing71 / 100
  • AI Fluency in CS: How to Build a Team That Actually Adopts AI with Cassie VaughnThe Customer Success Pro Podcast · on Consumption-based pricing70 / 100
  • CCT 355: Zapier Breach Lessons For Cloud Security and Setting Up TPRM Program in 15 MinutesCISSP Cyber Training Podcast · on Service accounts69 / 100

More from M365.FM

All episodes →
  • Beyond the Portal: The Strategic Architecture of Microsoft Graph and PowerShell68 / 100
  • Think Like an Attacker: Microsoft Security Exposure Management with Uros Babic [MVP-MCT]78 / 100
  • Stop Building Bots, Start Building Runtimes: A Field Guide to Microsoft Agents55 / 100
  • EXTENSIBILITY FIRST: Building .NET Systems That Survive Change with Miguel Castro [MVP]85 / 100
  • Microsoft Copilot Adoption: What Actually Works - With Chris Hinch [Microsoft]75 / 100
Explore the best B2B Engineering & DevTools podcasts →
All M365.FM episodes →