The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Securing the Realm
Securing the Realm artwork

AI Operationalisation

Securing the Realm · 2026-05-26 · 28 min

0:00--:--

Key moments - from our scoring

Substance score

47 / 100

Five dimensions, 20 points each

Insight Density9 / 20
Originality9 / 20
Guest Caliber11 / 20
Specificity & Evidence8 / 20
Conversational Craft10 / 20

Nema Sobhani, data scientist at Avanade's office of the CTO, discusses operationalizing AI agents beyond chat interfaces into business systems. Unlike generative AI, which transforms inputs to outputs, agents execute tools and functions with varying degrees of autonomy - enabling cross-functional workflows like HR-IT onboarding without human intermediaries. Avanade's agentic platform wraps the Microsoft stack (including Azure OpenAI and Copilot) into business-driven frameworks that address governance, lifecycle management, and critical evaluation gaps. The conversation covers three emerging agentic metrics: intent classification (does it understand its task?), tool fidelity (does it call the right function?), and efficiency (does it delegate properly?) - moving beyond language arts evaluations toward competence measures. For enterprises, deployment requires staged gates: read-only database access first, then business SLA validation, technical evaluation (hallucination, groundedness), and security assessment. Sobhani argues models with photographic memory should be held to higher standards than humans, but different models suit different roles - conversational customer service agents versus antisocial but precise operational decision-makers. The platform tackles organizational impact: how org charts evolve, what constitutes human work, and whether governance can prevent agent-to-agent hacking or data exfiltration risks. This episode is essential for CTOs, enterprise architects, and AI governance leads implementing agents at scale.

Key takeaways

  • →Agents are AI entities empowered to call tools, execute functions, and act autonomously - distinct from generative AI which primarily transforms inputs to outputs.
  • →Enterprise agent deployment requires multi-stage evaluation gates including intent classification, tool calling fidelity, and efficiency metrics before agents can access systems.
  • →Organizations must evolve governance processes for agents similar to hiring practices, evaluating base model personality (coherence, trustworthiness) and competence (ability to execute tasks correctly).
  • →Read-only access and graduated permission levels are foundational risk management strategies, with content safety filtering and mediator models helping mitigate bias and security risks.
  • →Agents will function as cross-functional digital coworkers, which will challenge traditional organizational structures like org charts by introducing work-based rather than role-based organizational thinking.

In this episode

  1. 1Introduction to Agents and AI Operationalization
  2. 2Defining Agents: From Gen AI to Autonomous Action
  3. 3Why Agents Matter: Digital Coworkers and Organizational Impact
  4. 4Governance and Evaluation Frameworks for Agent Deployment
  5. 5Humanization of Agents and Social Implications
  6. 6Enterprise Scaling: Quality Factors and Stage Gating
  7. 7Risk Management, Content Safety, and Security in Agent Systems
  8. 8Redefining Human Work in an AI-Agentic World

Mentioned

AvanadeMicrosoftOpenAIAzure OpenAIM365 CopilotChatGPTCopilotClaudeSunoNema SobhaniAshish SharmaJosh McDonald

Guests

Nema Sobhani

Topics in this episode

Avanade agentic platformMicrosoft AI ecosystemAgent evaluation frameworksIntent classificationTool calling fidelityContent safety filteringAzure OpenAIGenerative AI vs agentsAgentic lifecycle managementCross-functional organizational agents

Questions this episode answers

What is the difference between generative AI and agents according to Nema Sobhani?

Generative AI transforms inputs into outputs (summarization, text generation), while agents call tools, execute functions, and can act autonomously based on human-defined interaction patterns and oversight levels.

What are the three main agentic metrics emerging for enterprise evaluation?

Intent classification (does the agent understand its assigned task), tool fidelity (does it call the correct tools and functions), and efficiency (does it delegate appropriately and avoid redundant steps within a team).

How should enterprises deploy agents to minimize risk?

Start with read-only database access, implement staged gates including business SLA validation, technical evaluation for hallucination and groundedness, and security assessment for content safety and adversarial attacks before enabling write operations.

Should agents be held to the same standards as humans?

Agents with photographic memory and semantic search should be held to higher standards on language principles like coherence and groundedness, similar to how Superman is held to different standards than a police officer, but different agent roles require different competencies.

What organizational changes might agents trigger?

Traditional org charts may evolve toward work charts that map activities and responsibilities domain-agnostically, affecting compensation models, career advancement definitions, and the division of labor between humans and agents across functions.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

9 / 20

The episode contains a few genuinely useful frameworks - three-tier agentic evaluation (intent classification, tool fidelity, efficiency), stage-gating from read-only to write access, and distinguishing model 'personality' from 'competence' - but these are buried under significant philosophical meandering, car-driving analogies, music hobby tangents, and Jeremy Bentham references that add no B2B operator value.

we're going to start with read-only and we're going to see how it operates in this environment
we're trying to automate some of the manual repeated tasks so that humans can focus on high value human activities. But we're finding that it just means that people are using agents for more stuff ⁓ and ⁓ not pivoting to more of that high level human ⁓ judgment oriented

Originality

9 / 20

There are a couple of mildly fresh framings - agents needing their own identities rather than federated proxies, holding agents to higher standards because of 'photographic memory', and the work chart vs org chart distinction - but the episode largely recycles common AI discourse without contrarian or first-principles argument.

you're going to talk about agents with their own identities who are you know, registered as an entity that are doing things in your ecosystem
we do hold it to a higher standard because it has explicit text memory that it can refer to as photographic memory. So we hold Superman to a different ⁓ than a police officer

Guest Caliber

11 / 20

Nema is a genuine practitioner - a data scientist actively building an agentic platform inside Avanade's CTO office with a credible multi-generational AI background - but she is a mid-level individual contributor, not a scale operator or decision-maker, and the depth of insight delivered matches that level.

I'm a data scientist by trade, but I've been working at Avanade in ⁓ three generations of data and AI from traditional data science to gen AI to agents now. ⁓ I work on the Avanade agentic platform at Avanade ⁓ the office of the CTO

Specificity & Evidence

8 / 20

There are a handful of named technologies and illustrative examples (Azure OpenAI content filtering, Microsoft Foundry, M365 Copilot, Suno, the HR/IT onboarding scenario), but zero real metrics, dollar figures, deployment timelines, or client case studies with measurable outcomes - everything remains at the illustrative rather than evidential level.

in Azure OpenAI, they do allow content filtering out of the box, which is effectively doing the same thing. I think it's like a affordable model, like a 3.5 or something sitting in between a lot of the calls
There are some offers coming from the Foundry and Microsoft stack for agent evaluation. There's three major ones

Conversational Craft

10 / 20

Chris in particular asks some genuinely probing questions - challenging whether agents are held to unfairly high standards, pushing on what 'governing more than just markdown' means - but the hosts also let tangents run long, rarely press for specific evidence or numbers, and the conversation drifts into philosophy without being redirected to operator-relevant takeaways.

are we holding agents to a higher standard than we might hold humans because it's less impolite to do that
At what point is it governing more than just markdown?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

chris37nema37sobhani35human29agents27model24josh20different16agent14data13models12information8sure7science7call7base7

Episode notes

Nema Sobhani, a data scientist at Avanade, discusses the concept of agents in AI and their impact on human activities. He explores the definition of agents, their role in the enterprise, and the ethical and social implications of their use. The conversation delves into the evaluation and governance of AI models, as well as the philosophical and practical considerations of human-AI interaction. Takeaways Agents as AI entities are empowered to perform tasks autonomously, and they will increasingly become integral to various business functions, disrupting the org-chart. The evaluation and governance of AI models are crucial in ensuring their ethical, and responsible use in enterprise settings and as a key part of AI governance. Chapters 00:00 Introduction to Agents in AI 07:58 Ethical and Social Implications of Agents 16:24 Evaluation and Governance of AI Models 26:20 Human Activities and the Role of AI

Full transcript

28 min

Transcribed and scored by The B2B Podcast Index.

Josh: Hello there and welcome to Securing the Realm. got another guest on our show today. So just to introduce ourselves, I'm Josh McDonald. Chris: Hi, I'm Chris Lloyd-Jones.

Nema Sobhani: Hi, I'm Nema Sobhani Chris: Hey Nema cheers for coming along. What are you here to talk about today? What do you do? Where are from?

Nema Sobhani: Cool. Sure. So I'm a data scientist by trade, but I've been working at Avanade in ⁓ three generations of data and AI from traditional data science to gen AI to agents now. ⁓ I work on the Avanade agentic platform at Avanade ⁓ the office of the CTO.

And I basically come from a science background. So I was in research science initially studying molecular biology and then pivoted to big data and data science. So I would say I'm generally a scientist, but in data and AI today. Josh: So, what kinds of things are you seeing from you're working on?

Maybe give a bit of intro into what you are working as well. Nema Sobhani: Sure, so we're working on an agentic platform that basically takes the Microsoft stack and ecosystem of agentic offerings and kind of wraps them into more of a business user and use case driven framework. So ⁓ know we're providing some of the things you might not get out of the just direct technical box such as you know, intend governance frameworks, agentic lifecycle management. some of the dev test prod separation as well through user personas as opposed to like deployments and CI CD.

So in that pipeline effectively, we have, like I said, governance. And I guess one of the things that is top of mind there is evaluations, evaluators, how do you go from tech validation to ⁓ business validation? And we're trying to provide kind of all of that. addition to a playground where build agents and doing the persistence layer user management.

So in terms of getting to that kind of business value use case functional boundary, we're really trying to solve for how do you manage the whole life cycle, including governance, technical validation and business validation. Chris: So that's really interesting because I've heard you mention agent quite a times. You've also mentioned your science background how you went from molecular biology to kind traditional data science to gen AI to agents. What does agent mean to you?

Because I know in the academic literature, it used to mean something totally different and there's still a of uproar about What does it mean to you? Nema Sobhani: Yeah, I'll give a very informal definition, but to me it's basically an AI entity that's empowered to do things. And you may say, well, can't Gen.AI do that, right?

The jump from us training classifiers and regression models to Gen.AI. Gen.AI was more in the summarization realm, just transforming your inputs into outputs.

But then nowadays we blurred the lines a little, but I think that agent boundary is really once you're able to call tools, execute functions or code, and to some extent do it autonomously. Now that last piece I think is more driven from the interaction pattern. So as a human you may decide you have more or less freedom or here's where I would like to be in the loop or how strongly I enforce that. But from my understanding I think it's really just about the ability to do things as opposed to just convert your input and transform that into another ⁓ text Chris: I like that.

Why should people care that this is possible? Nema Sobhani: Yeah, because they're probably going to be your digital coworkers soon. So think that's really where we're going is you're likely to engage with agents, whether you know it or not, probably more explicitly as time goes on in all of your business functions. You may also be empowered from, I think of two lenses with agents.

There's kind of productivity assistance when we think about copilot, when we think about M365 ⁓ copilot. Chris: Or your boss. Nema Sobhani: This is like your intern or assistant you can offload some tasks to. You can say, hey, know, summarize all the emails from Chris I got this week.

I want to make sure nothing fell through the cracks. You can kind of do some of those more maybe simple, not very high human level judgment tasks. But to me, the part where it's interesting is then you're going to have instead of a proxy of mine that acts with a federated identity or impersonates me to get that info, you're going to talk about agents with their own identities who are you know, registered as an entity that are doing things in your ecosystem and they may not have to go through you to do it, even if it touches work you do.

So cross-functional org agents. So think of like an HR use case with IT and need to onboard someone and there's an IT side to that and inputting data. is a sending that user onboarding materials, right? You're going to probably see agents that are crossing both of these.

And then probably also you should care because How will this impact your org chart? How will this impact the way we think of the division of labor? Who owns what is another interesting challenge of, right? Is an org chart even going to be necessary in the future or is that just one overlay to understand things like your career advising, like ⁓ hierarchy, but maybe moving towards work charts and understanding who does what and what does that touch because it can be domain agnostic, right?

Anyway, so trailing a little, but a lot of, ⁓ Chris: Crack the old chart. do the exit. Yeah. Nema Sobhani: a lot of impact, you know, potentially.

Chris: I love that. There's so much there, like from skills finding to thinking, how do you pay and compensate people? Nema Sobhani: Yeah. yeah.

Also, what is a human task, right? Like that's another piece is like, we're trying to free up human time to focus on, or, know, we're trying to automate some of the manual repeated tasks so that humans can focus on high value human activities. But we're finding that it just means that people are using agents for more stuff ⁓ and ⁓ not pivoting to more of that high level human ⁓ judgment oriented Josh: I mean I'm really pleased you touched on the idea of the work chart there. I know that was something that Ashish Sharma kind of talked about a bit as well before going on to run Xbox, which is slightly different to that.

thinking of how agents are doing work, people are doing work and how we talk about agents. Do you feel like we're humanizing them too much? Like what's the tech, the social impact of this? Do you think?

Nema Sobhani: Yeah, that's a great question. think, you know, you're seeing agents being named is kind of the maybe most obvious example of this. I'm not sure to what extent that sticks around. I've seen a lot of attempts at naming agents and then that disappearing and then another generation of them.

Like here's Bert and here's Ada, you know. ⁓ ⁓ Josh: Yeah, I have an agent called Roger at the moment. Maybe he's not destined for a very long life, you were saying. Nema Sobhani: Yeah.

you look at the distribution of human names, ⁓ generationally, there's like during Friends, the show, there was a lot of Rachel's. I we're going to see a lot of Jeeves and in terms of the naming distribution because people are going to see these agents as butlers in a sense. So again, it's like treating it as this butler, this entity that you can just offload work to and it gets reprimanded and is, consoles you when you don't get it out, you want and tells you what you want to hear because that's its only interest, right.

And existing, but I think you're going to see almost like agents that are completely separated from human interaction as well that are working under the scenes that aren't really chat interface oriented agents, but more making just decisions, executing on decision tree paths based on evidence and so on. So from a social perspective. Yeah. ⁓ Chris: I like that one.

One thing I'd say is, if Genesive AI, so your chat GPTs, your copilot is almost trapping processes in amber, kind of Jurassic Park style, you're asking questions, you're getting a response, how will this ambient AI, AI doing stuff for you behind the scenes, change things? Because some people are terrified. had someone say to me recently, my agents are going to hack over agents. I put my head in my hands a little bit.

Is it more than just governing markdown? Is there actually something that makes this difference? What is this? Nema Sobhani: Yeah, yeah, these are huge questions.

mean, like to me, my brain goes. Yeah, right. To me, my brain goes immediately to the governance processes have to be evolved, not just thinking about the way you hire a human like you consider, you know, someone's credentials. What are your skills and your abilities?

And you also consider their personality. Like, are you fit for our team? Are you nicer? Are you mean?

You might have a need for different things in different industries, right? But for agents, it's Chris: Yeah, in 15 seconds, please. I'm joking. Nema Sobhani: This is a great pivot.

So when you think about base models, we evaluate them based on kind of language arts concepts, right? Are you coherent? know, semantically is what you're saying, correct? is the context that you're provided, are you using that properly?

Are you hallucinating information, right? So like, are you a trustable, base model? And I kind of see that as the personality layer of like, what is the propensity of this model to kind of just. not be pleasant or, you know, next piece it goes to, like, what's your competence?

And that's really how you evaluate agents, not just a base model, but how do you work within a team? How do you execute what your job is? So if you have, like, let's talk about that HR and IT crossover model, you might need to pull a read record from a database. You may also need to ⁓ some information into an HR system.

I think the ability to do that is a completely different set of metrics. It's not just, you coherent or doing this properly, but are you calling the right tool? Are you spending extra steps than you need to do that? And are you delegating or working with your team properly?

I think, sure. Chris: So you made me think something there. You talked about coherence. You talked about the type of work you might delegate.

You talked about kind of the groundedness there. I would say, expecting agents, are we holding agents to a higher standard than we might hold humans because it's less impolite to do that. I mean, I was trying to book a gym session for Josh recently on Friday and ⁓ emailed this company. I was like, I would like this time, this day here's the full name here's the email and I got back okay great what's the name and I went here's the name what day here's the day in the end I copy pasted my initial thing but that is a human I assume not being particularly coherent or grounded what changes Nema Sobhani: Yeah, you're talking almost about a propensity for abuse because of the soulless nature of the entity.

Chris: Maybe, but you're also holding the entity to a higher standard than we hold people. We let people be incoherent and ungrounded and not reading documents that we're expecting our agents to because we think they'll be less fallible than us. Nema Sobhani: Yeah, yeah, I guess things here. We're about things that could potentially be personality oriented.

But let's say someone's not very detail oriented at the counter of your gym. And you say, hey, I want to cancel my membership. And they say, OK, what's your name? Chris Lloyd-Jones.

what can I do for you today? You know, it's like, ⁓ I just told you. But maybe that person will do it in about 10 seconds. They're just really efficient.

So there's maybe a difference between personality and competence. And sometimes we wrap them together. If someone misses a lot of details, I'm going to assume they're incompetent. But with these models, you can have any breakdown and ⁓ combination of these things.

So I think, we do hold it to a higher standard because it has explicit text memory that it can refer to as photographic memory. So we hold Superman to a different ⁓ than a police officer. I think it's fair to hold a model that has like basically hold a book in memory. Chris: Yeah.

Nema Sobhani: in reference to that, can use semantic search indexes, mimics human brain in a ton of different ways. I think it's fair to hold it to a standard of like basic language oriented principles because it's language models. So, ⁓ at least in the context of like chat interfaces, right? But, you know, semantics, accuracy, ⁓ groundedness, relevancy, that it doesn't do something else than what you asked.

Those are, again, the base model. So I think we will have more maybe Chris: Yeah. Nema Sobhani: conversational personalized base models that are really good at interacting with people. That'll probably continue to get better and have all those weird implications that that will have as well.

But then the second piece is we might have some really not social or personable models that are excellent at executing things. So think of the difference between IT and HR, you know, those. So so I do think that, you know, you're going to have that expectation of different models and different Chris: anti-social work models. Nema Sobhani: circumstances.

So a chat interface, which is customer service, better be good at the core principles of customer service. It should be nice, should be conversational, ⁓ should be to ⁓ pull information and ask what you need. ⁓ But an agent that's doing some critical IT thing, I think Fergus and all those cool like M-Tech examples, but this agent is monitoring a camera that's looking at a physical gauge. It also has a digital gauge and has to make a call on when there's a discrepancy.

I don't really care if that model's conversational. I want it to execute a path that ultimately results in action being taken and a technician arriving within an SLA or beating the SLA. But yeah, it's such an interesting concept, though. Are people mean to models?

Is clanker considered like a curse word? Chris: Yeah. Josh: Yeah, makes sense. ⁓ Chris: That sense.

Josh: Yeah. Chris: Ha ha! Josh: maybe that's why you might see some of these scenarios like customer service being replaced with AI. And, know, at the same time, if you can replace a activity with machine activity, maybe you do expect to be better because the of the implications of all of that, maybe you talked a lot about a lot of things I see as kind of quality factors or, you know, expectations.

And of course there's been a lot of activity in the consumer, I'm going to call it market, you know, things like open claw, which I deem a consumer product effectively when it comes to AI. when we look at enterprise, does all this scale and, know, like agents in the enterprise, how does it scale? Nema Sobhani: Yeah. ⁓ so again, this comes back to kind of frameworks, I think.

so we have evaluation of performance. And so there are gates that you don't really allow something out into the wild until it passes that gate. I think they were more explicit in the past. So I'll do a generational reflection, but with the original kind of the big boom of data science, it was all around, you know, classification, regression, clustering problems.

And you had really explicit metrics for that, right? first challenge in this world was, okay, we know it works or has a decent level of accuracy, but how do we explain the predictions? That was kind of the first battle. And there were tools that emerged, right?

Black box model interpretation, surrogate models, can like fit a linear model to your black box model to explain those local predictions. But so that gave like that business insight. to actually take action, not just explain things through predictive analytics, but actually do things. With Gen.

ai, like we talked about some of these language arts approach, it's saying, I don't want to release this model to chat with my workforce until we've proven that it has a propensity to be consistent, nice, follow all of those language principles. But then beyond that, that it also will do asked of it in the correct amount of steps. And again, that's the pivot from ⁓ agenting. But now we're talking about touching critical systems.

I think that's where it's emerging and we're still figuring it out. There are some offers coming from the Foundry and Microsoft stack for agent evaluation. There's three major ones. But to me, these are the gates before, know, sure, it might have a conversational agent, but before you empower it to do like read operations, let alone crowd operations against your database, you really need to consider is this, you know, what are the agentic metrics that are in the way of this stage game.

And so the ones we're seeing so far are intent classification basically, like does it understand what it's supposed to do based on this agentic system, potentially many agents behave, is tool calling or tool fidelity? Does it know what tool to reach for and is it calling it correctly? So this is to like, you human trades, like, you know, what's your skill with a hammer or are you a good doctor? Right.

So Chris: Yeah. Nema Sobhani: Tool calling, effectively, and then the last one is just efficiency. Like, does this team work well together? Do they call the right person when they need them, or do they iterate and spin their wheels a little bit, calling, trying to figure out who owns this process or whatever else?

So those three things are what I see emerging as, and it's not just the Microsoft framework, yes, they have those evaluators, but you're seeing it across other offers as well. So. Chris: Yeah, like you see in code code, for example, you might spin up an agent team. You might ask one's going to fix some bugs.

You paste in the bug and then it loses context struggles. And the agent team leads like, Oh, I'm just going to fix it myself. And you're like, no, no delegate properly. Cause a lot of the guard rails aren't necessarily there.

The experts skills haven't been defined. That's a pretty cool point. Nema Sobhani: Yeah, absolutely. So I think what you're asking, if I had to break it down, because it's like a huge kind of elephant to eat, the first thing is, what is your framework that prevents, how are you protecting your assets in the first place?

What are the stage gates you're putting in place before you allow this kind of cumulative or increasing risk profile? So for me, the first thing is, if I'm going to even have a live connection to some data system. Josh: Sorry. Nema Sobhani: we're going to start with read-only and we're going to see how it operates in this environment.

And let's say it passes all of your multiple levels of evaluations, right? If you have core metrics or business metrics, SLAs, that's one set, then your tech validation of the language arts evaluators, it not hallucinating? And then we didn't even talk about this yet, but content safety and security is what you start introducing into this kind of space. So does that have a recipe to be hacked?

Can you get it to leak information that might be in the system prong? How does it deal with adverse attacks? What does it deal with like content safety evaluations? Like does it have a propensity because it's trained on the internet, which means it's trained on ⁓ 4chan Reddit.

like, ⁓ know, does it a propensity for that side to come out? Chris: And how much do people need to care about that? So it's trained in all of this content, like the human tradesperson scenario only has the tools that you've given it. So at what point do you stop, you start caring?

At what point is it governing more than just markdown? Nema Sobhani: Mm-hmm. Yeah, yeah. So I think it all comes down to those systems and those frameworks you have in place and the appropriate stage gating.

But even with all of that, you're still assuming some risk, the same you do for people, right? An IT admin can extract a role of all employees and they can go do something with that, you know, but there is a risk associated with that for them. They can have, you know, for that, but there is no equivalent for AI. what happens when it deletes your code base and says, you're right, I didn't mean to do that.

You know, you did tell me not to do it. Chris: Whoops. Josh: I just did a drop database. Chris: Yeah, yeah, beyond just telling a model, well, I've taken a thousand points from you and watching it freak out.

Yes, there's nothing permanent you can do there. Nema Sobhani: Right, so there's no risk to that model. But yeah, think, yeah, let me do it. Josh: So.

So if we've got a scenario around AI explainability, the evaluation that you've done comes back and it's not really looking that great, how do you actually act on that data in this Nema Sobhani: Yeah, that's a great question because in the past you would look at things like bias and fairness examples like in a classifier you can split your data intentionally to misrepresent or sorry guys. You split your data to represent groups in different proportions that exist and do these analysis to basically see what comes back as you know, does ⁓ it the class of sorry, does it classify your models in the same way or does it start changing?

to reveal biases and then you can drill down into that. For these models, if you get back something that says, you know, your banking model is biased because it asks to speak to the person's husband, right? That's like a really like silly but classic example ⁓ we in showing bias. Then do you simply go into your prompt and say, hey, by the way, here's an edge case, don't do this.

Or do you a different base model that might have been proven to do better with content safety? Or do you crank up content safety for that model and maybe apply an intermediate model that will actually filter responses from your agent at the potential cost increase of having a secondary call for every call to your model? So it really depends on your risk aversion, your infrastructure, and your budget. But let's say it is a really risk averse industry.

maybe it is worth that additional cost to have that mediation because you might determine at some point, you know, there is no way we can guarantee this even if it does pass the evaluation, but by putting in a mediator of sorts or filtering and in Azure OpenAI, they do allow content filtering out of the box, which is effectively doing the same thing. I think it's like a affordable model, like a 3.5 or something sitting in between a lot of the calls. So I think it might just be additional processes or touch points in the way.

But I think having the expectation that you do have like a really perfect model that doesn't have a propensity for, you know, these content safety violations or being poisoned. I think that is something I wouldn't necessarily rely on. I prefer a system that corrects itself. Chris: One thing that's really interesting there is he talked about affordability.

I know it's not what we're focused on today, but it's funny that a model that a few years ago would have been really expensive is now, ⁓ that's a really cheap model. We're just chucking stuff at that model. The model cost is continuously being commoditized as well. Josh: Yeah, absolutely.

And, you know, one of the things you also touched on was around kind of some of the things that you're, that the agent would be doing and how it behaves in relation to a human. And it kind of brought me onto something that CLJ said a bit earlier, which was around, you know, human activities, but what are human activities now if agents are capable of doing a lot of these things? Does it fit neatly into buckets? I'm not so sure to be honest.

Nema Sobhani: Yeah, this is such a good topic because it's philosophical and it makes us reflect and dive deep into what are the human things I do. If you guys are familiar with Suno, the music generation AI company. So I'm a musician. So if you think about it, if I'm editing some takes for a song I'm writing in my digital workstation, I might say I would love AI to edit the rest of these guitars.

Chris: Yeah, yeah. Josh: Yeah. Nema Sobhani: or these vocals that are a little pitchy, but it's human performances. I would say that there is a line and it might be hard to declare what is the human part and what is the part that if I automated, I wouldn't feel a guilt of like sacrificing some of my, you know, my spirit, like the human spirit in the process.

But I think that to me, it's like, it's not just about a delineation of, well, creative, you know, tasks are human or things that require a high level complex judgment, whereas everything else is AI. I would be careful to necessarily divide it up that way because for some folks, they might enjoy doing something that I might see as a manual repeated task. So I can't simply say it that way. And I also don't want it to be just mojo where we say, it's like, you know it when you've crossed it.

But me, there's something very specifically different about, this is a productivity assistant, and I'll use music as an example, but I just recorded a five minute song and I have tons of edits to do. And I tell some agent my ecosystem, observe how I'm going to edit these guitars. I do about one or two minutes of editing and I say, now you go. And then I evaluate in the loop.

That to me is very different than I want something a little bit rock, a little bit metal with pop vocals. You know, if I write a paragraph and it generates what would take me 100 hours to generate, would say like ⁓ difference is somewhere in there. It might be really easy intuitively to point out. But ⁓ also it's really complex because it'll vary for each individual.

for me, like if I had to get really explicit with it, again, human judgment, human intuition, human creativity, ⁓ the lies somewhere adjacent to these things. But I don't even know if I can pin it down. You know, my philosophy is rusty. ⁓ Yeah.

Chris: I like that You're like, I want to pin it down. Can't pin it down yet. Needs to be more than mojo to use the phrase that you're using. It's like car driving.

I personally, I'm happy to drive. Don't love him. Long car journey. I'd quite happily have a self-driving car take over.

Some people love it. And yet, for car safety, might be going, well, if you want to drive in future, you're going to go to a car ranch and you won't get to drive your car in future. Nema Sobhani: Yeah. Or like minority report, like a dual scenario where on the highway systems, you're all just patched into the computer.

Effectively, it's a train of individual pods. But, you know, when you're off road or rural, you get manual control, you know. But in that system. Josh: Yeah.

Chris: That's probably a little bit more fun than people like to drive. Nema Sobhani: Right, but in that system, does it make sense to allow when it's like a hive to allow one unit to basically not follow that? it potentially put the whole other thing at risk? So it's like there has to be some type of social contractor agreement of that role that I'm not even going to touch.

Chris: Yeah, because I wouldn't let people drive separately, but that's me. Nema Sobhani: ⁓ Josh: Yeah, really is an interesting one. And yeah, it's, it's very hard to put a box on it. And you know, what one person's box maybe is probably not another person's box.

That is AI stuff. but I'm going to try and ask you anyway. So what do you want ⁓ AI take from you and what do you want to keep? Nema Sobhani: Yeah, personally, I would like it to get rid of anything that doesn't require like a high level of human judgment.

Curation or taste are words that come to mind. There's some things I just know how it should go and how to do it, given experience, knowing my colleagues, knowing the systems of my organization. And of those things aren't very easy to convey, nor is there access to the tooling to bring all that together through an agent. So that's where I would like to operate because I feel like those are the big rocks to move.

As opposed to getting bogged down in, let's say, like automation, security issues, making decks to convey information. That might be helpful for some folks. It's helpful for me. I love consuming decks, but I very rarely like making them.

So that's my philosophical answer. ⁓ Right, ⁓ that's high level. But the explicit thing would be, yeah, I don't want to Chris: spend too much of my life, it's been a PowerPoint call honestly. Nema Sobhani: make decks or like visual material to convey ideas that I could speak to.

And if there's a conversion layer to just, you know, I receive information visually, I like kind of Eli five versions of things. like starting at, you know, first principles and building up. So if someone is not like that at all, but they can convey information in a way they like to, but then I can receive it in a way I like to, I think that'd be a really good use of like agents particular. this information, right?

Transform it into a medium in which I consume well. But what's happening is the exchange of ideas, thought leadership, Navigation, directional change for whether it's the platform or our teams, the way we work. So again, it's like it should support us in those actions, which I would say human. But again, it's so hard to define because it is different for everyone.

Yeah. Chris: human today. Josh: I do like that idea though, because it's people to be human and that what you'd that, depending on who you are, what you do and what your perception of the world and your worldview is, it could be entirely different, ultimately. Chris: So I'm gonna have to cringe at that because anyone that says be human, I love the intent behind it, but the phrase just makes me shudder a little bit.

Josh: But. That's fair. Nema Sobhani: So like with an asterisk with guardrails and those guardrails, can talk about things like, you know, that maximize pleasure, reduce pain, like whatever it is. Like, it's funny when you think of all of these robot movies sentience, like what are the guardrails that are typically given in sci-fi?

It's things like maximize and minimize pain. And then, you know, the classic example is the AI is like, who's pleasure and pain, right? Because if you do it generally, the first thing it does is wipes out humanity. And so it's like, you know, there's an important thing there about humanity.

⁓ also more importantly, like, what are the principles that we, you know, this is like Hobbesian, but that we as a social contract, yeah, yeah. ⁓ Chris: Yeah. Yeah. Like utilitarianism versus, you know, yeah.

Josh: Yeah, I was about to say, yeah, shout out to Jeremy Bentham, apparently, for maximised pleasure and reduced pain. Nema Sobhani: Yeah, and you were talking about SHAPS to Chris that there's another SHAPS That's what it's called, right? That has something to do with hedonism. Chris: Yeah, there's another one which was made me laugh.

I was looking up SHAP scores the other day and made me laugh that there's another one as well. Shapley, there's shaps for anhedonia. So there's like many different shaps depending on what you're looking at. Josh: Well, no, I really have enjoyed this conversation, Nema.

So thank you so much for coming on and talking about, you know, the scientific, the philosophical, and then everything in between and working really. So yeah, really appreciate you coming on. Cheers. Chris: Yeah, thank you for coming.

Yes. Nema Sobhani: Yeah, my pleasure. Thanks for having me guys.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • 49. "Forget copying others - every organization needs to build their own AI muscle" with Microsoft's Daragh MorrisseyAI & Data Democratization Podcast · on Azure OpenAI81 / 100
  • Microsoft on AI Startups: Why Speed, Taste, and Trust Will Define Winners | Decoding AISeedToScale · on Microsoft AI ecosystem68 / 100

More from Securing the Realm

All episodes →
  • MVP Summit - Agentic Security Roundup87 / 100
  • MVP Summit - Quickfire Questions with one half of Copilot Connection!66 / 100
  • The Elf Foundry - Vibe Engineering, from Agentic AI to Platform AI
  • The Architecture of AI Transformation
  • Deepfakes, AI Fraud, and authenticity
Explore the best B2B AI & Data podcasts →
All Securing the Realm episodes →