
The Road to Accountable AI · 2026-05-21 · 33 min
Key moments - from our scoring
Substance score
70 / 100
Five dimensions, 20 points each
Holistic AI's Emre Kazim traces the evolution of AI governance from its 2020 origins through 2025, revealing a fundamental recalibration of how enterprises approach responsible AI. Originally inspired by the Knight Capital trading algorithm disaster, Kazim and co-founder Adriano Koshiyama published "Towards Algorithm Auditing," establishing a framework around technical variables like explainability, privacy, bias, and robustness paired with traditional governance structures. However, Kazim now argues the analogy to data governance and GDPR-style compliance was incorrect - AI governance increasingly mirrors cybersecurity's incident-driven, technically intensive model rather than documentation-heavy compliance regimes. This shift has profound implications: ownership moves from legal and policy teams to technical organizations, assessment requires deep interrogation of systems themselves (red teaming, bias testing libraries, explainability tools like Agent Graph) rather than form-filling, and the governance imperative is enabling rapid AI deployment rather than restricting it. Kazim addresses why human oversight as a primary risk mitigation strategy has eroded due to system sophistication and real-time monitoring challenges, and predicts that best practices emerging from industry incidents will cascade faster than regulatory frameworks can accommodate - particularly given generative AI's explosive development between 2023-2024.
Emre Kazim's original thesis that AI governance would follow data governance and compliance models (like GDPR) has shifted to a cybersecurity-driven model emphasizing incidents, technical assessments, and real-time performance management rather than documentation and regulatory compliance.
Modern AI systems, particularly agentic systems, are too sophisticated and dynamic for humans to meaningfully oversee in real-time, making traditional human oversight an inadequate control mechanism.
It includes specialized tools like explainability functions for agentic systems (e.g., Agent Graph), bias testing libraries, benchmark assessments, and deep interrogation of system behavior - going far beyond questionnaires or impact assessments.
Best practices emerging from real-world incidents and industry responses will cascade quickly across enterprises and become de facto standards, rather than waiting for abstract regulatory frameworks.
After an explosion of pilots in 2023-2024, companies are now seeking to demonstrate actual value and ROI from AI systems rather than pure experimentation, driving demand for governance tools that enable faster, safer deployment.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode offers substantive re-framing of AI governance as cybersecurity-driven rather than compliance-driven, with specific technical examples (Agent Graph, bias testing libraries, hallucinations). However, there is considerable throat-clearing and narrative meandering (20+ minutes into a 33-minute episode before concrete examples), and several concepts are presented without deep operational detail that would help a B2B operator implement them.
the AI governance space and market, if you will, is gonna be defined more by more in analogy to cybersecurity than to governance in the way... to compliance in the way that I thought it was going to be
it's moving more into incident management, performance management, you know, um, rather than, "Oh my God, we've got this compliance thing that we need to tick off
The core thesis - that AI governance mirrors cybersecurity's incident-driven evolution rather than data governance's compliance regime - is genuinely fresh and contrarian to prevailing regulatory thinking (EU AI Act). The vanguard vs. democratic adoption model and the observation that the problem is becoming 'tech versus tech' rather than 'law versus tech' are non-obvious. However, once stated, these ideas are not deeply explored with new evidence or challenged rigorously.
I think the EU AI Act is not a direct, you know, comparison to GDPR. I think AI does something else
I think it's moving more into the deeper tech space than I think people had anticipated, or even I had anticipated
Emre Kazim is a genuine practitioner - co-founder and CEO of a purpose-built AI governance company who has lived the evolution from academic thesis to market reality since 2020. He has published in top-tier venues and won OpenAI-backed hackathons. However, the episode lacks a second guest or external voice to create productive friction, and Kazim's insights are largely uncontested monologue rather than debated.
So I'm co-founder of the company, and I co-founded the company along with Adriano Koshiyama
We won a prize. We were awarded this prize by one of the co-founders of OpenAI. Um, it was a hackathon, actually, 600 teams, uh, competed
The episode includes a few concrete examples: Knight Capital ($440M trading algorithm loss), Agent Graph tool, open-source bias libraries, GDPR vs. AI Act, and the distribution of companies across traditional ML, generative, and agentic systems. However, most claims lack supporting data. The discussion of '23 - '24 pilots and 2025 reckoning is asserted but not backed by market data, customer counts, or measurable shifts in demand.
So it was a trading algorithm that glitched, and I think it wiped out $440 million from the business
There was a piece of work that we did, uh, I'm happy to say, was on explainability of agentic systems. Mm-hmm. Really cool piece of work actually, um, Kevin. It - We won a prize. We were awarded this prize by one of the co-founders of OpenAI. Um, it was a hackathon, actually, 600 teams, uh, competed, and it was just about this really cool, we called it Agent Graph
Kevin Werbach asks intelligent, probing follow-up questions (the cybersecurity analogy push-back, the research-in-industry concern, the infinite regress problem) that create some productive tension. However, he often yields quickly to Kazim's lengthy answers without sharp redirects, and several moments lack deep follow-up. The conversation is respectful but not confrontational enough to force precision or reveal gaps in reasoning.
Can I push back on you a little on that- Please, please ... on the cyber analogy? Absolutely
Although, you know, we're going to need to wrap up after this. The, the, the challenge is w- when you have these, uh, applied companies especially, that have valuations where they're raising in the hundreds of millions or billions- Yeah ... of dollars, uh, then people who otherwise would be in academia and doing research increasingly are doing their work- Behind closed doors
Computed from the transcript - who did the talking, and the words that came up most.
Holistic AI was one of the first companies built specifically to govern, audit, and red team AI systems. As co-founder and co-CEO Emre Kazim explains, its original thesis was that AI governance would mirror data governance: a compliance-driven regime. He now believes the better analogy is cybersecurity: a more technical, incident-driven discipline where best practices emerge from real-world events and propagate across industry, rather than descending from abstract regulatory frameworks. Kazim argues this shift has significant implications for who owns AI governance inside enterprises, what skills they need, and why documentation-and-reporting vendors are unlikely to capture the core of the market. Kazim also makes the case that human-in-the-loop oversight, long treated as the default answer to AI risk, has become untenable as systems grow more dynamic and agentic. He distinguishes between two enterprise adoption patterns: a democratic model in which every employee has a copilot, and a vanguard model in which a small number of mission-critical agentic systems drive most of the value and demand most of the governance attention.
Transcribed and scored by The B2B Podcast Index.
This file was generated by Descript Kevin: Hi, I'm Kevin Werbach, Professor of Legal Studies and Business Ethics at the Wharton School of the University of Pennsylvania. For decades, I've studied emerging technologies, from broadband to blockchain. Today, AI is promising to transform our world, but AI needs accountability, mechanisms to ensure it's developed and deployed in responsible, safe, and trustworthy ways. On this podcast, I speak with the experts leading the charge for accountable AI.
Holistic AI was one of the earliest companies built specifically to govern, audit, and red team AI systems. In this episode, co-founder and CEO Emre Kazim talks about how his original thesis that AI governance would look like data governance anchored in compliance regimes has given way to something he now thinks looks more like cybersecurity, driven by incidents, technical assessments, and an arms race between AI capability and AI oversight. We also get into why human in the loop is no longer viable as the primary means of addressing AI risks, the vanguard model of enterprise AI adoption, why best practices will become the de facto standards for responsible AI, and why he believes that meaningful research capacity will be the price of entry for anyone working in the AI governance space.
All right, Emre, welcome. Thank you for joining me. Fantastic. Thank you for having me, Kevin.
Uh, so you co-founded Holistic AI in twenty twenty, if I'm right. That was before there was ChatGPT, before the EU AI Act. Uh, there was, you know, a lot of activity around, uh, responsible AI, but, but things have changed a lot. So what is it that, that originally got you started that, that made you see the need for this kind of service?
Emre: Kevin, yeah, this - that's actually a great question because it's been, uh, uh, the - it's been so long now since we, uh, got going, um, that there's a... Sometimes when you look back in retrospect, you start to realize how many different kind of events precipitated certain kind of responses and other events which we kind of foresaw and act- and predicted and so on and so forth. So let me give you a bit of a primer in terms of how Holistic AI, um, spun out and what our kind of origin story is.
So I'm co-founder of the company, and I co-founded the company along with Adriano Koshiyama. So Adriano and I met in the computer science department of University College London. So, uh, I'm sure your listeners are familiar with UCL. Um, it's the home of DeepMind, uh, Google's DeepMind.
Um, we - I tr- I claim Hinton as one of our Nobel Prize winners, but certainly Demis, uh, everyone was here at, was at UCL, and there was a lot of interest in foundational, uh, AI work and the generational synthetic data and so on and so forth. And in that context, there was an emerging and, and, and a very important conversation about trust Which was if we're going to use these systems or, and meaningfully harness their potential, then we have to have confidence in those systems.
So, um, Adriano was doing his PhD at the time. I think Adriano started, um, in 2018 I met him, but he started before then. I started my postdoc in 2018. That's when I ended up at UCL.
And the general conversation that we were having at that time was, how do we engender trust in algorithmic systems? And the use case that we were focused on was the example of a trading algorithm. Um, it was, the company was called Knight Capital. Kevin, I'm not sure if you're familiar with this use case.
So it was a trading algorithm that glitched, and I think it wiped out $440 million from the business. Hmm. So it was this- Oh, yeah ... huge, yeah- Now, now I recall.
Yes, yeah, yeah. It was this huge kind of use case in financial services, and there was this awareness that, okay, look at how mission-critical these AI systems are. You know, we have to have confidence that these systems actually work, and they're reliable and safe and so on and so forth. What was interesting is that at the same time, that was very much about operations, but at the same time, a social awareness was emerging around issues of bias and explainability.
So we weren't sure, okay, how are we gonna go about doing this? And we thought our original thesis was that if we create a certification standard and a, an algorithm meets that standard, then it is, you know, ergo trustworthy. Now, the issue was, you know, what kind of certification standard? What, what was the standard that we were going to, uh, generate?
And also, why were we in a position to forward that standard? So what we did, as academics, uh, we had the power to convene, and we brought the great and the good in the UK together, so then we had regulators, we had the civil service bodies, um, and then we had industry, and then we had ourselves as academics. And we thought, okay, what we'll do is create a certification standard that all of this cohort or group agree to, and then we'll kind of publish it as a public statement and then we can get going.
And what happened pretty quickly is we realized that certification is an endpoint rather than a starting point, and the whole problem got flipped on its head, and two streams emerged. One was conversations about setting up a kind of national synthetic, uh, data center, and we were gonna basically set up something where you could test your algorithms against these data sets, and then you would know if they were biased or whatever it was be-, um, because we knew the data wasn't, so you'd be able to test the actual algorithms against that.
And the second work stream was what we referred to at the time as algorithm auditing, which was saying, okay, let's just take a single system and let's say, how does one go about auditing an algorithm? And what does it even mean to audit an algorithm? So, uh, Adriano led the paper, I was second name author on this And he wrote, which I, what I consider to be the foundational text in the field, which is Towards Algorithm Auditing. And the paper, Towards Algorithm Auditing, became a really powerful framework to say, "Okay, this is how we would theoretically go about, um, assessing an algorithm."
And the core thesis was there were technical variables such as explainability, privacy, uh, bias, and robustness, and robustness was a kind of, uh, umbrella term for re- reliability, reproducibility, security, and so on. And that those technical verticals, we were calling them verticals or variables, were controlled for by just like good traditional governance, documentation, compliance, taxonomies of responsibility, accountability, and so on and so forth. Um, and then that was the kind of core thesis.
Mm-hmm. Now, the problem is you can't study medicine in the library, you know? Of course. You study medicine in the hospital.
And, uh, Ajorn and I were like, "Okay, if we're actually gonna have impact, we've gotta do this in the real world." So we spun the company out, and we got going. And what we started to do was say, "Okay, how does one go about, in quotes, auditing an algorithm?" Um, and you know, we used this language of auditing because we wanted to evoke a kind of seriousness within the community because AI ethics and responsible AI, um, quickly became quite fluffy as terms.
Um, so we were trying to make an analogy to say, "Hey, this is as serious as, as financial auditing," even though the analogy is not that great. Mm-hmm. So the kind of concept and the term changed and morphed, and we can talk about that perhaps, into things like, you know, is this AI GRC? You know, is this AI compliance?
Is this AI governance? And so on and so forth. But actually, the core of it was- Mm-hmm ... you know, this is a, the foundational problem of this space, which is we can have all this technology, but if we can't use it, what's the point?
And how can we create the kind of confidence and trust in that? And that moved from an academic exercise- Mm-hmm to, you know, maximizing our impact through, through a business. Kevin: Fast-forward to today then, what has remained of that initial vision, and what's changed? Emre: So I think at the...
So one thing that's changed is I think, um, I think at the start, I believed that what we had... So, and I still believe this thesis, which is that 2010s was defined by data, and the 20, uh, 20s have been defined by AI. You know, a little primer here, I think the 2030s are probably gonna be defined by robotics. Mm-hmm.
And, um, which is this kind of greater and greater embodiment or- Mm-hmm ... you know, autonomy of systems. And we thought, I thought that the way that data, w- a primary way in which data was governed was through legislation, and for example, we had the GDPR, and the GDPR, it m- created a huge market. I think that, um, that was ultimately wrong.
I think it has not turned out to be the case. I think the EU AI Act is not a direct, you know, comparison to GDPR. I think AI does something else, and I think that the AI governance space and market, if you will, is gonna be defined more by more in analogy to cybersecurity than to governance in the way... Uh, sorry, to, than to compliance in the way that I thought it was going to be.
Now- That, that's really interesting. Could you say more what that analogy means? So let's say, so I always thought, okay, what is AI... Uh, let's call it now AI governance.
Let's, that's the category. And when you're doing category creation, uh, you know, no- it doesn't exist. You have to think, "Okay, what are the closest analogies here?" And I thought about, you know, a bunch of ways.
One way I thought about was, okay, we've had cloud transformation, and then we got cloud governance. You know, does this look like that, you know? We had data transformation, and then we got data governance. And when you look at data governance, yes, there's things like providence, there's things like, uh, privacy, there's, uh, not pri- uh, compliance and privacy is not the sum total of data governance, of course.
There's just a whole regime there. And then there's cybersecurity, right, which is a different kind of problem, uh, and a different form of governance, if you will. And my thesis was of this cohort, this would look more like data governance. And a big part of data governance, notwithstanding not the, not, not in its totality is compliance, was the GDPR.
And actually, I thought AI is a companion or sister to data, that it has such an intricate relationship with the two that it's more likely to be the, the AI governance, whatever this cashes out as, is more likely to be closest to data governance. In fact, we had to make a- And presumably that was Kevin: the, that was the thesis of the European Commission as well- Absolutely ... in putting together the AI Act. Absolutely.
Emre: Absolutely. So, so a couple, you know, we wrote a lot of papers on this. Like w- e- right at the start in 2020, I believe the EU published a, their two principles upon which they were going to, uh, govern or, uh, um, uh, legislate for AI, and it was a, a ecosystem of, uh, competitiveness and an ecosystem of trust. And I thought, okay, yeah, these are the legal frameworks and so on.
Now, the cybersecurity side, it didn't f- it didn't, you know, it didn't seem to have the kind of implications of privacy... Or sorry, of like things like bias, explainability- Mm-hmm. Right ... a lot of the concerns, the social concerns that people were.
Mm-hmm. It was more like, okay, the company needs to be protected from certain kinds of incidents, if you will. But as we've moved and, and, you know, Kevin, I'd be keen to hear your views on this. As we've moved, it feels like it's moving more and more into incident management, performance management, you know, um, rather than, "Oh my God, we've got this compliance thing that we need to tick off And, you know, we n- we need a bunch of people with legal and policy backgrounds.
Uh, so it feels to me like that's where I probab- where it is becoming more technical. Kevin: Interesting. Uh, Emre: incidentally for us, Kevin, it's, it's actually a good thing because we come- Yeah ... from foundational research.
Yeah. Um, and it turns out to be non-trivial to red team an algorithm, you know? Or to, to really think about what are the latest benchmarks, or how can we innovate in this space. But, um, but yeah, I think it's moving more into, in, in the, into the deeper tech space than I think people had anticipated, or even I had anticipated.
Kevin: Yeah. It, uh, fascinating. So that's what I was going to ask you, is what would be different for organizations taking this cybersecurity mindset or for a, a company like yours that's providing these services compared to y- where you started and many people started was this is like auditing, we've got a whole literature and an experience and infrastructure around financial auditing. So what does it look like in Emre: practice now?
I think a couple things. One thing is, um, it's not an extension of data governance. Kevin: Mm-hmm. Emre: Ergo, it's not just adding a few questions, um, to, you know, your data governance regime.
Mm-hmm. So it has a huge implication in terms of who owns this within the business. Mm-hmm. I think secondly, the skill set of the kind of people who are involved in that kind of, uh, concern is very different.
So it's not so much legal and policy, I think it's now gonna become more and more technical, or people with technical kind of literacy and so on and so forth. I think the third thing is that it dep- it, at the core of this, of, of the, let's say, the problem of governing AI is it's... Data governance is quite conservative, right? It's about conserving, protecting, you know, ensuring, um, a certain kind of, um, regime of, of protection, if you will.
Whereas AI fundamentally is about the exploitation and opportunity that it presents, right? So companies are racing to, uh, adopt AI, and almost like the winners of that race are gonna win this mar- are gonna win the market generally. So you have this immense pressure to say, "We need to use these systems." So the governance aspect is actually about saying, "How can we ensure as quickly as possible we can get the rockets off the ground?"
You know, it's not just about protecting from the outside, "Okay, let's make sure no one can access this," and so on and so forth. So I think at the, uh, it, it's much more technical. So, like, I thought the market would have landed more on documentation and reporting, so we expect... And I think a lot of vendors just emerged and did that, basically.
You know, "We provide documentation, we provide..." Yeah, sure. You know, so does everybody else, you know? But actually at the core it's saying, "Can you interrogate the systems themselves?"
Mm. So I've been looking at, um, two concrete examples. One is, uh, our thesis on human oversight- So, you know, Kevin, you remember, I'm sure, you know, it used to be the kind of panacea for everything. You know, we just have make sure you have human oversight or there's more human oversight.
Kevin: Yeah. Yeah, yeah. And- And everyone always says, "Human in the loop, that's the answer." That's it, you know?
I still hear that all the time. Emre: And then it's like, you know, we've got this elephant in the room, which is in what meaningful way can a human, um, you know, let's say with these agentic systems, meaningfully oversee such a system, you know? And then you've got this kind of deeper problem of the systems are so sophisticated and dynamic that you... What about real-time monitoring?
Like, what does it mean to actually have oversight? So that's an aspect that has been significantly eroded, I believe. You know? The second part is the actual what does it mean to assess, and I, I strongly believe that it requires a meaningful technical assessment of the system.
So not just ask - So if you remember the, for example, the data protection impact assessments would just be a form that someone filled in, you know? It's just simply, um, it's just inadequate for these systems. Kevin: What exactly are those technical capabilities? You, you mentioned red teaming before, but, but give us a little bit more detail about what that means in practice.
So, so Emre: we'll take a few examples. Um, let's say for, let's say, let's say we were inter- Let's say you looked at... There was a piece of work that we did, uh, I'm happy to say, was on explainability of agentic systems. Mm-hmm.
Really cool piece of work actually, um, Kevin. It - We won a prize. We were awarded this prize by one of the co-founders of OpenAI. Um, it was a hackathon, actually, 600 teams, uh, competed, and it was just about this really cool, we called it Agent Graph.
Uh, anyone can find it on our website, and it's part of our product. And it was basically an explainability function or tool for agentic systems. And, you know, that's the kind of stuff that really comes out of the academy. You know, it comes from foundational research.
It's bleeding edge. Uh, it's the kind of stuff which is, it's, it's not meaningful to just ask a few heuristic questions- Mm-hmm ... um, and say, you know, "Who, who made this decision?" Or, you know, "What was this output?"
It's about actually can you interrogate this system technically? So that's one example. And then there's like a whole suite of... So you've got a social definition of bias, let's say, uh, but then you've got a whole suite of technical tests to assess bias.
Of course, nothing can, can be done in, in a vacuum. Um, each one is contextualizing the other, but you do need to assess these technically. So we have a, a large open source, um, bias or debiasing library. So these are the kinds of things that I'm, I'm referring to.
We still now industriously publish in, um, leading, uh, computer science or AI conferences and journals, uh, because we think fundamentally that's probably where the, the direction of travel is, and we believe in it as well. Mm-hmm. Kevin: How much of it, the discontinuity is there with those technical responses with generative AI? Because it, for example, there's, there's a lot of bias testing algorithms which have been around now for a decade, but when you've got- Yeah ...
a non-deterministic system- Totally ... then it's become a different environment. Emre: Yeah, totally. So that's an interesting question, and if I may, Kevin, I'd love to indulge a little bit about this kind of drama with the standards- Mm-hmm ...
and regulatory bodies in our ecosystem- Mm-hmm ... writ large. So I think we were desperate as a community to say there's the, you know, somebody come out with the framework that we can just... That will be our target to assure against, right?
And I think what happened was, rightly so, um, things were moving so quickly that these bodies themselves were having to react in a way which was, you know, f- uh, tr- uh, tracking the bleeding edge of these- Mm-hmm ... uh, systems and then saying, "Okay, how do we assure it?" And I think, and I give absolute credence to this, I think that it, it's been almost an impossible task. You know, we've got o- There was a period in particular between '23 and '24 where it felt like almost every month or every other week, some huge seismic event would take place in our ecosystem, new capacity.
I mean, the generative models weren't around in anything like the way that we're talking about them when the EU did its first draft of the, of its legislation, and then all of a sudden they had to kind of respond to general purpose systems. So this is where I think the analogy with cybersecurity comes in. I think the way that cybersecurity works is that you have, like, incidents or events, and then this has this huge implication on the industry, and then there is a solution, and then that cascades quickly across- Yeah ...
um, uh, enterprises or, or, or companies writ large. And I think basically that kind of best practice becoming normalized is probably the way in which it's going to hap- it's, it's going to work. So we're gonna see practical responses to, to things that just emerge. It's gonna come- Mm-hmm ...
from the academy, it's gonna come from industry just responding, and then they're gonna become the standards. I think it's gonna actually flip on its head. I think best practice will become the de facto standard because I'm not sure we can really do this in abstract, you know? Kevin: Can I push back on you a little on that- Please, please ...
on the cyber analogy? Absolutely. Um, certainly, you know, there, there's a lot of great technical capability, uh, in cybersecurity. Um, but you know, if you look, especially you look across enterprises, so- Yeah ...
so not just, you know, military, you know, super high security kinds of contexts, um, you know, there's incident reporting. There are very good firms when you have a data breach that will go and clean things up. But a- at least what I see is, is we don't really see a race to the top. There's still many companies that are saying, "Okay, this is a cost of doing business."
Yeah. Uh, and, and they're not really adopting the best practices. Doesn't seem like the incentive structure is really there broadly enough. So i- i- if you...
You're nodding your head, so if you agree with that, how does- Yeah ... AI governance avoid that fate? Look, Emre: so I agree. I, I, I'm not...
Uh, because I think the ana- that the analogy with cybersecurity breaks down as well. Mm. Um, it's just that I was just trying to make it in comparison to data governance. Kevin: Well, I think it's a good analogy- Uh, yeah ...
it's just cybersecurity is not necessarily a positive model. Emre: Uh, yeah, sure. Then, you know, so, so, so absolutely, I agree. So I can think, uh, in my, in my mind, I can think of different categories where you would have sufficient...
You'd have market motivations and other categories where you would not have market motivations. Uh, and I think probably as we move forward, institutes like NIST and so on will probably have to work most aggressively on promo- o- on focusing on the areas where there isn't a market motivation to, for a company to, to behave in a particular or, or to govern in a particular way. But I must talk... I'm talking about really basic stuff, like for example, the hallucinations.
I can just... A m- an immense market motivation to get that right. Yeah. You know?
Uh, bias is not... People, I think, fixate on biases in demographic terms, but we're talking about, like, a algorithmic bias, like in terms of its... So we've got huge, um, market motivation to get that right, you know? So I think there is enough.
It... That's not the same in, in a lot of cybersecurity, of course it isn't. Um, you know, whereas I think there is a, there are more powerful forces in- Mm ... in our industry, uh, for sure.
Kevin: What are you seeing right now with the companies that, that you talk to or, or that work with Holistic? What's, what's motivating them? What are the biggest problems that they're bringing to you Emre: today? I think overwhelmingly...
So my, my experience, and I'd love to know if it, if it's, uh, analogous to yours, is... Or, or in, or coherent with your experience. My, my experience was that in '23 and '24, there was this explosion of pilots. Mm.
Like a huge amount of kind of speculative work. And I felt like at the start of '25, there was a kind of reckoning, um, which was, you know, what's going on here? You know, where is the return on value on these systems? You know, what are we actually seeing here?
And we're starting to see more and more of this sense of reflecting a bit more and saying, "Okay, can we actually derive value from these systems?" So there has been... I think we've moved past that pure experimentation phase. Uh, I don't think we saw the kind of productivity gains that the industry was hoping that it would have gained, but I don't think it was a wasted labor.
I don't think it was wasted labor. I think that capital that was invested, uh, both in terms of human labor and actual capital, I think it's gone into knowledge rather than production. Uh, and that, and that means, you know, it's almost like a J curve. So ultimately, I think now we're gonna start to see the, the upside of it.
So the kind of... What we're seeing generally is a far higher level of maturity in terms of awareness of this problem. Um, they've had some periods of experimentation, um, and now they're really keen to actually motivate and push this forward. Um, typ- so the distribution is usually...
Actually, a lot of companies are still using traditional, uh, machine learning systems. Then they've got a c- a cohort of generative systems that have, have been quite reasonable and worked, and now they're really experimenting with agentic systems. Um, not, not generally not in production, but, uh, you know, really keen to move in that direction. So it's - That, that's the kind of distribution.
Kevin: Yeah. How does that impact on their interest and commitment to AI governance? Oh, Emre: I think it, it's the far... There's, it...
So what I'm finding is there are different constituents that are concerned- Mm-hmm ... with AI governance. Mm-hmm. So one of the constituents is whoever owns AI transformation in the business.
All right, so fundamentally, they are responsible for undergoing... It's for motivating this transformation, and governance is just central to that. For them to have a single pane, um, view on the use of AI across the business, to be able to manage this, to be able to have, you know, a view on who's using this and who isn't using this. So fundam - Let's say that's a kind of AI governance as, uh, something that falls out of a technical track, you know, which is just about transfor- tran- AI transformation across the business.
And of course, the sec- second cohort is those very much motivated by compliance or policy, you know, and saying, "Look, we need to be, we need to be able to... We really need to be on our brief on here. You know, how are we gonna be accountable? How can we document what we're doing?
How can we ensure that we have sufficient audit trails and so on and so forth?" Anticipating the AI Act and other frameworks. So they're the two kind of general, uh, camps. I don't think, um, the persona that owns this has solidified.
Um, you know, Kevin, it's funny, I thought at the start of the decade in 2020, I thought by the middle of this decade, so last year, we would have had, we would have settled on a lot of the core issues like what does it mean to... What... Like for example, the documentation I thought would've been solved, you know? Um, and then I thought the second half of the decade would be codification of technical standards.
Kevin: Mm-hmm. Emre: But actually, the, the market or the general space is still very indeterminate. Uh- Yeah ... it's, there is - We're starting to see the kind of repeatability in personas and so on, but actually, it's still quite, um, it's still very, very dynamic.
Mm-hmm. Kevin: Well, so what do you see from this vantage point of 2026? What, what do you think now we can expect in 2030 in this area? Emre: Yeah, I do.
I think we're gonna base... I think we would have s- I think in quotes, we would've solved the problem of productivity Actually coming from these systems. So like, um, I don't think it's gonna be... I think there was, there was a thesis that Copilot was the approach, which was every single person is supercharged with their own AI.
I think it's more gonna be there's gonna be some very, very powerful use cases- Mm-hmm ... and we're gonna see immense value from them. Mm-hmm. And that companies would have solved that, like we would've been right in our, fully in our stride, um, in, in that particular path.
So I think it's probably not so much the in quote- I'm talking about in the enterprises. Um, it's not so much I think about the kind of transformation of every single person having, uh, a system, not the democratic model, as I call it. I think it's more the vanguard model of, you know, companies are gonna invest heavily in these, you know, mission-critical, probably agentic systems that are gonna motivate huge change within the business. Mm-hmm.
And the productivity will come from them. And the governance in that context is very different to the governance in the more generalized sense. So the, the more generalized sense, I think we'll continue to have the problem of shadow AI, who's using it, what are they doing, and so on, and that's it. And then on the kind of more vanguard models, let's call them, or, or, or use cases, it's gonna be very much red teaming and saying, "Okay.
How, however deep we can go, let's go." You Kevin: know? Mm-hmm. And how much can red teaming be automated and scaled, which presumably is essential, especially with go-to a-agentic models?
Emre: Yeah. So I think what we'll have is probably we're in a, it's like an arms race, uh, from... So it's funny, is it - So I guess one of the themes of this conversation is it, what, is that I thought it was gonna be the law versus tech, right? And as it's progressed, it seems to be tech versus tech.
So I think what we're seeing is basically it's gonna be the rate at which we're gonna see innovation in AI is gonna have to be mirrored and will be mirrored in the rate at which we're seeing innovation in governance. Mm-hmm. I think for the long term, you know, for the next ha-half a decade, decade, it's gonna, it, it's impo- it's gonna be impossible to play in this game without meaningful, uh, research capacity. So I just think it's gonna be...
That's, that's the m- that's the direction in which we're gonna be seeing this. It's, uh, it's... We, we have to anticipate more, uh, and significant innovation in this space. You know, I mean, we would be...
It, it's not static. You know, it's not that we just had the LLMs and then now we're, you know, responding to it. It's like, you know, every year something dramatic is taking place. So I think it's fundamentally gonna be whoever's on the core face of the research and the innovation in governance is gonna be the only ones that can meaningfully govern, uh, these systems.
Kevin: You mentioned early on that, that human oversight, uh, is, is no longer going to be, be viable or just as a, as a simple concept. So then what- Yeah ... what replaces that, that serves the needs that people were talking about human oversight for? Emre: So, so, so I do believe there will be, um, human oversight, but I think it'll be at a different level of, of abstraction.
You know? So at the moment it's, I guess, I'm trying to think about what the appropriate, um, mental model would be. It's a bit like saying, you know, we don't, we don't scan for planes using our eyes, right? Mm-hmm.
Um, we have radars, we have so on, and then we have somebody viewing the radar, right? So e- I see it a bit like, I see it like that. I think it's at a high level of, of abstraction. We'll see the kind of human oversight.
It won't be on the every single system or somehow... I think we're gonna basically move towards AI governing AI. I think effectively that's what we're going to see. We're gonna see an increase in automation and an increase in the use of, um, things like guardian agents, sentinels.
Mm-hmm. Um, you know, really quiet areas, which I thought, you know, were a lot further away are actually closer than, than one imagines. So I think that's basically the paradigm we're moving towards, AI governing AI. Kevin: Hmm.
Which presumably makes your point about the need for research capability, because obviously you've got to build AI governance systems that don't have the same problems as the AI itself. Emre: You've got the problem of infinite regress, right? Yeah. So like you need to, how does that system govern that system?
How do we govern the system governing the system? How do you have confidence in that system and so on. And there's regimes in which you can do this, but, but absolutely, I have no doubt in my mind, and it's quite extraordinary because, you know, as an entrepreneur, when you leave the academy, uh, it's a d- it's a completely different, um, domain of expertise and incentive structure. But when you're - But we, we seem to be in the most unique, uh- Mm-hmm ...
setting, Kevin, where like the f- that research is driving, foundation research is still driving the, the market or the industry, which is an extraordinary thing. You know? Um, two company, uh, from our home institution, UCL, um, two companies just raised a significant amount of money. Uh, one raised one point one billion, the other raised five hundred million.
Uh, I think, oh, you probably saw the news recently. And for la- effectively labs, you know? Mm-hmm. And it's, that was quite striking for me, but I, but it, but it kind of makes sense, you know, that, um, that still there's so much foundational work, uh, and the kind of foundational work that is taking place is coming from, from research.
Hmm. Kevin: Although, you know, we're going to need to wrap up after this. The, the, the challenge is w- when you have these, uh, applied companies especially, that have valuations where they're raising in the hundreds of millions or billions- Yeah ... of dollars, uh, then people who otherwise would be in academia and doing research increasingly are doing their work- Behind closed doors in organizations that are not necessarily publishing Yeah How do we get out of that dynamic?
Emre: Yeah, I think that's a, that's a great que- great question. I don't know. Yeah Because they are producing fantastic, there are companies producing fantastic work, and at the same time, exactly as you said, so like I noticed LeCun left, uh, Meta, right? And he's the classic example of somebody I would've expected to go back in, properly back into the academy, but instead he basically raised another billion dollars, I think it was a billion euros or whatever it was, to set up another lab.
So, um, there is a whole conversation to be had about, you know, what is taking place there. Can you, can you do meaningful, purely speculative interest-driven research in the context of a commercial, um, uh, entity? Um, I mean, um, that's a, a greater question, if you will, but I do think, I mean, the US is the best model for this, you know? Uh, we've got so many examples in the US of where, um, fundamental and foundational innovation is taking place in industry, you know?
So I'm not as, I'm not, I'm not, I'm not ... It's, this seems to be a unique dynamic, um, but I think that's an important question to ask. Mm-hmm. Absolutely.
Kevin: Well, it's, as you say, a fascinating time in many ways. We could keep going for a long time, but we need to wrap up. Emre, thanks again. It's really great to talk to Emre: you.
Thank you so much. Thanks, Kevin. Kevin: This has been The Road to Accountable AI. If you like what you're hearing, please give us a good review and check out my Substack for more insights on AI accountability.
Thank you for listening. If you wanna go deeper on AI governance, trust, and responsibility with me and other distinguished faculty of the world's top business school, sign up for the next cohort of Wharton's Strategies for Accountable AI online executive education program, featuring live interaction with faculty, expert interviews, and custom designed asynchronous content. Join fellow business leaders to learn valuable skills you can put to work in your organization. Visit execed.
wharton.upenn.edu/acai for full details. I hope to see you there.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.