The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/The Road to Accountable AI
The Road to Accountable AI artwork

Rumman Chowdhury (Humane Intelligence): The Need for Discernment

The Road to Accountable AI · 2026-05-14 · 36 min

0:00--:--

Key moments - from our scoring

Substance score

70 / 100

Five dimensions, 20 points each

Insight Density15 / 20
Originality14 / 20
Guest Caliber16 / 20
Specificity & Evidence12 / 20
Conversational Craft13 / 20

Rumman Chowdhury, CEO of Humane Intelligence Public Benefit Corporation and pioneer in responsible AI, argues that despite narratives of unchecked AI acceleration, meaningful governance mechanisms are emerging in the US through independent verification organizations (IVOs), state-level legislation like Virginia's recent mandate, and growing industry adoption of AI oversight bodies. However, she emphasizes that the assessment ecosystem remains immature - lacking standardized methods, adequate tools, sufficient model access, legal protections, and critically, enough trained evaluators. The distinction between frontier model companies (prioritizing speed and catastrophic risk reduction) and enterprise deployers (seeking assurance and governance) shapes fundamentally different evaluation needs. Enterprise clients using AI for internal processes demand context-specific assessments that generic benchmarks like MMLU cannot provide. Chowdhury is building modular tooling through Humane Intelligence to enable scalable, customizable human-in-the-loop evaluations rather than locked-in SaaS platforms. On labor market disruption, she reframes the debate: rather than an 'AI job apocalypse,' the real challenge is ensuring young workers develop discernment - the domain expertise and judgment to distinguish between adequate and exceptional AI outputs. Without deliberate workforce integration strategies and policy-backed safety nets, companies risk hollowing out entry-level hiring pipelines.

Key takeaways

  • →Independent verification organizations and state-level mandates like Virginia's represent the first meaningful regulatory mechanisms in the US, moving beyond voluntary commitments to enforceable governance standards.
  • →The AI assessment market is immature across six dimensions: tools, methods, model access, legal protections, institutions, and available evaluator talent - requiring both technical innovation and workforce development.
  • →Generic benchmarks like MMLU are useless for enterprise AI deployment; effective evaluation demands context-specific testing customized to each organization's actual use cases and stakeholders.
  • →AI won't eliminate jobs but will automate entry-level work without domain expertise requirements, making 'discernment' - the ability to spot flaws in AI outputs - the critical skill differentiator for new workers.
  • →Companies that fail to deliberately integrate AI into workforce strategy risk becoming trapped in short-term efficiency gains while undermining long-term competitive advantage and losing the next generation of trained talent.

In this episode

  1. 1The Current State of AI Regulation and Accountability
  2. 2Independent Verification Organizations and Standards Development
  3. 3Enterprise AI Governance Challenges and Scaling
  4. 4The Immature Landscape of AI Evaluation Tools and Methods
  5. 5Contextual Testing and Responsibility Between Frontier Labs and Enterprises
  6. 6Customer Risk and Conservative AI Deployment Strategies
  7. 7AI Job Displacement and the Need for Discernment in the Workforce
  8. 8Leadership Strategies for AI Integration and Workforce Development

Mentioned

Humane IntelligenceTwitterAccentureParityVirginiaFathomClaudeAnthropicWharton SchoolMITSECRumman Chowdhury

Guests

Rumman Chowdhury

Topics in this episode

agentic systemsEnterprise AI deploymentAnthropic Claude modelIndependent Verification Organizations (IVOs)Virginia AI governance legislationHumane Intelligence Public Benefit CorporationAI evaluation benchmarks (MMLU, SimpleBench)Generative AI evaluation methodsHuman-in-the-loop testingAI governance and oversight bodies

Questions this episode answers

What are independent verification organizations (IVOs) and why does Rumman Chowdhury advocate for them?

IVOs are independent third-party bodies tasked with auditing and verifying AI system compliance with standards and regulations. Chowdhury has advocated for them for years; Virginia recently became the first state to pass legislation requiring IVOs, with Dr. Gabriella Waters at Virginia State University's Center for Responsible AI now tasked with defining their structure, feasibility, and liability frameworks.

Why aren't standard AI benchmarks like MMLU useful for enterprise companies deploying AI?

Benchmarks like MMLU measure general capabilities but lack context-specific relevance; an auto manufacturer's definition of fairness or safety differs fundamentally from generic measures. Enterprise clients need customized evaluations reflecting their actual use cases, stakeholders, and risk profiles.

What is 'discernment' and why does Rumman Chowdhury see it as critical for the next generation of workers?

Discernment is the domain expertise and judgment to identify flaws, limitations, and contextual misapplications in AI-generated outputs. As AI automates entry-level work requiring no specialized expertise, workers without discernment cannot distinguish between adequate and exceptional AI work, making this skill essential for career advancement.

What is the 'augmentation trap' in workforce AI integration that Chowdhury references?

The augmentation trap, from a recent MIT paper by Sinan Aral, describes how short-term efficiency gains from deploying AI tools can mask longer-term competitive and workforce development problems if organizations don't deliberately integrate AI into workforce strategy.

How is Humane Intelligence's approach to AI evaluation tools different from typical SaaS platforms?

Rather than a proprietary SaaS platform that locks clients into specific vendors, Humane Intelligence is building a modular, open orchestration environment similar to GitHub or Hugging Face, allowing organizations to mix and match tools and avoid being dependent on vendors as technology evolves.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

15 / 20

Rumman provides substantive, actionable insights on AI governance maturity, the distinction between frontier labs and enterprise deployment, the concept of 'discernment' vs. critical thinking, and the 'Augmentation Trap' framework. However, significant portions involve conversational filler, repetition of themes, and throat-clearing that dilutes density. The insights are real but not consistently packed.

it is an incredibly immature market. Uh, I've been calling it vibes-based evals
The short-term efficiency gains of putting an AI system to do work actually has the long-term impact of cognitive decline for the worker because they're not incentivized

Originality

14 / 20

Rumman introduces genuinely fresh framings: discernment as distinct from critical thinking, the Augmentation Trap, AI-as-tool-of-mastery vs. tool-of-production in education, and the critique that intelligence benchmarks are Industrial Revolution relics. These are not standard circulating frameworks. However, some territory (IVOs, responsible AI governance, job displacement debates) is well-trodden, and the conversational structure sometimes defaults to repeating familiar responsible AI talking points.

I argue that notion of intelligence is actually antiquated 'cause it's built for the Industrial Revolution that we have now automated away
AI is often presented as a tool of production...Instead...if we think of AI as a tool of mastery

Guest Caliber

16 / 20

Rumman Chowdhury is a legitimate practitioner and leader: former Director of ML Ethics at Twitter, founder of Parity, CEO of Humane Intelligence, currently building a public benefit corporation focused on AI evaluation. She has testified to Congress and state legislatures, and works directly with enterprise clients on governance. She is not a career podcaster or pure theorist, but a working operator. Her credibility is strong, though the episode reveals she is more of a systems-thinker than a deep technical expert.

She previously served as Director of Machine Learning Ethics, Transparency, and Accountability at Twitter, founder of Parity, an algorithmic auditing platform, and Global Lead for Responsible AI at Accenture
I testified to the State of New York on the need for e- e- evaluations and audits

Specificity & Evidence

12 / 20

Rumman offers some concrete examples (Virginia's IVO legislation, the Augmentation Trap paper by Sinan Aral, Harvard seminar structure, insurance client testing) and frameworks, but largely avoids specific metrics, dollar figures, or named deployment case studies. She frequently speaks in generalities about 'companies I work with,' 'enterprise clients,' and 'the non-tech sector' without disclosing details. The discussion of governance structures remains somewhat abstract.

There's a legislation that just passed in Virginia. Uh, so what that would be is like what I've been advocating for, I don't know, 10 years at this point
There's this great paper that just came out of MIT, Sinan Aral and, and his grad student...It's called the Augmentation Trap

Conversational Craft

13 / 20

Kevin asks solid, probing questions that push Rumman to clarify her positions (e.g., on motivation without regulation, bridging governance and integration gaps, escaping the augmentation trap). However, he often accepts her answers without sharp follow-ups, allowing some hand-waving to pass unchallenged. The conversation is collegial and substantive but could be more adversarial. Kevin does not press on contradictions or demand specifics.

Is it realistic to rely on these, uh, private sector or, or public-private types of governance mechanisms without having mandates?
So is there a way to bridge those gaps? Am I, am I identifying a real problem?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Most-used words

rumman30kevin29intelligence17cause13discernment13world12hype12environment11different11tools11governance10tool10back10education9seeing9point9

Episode notes

Kevin Werbach speaks with long-time responsible AI leader Rumman Chowdhury the current environment, in which substantive standards and oversight efforts for AI are taking shape amid a larger anti-regulation wave. Chowdhury distinguishes sharply between frontier labs, where the posture is largely "AI at all costs," and the non-tech enterprises she works with, who are wrestling with how to scale governance bodies that originally reviewed single AI implementations to hundreds of systems, third-party procurement questions, and agentic workloads. She describes the current evaluations market as immature on nearly every dimension, and explains why generic benchmarks rarely translate to enterprise contexts like insurance or auto manufacturing. The conversation then turns to AI's impact on work and education. Her concern is that companies pursuing short-term efficiency by cutting entry-level hiring will face what MIT researchers Caosun and Aral call the "augmentation trap," in which workers' cognitive skills atrophy while new workers never develop them.

Full transcript

36 min

Transcribed and scored by The B2B Podcast Index.

This file was generated by Descript Kevin: Hi, I'm Kevin Werbach, Professor of Legal Studies and Business Ethics at the Wharton School of the University of Pennsylvania. For decades, I've studied emerging technologies, from broadband to blockchain. Today, AI is promising to transform our world. But AI needs accountability, mechanisms to ensure it's developed and deployed in responsible, safe, and trustworthy ways.

On this podcast, I speak Rumman: with the experts leading the charge for accountable AI. Kevin: My guest is Rumman Chowdhury, for many years one of the most influential voices working in responsible AI. She co-founded Humane Intelligence, a nonprofit that pioneered community-driven evaluations and red teaming of AI models, where she served as CEO before stepping down to focus on a new public benefit corporation bringing forward this AI evaluation work. She previously served as Director of Machine Learning Ethics, Transparency, and Accountability at Twitter, founder of Parity, an algorithmic auditing platform, and Global Lead for Responsible AI at Accenture.

Our conversation ranges from the specifics of AI assessments and governance processes in organizations to broader questions about AI's impact on jobs and education, and ultimately why we need to reframe our perspective on thinking itself. I enjoyed speaking with Rumman, and I'm sure you'll enjoy listening. Rumman, thanks so much for joining me on "The Road to Accountable AI." Rumman: Thanks for having me, Kevin.

Kevin: Uh, you were really one of the, the pioneers of this field of responsible AI, and, and obviously, you know, a lot has changed since those early days. We're, we're in a period right now where there's been a lot of pushback and pullback from some of the at least regulatory activities around, uh, AI. Uh, just in general, how do you see the environment that we're in right now in terms of these issues? Rumman: Yeah, I, I'm, I'm glad you're framing it as a, a push/pull.

I think the, uh, maybe superficial narrative is we are in an era of pure AI acceleration, and everyone's just all in all the time. But if you scratch beneath the surface, you do see actually a lot of very sensible regulations, standards, groups, and bodies being set up, uh, that are trending in a really positive direction. When I say positive direction, to me that means very applied, as in like how can we implement reasonable audits? How do we create good governance?

How do we set standards that we can adhere to? And then how do we start giving some teeth around all of these things? And I think a lot of people would be surprised to hear that, yeah, it's happening in the US as well as in other parts of the world, so I'm - what I'm not talking about right now is necessarily the EU AI Act, but actually, you know, things that I'm working on that, that we're seeing in the United States as well. Kevin: What are you seeing in the United States?

Rumman: Yeah, so there, there's a few things. Um, one is I'm part of an AI assurance industry trade group that was established by this group called Fathom, as well as some others. Um, independent verification organizations, IVOs. There's a legislation that just passed in Virginia.

Uh, so what that would be is like what I've been advocating for, I don't know, 10 years at this point. Sometimes I forget that I've been in this field for 10 years. But I would say probably the first time I testified to Congress calling for independent bodies and, and that actually has a very specific meaning, was probably about... It was pre-CoVet, so like five years ago at this point.

Now we are actually seeing Virginia as the first state passing, uh, the need for an IVO, which is an independent body. Now, my colleague Dr. Gabriella Waters, who runs the Virginia State University Center for Responsible AI, is, uh, is actually tasked by their Joint Commission on Technology to put some structure around that. They can claim things.

Uh, what does that look like? What is the feasibility? What is the responsibility? Most importantly, where does the liability lie when you find issues?

Kevin: Is it realistic to rely on these, uh, private sector or, or public-private types of governance mechanisms without having mandates? Rumman: Well, an IVO would be a mandate. That's kind of - That's what I think is so appealing about it to so many people. I think you're, you're right to sound a bit skeptical.

We've, we've had voluntary commitments and, you know, all sorts of nice glossy media moments for, for quite some time now. Uh, and I think we are very much past the point of that. I do think that, you know, we do need standards. We do need, um, some guidance.

But at the same time, work- the work of people like us will never really be done, right? When a standard is set, that means the, the water level is higher, and then our work moves up higher from that, right? Like most of us are now in the business of compliance. Uh, compliance then becomes the job of legal teams and the associated technical parties working with legal teams for people working on responsible AI, AI safety, et cetera, then the bar just gets lifted, right?

Then now we can think of the next stage of, of big problems and, and issues to deal with. All that is to say, you know, we will not magically go from a world of zero regulation to a world of perfect regulation. We will go from a world of zero regulation to some regulation, some standards, some norms, some public pressure, a wide range of things, and our job is to uplift all of them. Kevin: Mm-hmm.

Uh, and I guess part of why I'm asking these questions is just trying to get y- your sense of the environment, uh, because th- there is that pressure to go fast, and, uh, some of it is just the, you know, the, the acceleration is, but some of it is companies- Mm-hmm saying, you know, this is, we see as an imperative. In that kind of environment, uh, you know, how do people like you and others make the case to need to focus- Mm-hmm ... on these issues? Rumman: Well, I'm, I'm glad you're bringing up industry, 'cause I see industry kind of in two parts, right?

Mm-hmm. Th- there's the tech industry, frontier model companies for whom exactly it's AI at all costs, and, you know- Mm-hmm there's kind of some language around safety. It's usually very much indexed on catastrophic risk, uh, and little else, while we're seeing explicitly these companies funding anti-regulation movements- Mm-hmm ... you know, all around the US.

On the other end, a lot of the companies I work with are actually, uh, not AI native companies. They're companies that are giants in their own industry- Mm-hmm ... and it could be, you know, retail, auto manufacturing, insurance, finance. They're all interested in AI, but the way I put it is, you know, it's not their baby, so you can say it's ugly.

You're not challenging their core product if you say, "You should do this differently," or- Mm-hmm you know, "You need for better governance," when actually they're seeking it. So with most of the companies that I work with today, what we see is that they've spent the last few years building out some sort of an AI governance or oversight body. That body is usually consi- consists of widely, like, six to eight people, you know, ranging from senior leadership and everything from, like, legal to product- Mm-hmm AI, cybersecurity.

The new one I'm seeing is there's a lot more HR in the room because of so much AI being- Mm-hmm ... used in, in the workplace. Uh, that being said, you know, that, what that group has been doing is reviewing every AI im- internal AI implementation one at a time, and while it is very loved in the organization, everybody feels comfortable with that modality, right? They feel that they're getting the input they need.

They're improving their products. It's just not sustainable or scalable. So the tension that I'm helping companies navigate is we love the system that we have. We've been governing AI, and we're happy with it, but it's not gonna work for a world in which we want to have hundreds of AI systems.

Mm-hmm. It's not answering the questions of third-party procurement and liability, especially as- Mm-hmm ... we start to see a little bit more shape around these laws, and importantly, we don't know what to do with agentic systems, so- Mm-hmm ... it, it's a completely different creature.

So for them, they wanna scale, but they wanna scale their governance as well. Kevin: And what are you telling them about how to do that? Rumman: Uh, well, I'm specifically working on some products and tooling, so it's Human Intelligence Public Benefit Corporation. Um, so two things.

One is specifically it's a public benefit corporation because I do see our mission while being a for-profit to, you know, to serve individuals and organizations that want to see AI being used better in the real world. And second is, one thing I'm working on is developing like a modular tooling environment. Mm-hmm. So it's an orchestration environment that has a wide library of things that people can use.

Think of it more as like a GitHub or a Hugging Face environment, so it's not like a SaaS platform. Uh, interestingly, like obviously started building the public benefit corporation last year. Now, you know, everyone's talking about SaaS is dead, and I'm like, "Well, good thing my instinct was I'm not building a SaaS platform because everybody in responsible AI and AI safety has a SaaS platform." Like, w- and I saw this coming and, 'cause the big thing is this field moves so quickly, companies are not willing to lock into a specific vendor- Mm 'cause they worry that the tools of that vendor will be outdated as technology updates.

Like I'm tool agnostic, so we've, we've built out a way of doing at scale human-in-the-loop testing, running evaluations, and also giving feedback. One thing I've been doing a lot is customized evaluations and, and customizing evaluation output. So what I mean by that is, you know, as I, a- again, like I said, there's like the Frontier Labs world, and then there's everybody else. Yeah.

In Frontier Labs world, MMLU and SimpleBench and all those benchmarks are amazing, and everyone just talks about them all the time. I have never had a single client outside of tech say that those benchmarks are helpful for them, 'cause it's not, it's not, right? If you are an auto manufacturer, some other person's definition of contextual is not yours. Um, so I'm helping companies with using, doing testing and creating reusable assets.

So all that is to say- Mm ... like the thing I'm trying to balance is context and scale. Kevin: Mm-hmm. Yeah, and I'm, for these purposes, uh, I agree with you.

We should focus on the, the enterprise side, you know, those who are, you know, maybe doing some development, but really more deploying these foundation models and other tools. Um, and the question I have there is how well matured is this ecosystem of assessment, where if they understand that there's a, a set of issues they may have, whether it's, you know, bias or accuracy or, or whatever else, uh, are, are we at the point where the, the tools are there, uh, and it's a matter of them just developing and implementing them?

Or, or is there still a lot of customization or a lot of figuring out how you can effectively assess? Rumman: G- great question. It is an incredibly immature market. Uh, I've been calling it vibes-based evals, right?

Like, you know, we, we... And, you know, uh, and I think most people listening to, to, to this show will, will know the difference, right? Generative AI introduced probabilistic output, not statistical output. Yeah.

So pretty much roughly three years ago we were thrown this brand-new problem- Mm-hmm ... for which none of the old tools- Mm-hmm ... really worked. So we have been scrambling to do evaluations while figuring out how to do evaluations, and I, you know, I would say we spent the first few years just scrambling behind the company saying- Mm "Hey, hey, hey, we're gonna do these evaluations.

There are things you've missed or things you're not focusing on." And now, not to say we have breathing room, but we can now take the time and say, "Okay, we also have to design better methods." So, you know, there are... It, it's immature in every sense of the term.

Mm. Like, one, we don't have the tools, two, we don't have the methods, three, we don't have the model access, four, we don't have the legal protections. Mm. Five, we don't have the institutions, and six, and maybe most importantly to me, which is why it create- started the nonprofit, is- Mm ...

we just don't have enough people doing this work. It's interesting, I, I testified to the State of New York on the need for e- e- evaluations and audits, et cetera, and some of the pushback from companies is that, "Oh, there isn't really enough of a need," and I, I think I am now on, on record saying, "Really?" 'Cause I'm really tired, and I need more people doing this work. Kevin: Um, well, it's great you're doing it, but, but I, I agree with you and, uh, you know, it's certainly the reason that, that people like me are working to, you know, not necessarily train professionals, but, but at least raise the overall, uh, ambient level of knowledge among people that, that we're sending out from a place like Wharton.

Can we overcome some of those, those barriers, though? So, so for example, in terms of the, you know, the non-deterministic nature of these systems, i- is it fundamentally possible to, uh, understand things like bias and, a- when we're in this environment where you don't necessarily know what's going to happen? Rumman: Yeah. Y- you know, we- We always like to use, you know, the term in context has become- Mm-hmm ...

very, very popular. And- Mm-hmm ... and that's always been the case, right? So we can create lots of guidance and rubrics and measurement tools- Mm-hmm ...

but it is up to the evaluator to have to understand the context in which it's gonna be used. Mm-hmm. And again, kind of why, why I'm, I'm often actually very excited to work with non-frontier labs. The frontier labs are making general tools, and it's hard to, like, narrow it down to a vertical.

I love seeing the thing in action. Mm-hmm. So if I'm working with, for example, an insurance client, you know, they're specifically testing an insurance tool, and I know exactly who the stakeholders are, who the end user is, and what the situation is, then I can craft, you know, what, what the testing should be and what I should look out for. Which is, you know, very different, uh, from the task that the frontier labs have at hand, which is, like, think of every bad situation- Yeah ...

ever in any setting of anybody, and I don't envy them that. But I do think then there's almost this handoff from what... A- and there needs to be, again, this, like, industry understanding, that there is a certain amount of work the frontier labs will do and actually can do, and there's actually some work that they cannot do. So what the frontier labs cannot do is, again, for this insurance client, figure out, okay, you're gonna be using whatever, Claude- Mm-hmm ...

in this insurance setting. We need to make sure as Anthropic that it's gonna work. But that, that I would not place that as- Mm-hmm ... a- as their responsibility.

So, you know, these things are, are very in-depth, they're very contextual, and the struggle that all of us have right now, everybody's grappling with this, again, like, how do we scale? How do we scale our testing, right? Mm-hmm. There is this industry push, and I don't think it's a bad push to integrate AI into all these different things.

It really can improve efficiencies. Like, I've been using a lot of agents for all sorts of things, uh, and I, I really like it. But then h- how do we take this good work and not, not stifle it by not having good, good governance? And- Mm-hmm ...

you know, you've, you've probably heard me say this, I've said this a whole bunch of times, it's a fallacy that regulation stifles innovation. Every company I've ever worked with is actually looking for guidance on what to do- Yeah ... and how to do it right. Otherwise, they're reinventing the wheel with every new project and every single time, and it's really not, not a helpful place to be.

Kevin: But in this environment where, especially in the US, there, there's not a lot of regulation, what, what is it that motivates companies to do this investment, especially given that uncertainty you've talked about? Rumman: Yeah. And it's really just assurance. You know, and the easiest definition of AI assurance is, is your AI model doing what you promised it's doing?

So for them it's a, it's a customer- Uh, it's a customer risk. What I've noticed i- what I've noticed for the past few years is that AI company ... Uh, sorry, a lot of enterprise AI rollout is internal, so they're using LLMs for backroom, back office processes, you know, for internal stuff, and they're very, very wary to use anything customer facing, and this is a reliability and assurance problem. They're okay if the model says or does something wrong in the context of, you know, reading some paperwork.

You know, they can absorb that liability, and there's probably very low liability because it's happening in the context of their workplace, but they will not put it in front of a, a customer because for them, again, like, uh, their, their baby is their tool, their product, their service. AI, as long as AI can help advance that- Kevin: Mm-hmm ... Rumman: they're happy to use it, but if there is undue risk being introduced, they're not, and that's why we have seen very, very careful rollout.

And I think last year everybody was talking about how, you know, whatever, some, some very high percent of pilots never moved to scale, and that's really a lot of what it is. We cannot, we don't have good ways of monitoring these systems at scale quite yet. Kevin: And I know you've also been looking at, at broader issues. Um, you were part of a debate with Chris Hughes, who's actually doing a PhD in my department at Wharton.

Ooh, really? Where, where ... Yes, where the two of you were arguing, as I understand it, a- against the proposition that there would be this tremendous AI job apocalypse. And so I'm, I'm curious your, your thoughts, because obviously th- this concern about, uh, AI completely disrupting labor markets is top of mind for a lot of people.

Rumman: Yeah. Yes, so that was a, an immensely fun debate. I, I really loved it, and actually the ... like, couple of weeks later I was up at MIT, and I- Mm ...

and I visited Professor Johnson's class. Mm. And so it's funny. Uh, I also showed up late, and I'm like, "Oh, this reminds me of undergrad."

Um, so he teaches this amazing class with Gary Gensler, who used to lead the SEC- Mm ... and, and it's really about, like, AI and AI innovation. Um, but yeah, I ... and I think, you know, a lot of us were kind of making the same argument that it's not ...

W- work is not a finite pie of things, and AI is not coming to eat more of the pie than we can eat, right? It is an ever-expanding- Set of tasks and new technology creates new jobs and new tasks. The, the example I used to give is like, hey, remember in the '90s there used to be this job called a webmind. It was kind of one person who knew some HTML, and now you not only have entire web teams, you hire external firms for your branding and your photography and all this other work that you do.

Again, you know, 1998, if you had told somebody that one day this will be, you know, a very significant expense your company spends and they hire dozens of people and external firms to do this, they would just kind of think it's funny. And, and that's, that's really what I'm getting at. You know, some jobs are- Mm are going to be automated away, and it is our responsibility from a policy perspective to ensure, this is what Chris is arguing, I know Chris is working on this, from a policy perspective- Mm-hmm ...

to create the right kinds of safety nets for people. Mm-hmm. The right kinds of safeguards, and in particular, there are a couple of audiences that we have identified, and one in particular is young people graduating with- Mm degrees and entering the market. Why?

Because AI does the, uh, a, a level of work that does not require domain expertise. Mm. So what, the, my 2026 word of the year is discernment. A lot of people have been using critical thinking.

Mine is actually discernment. Mm. And what do I mean by that? What I mean is, like, I use an AI agent to design a website.

I have no design skills. That website looks amazing to me. If I showed it to a web design professional, I can guarantee you they would immediately point out little issues, problems, et cetera, right? And we're all in, in, those of us who have been in the job market, been in our- Mm-hmm ...

our professions for a bit- Yep have that, we have that discernment. Kevin: Yeah. Rumman: That, again, to the median person who does not have that expertise, the output of AI looks really interesting, it looks great, it looks wonderful, and maybe, you know, I'm fine with my super generic AI agent produced website. But if I were trying to go for something more innovative, deeper, et cetera, I would still have to call a professional.

Now, if we're not hiring young people into entry level jobs, we're gonna have to do something to ensure that they're getting upskilled and able to be the next generation with that discernment, 'cause I am not planning on having my job in 40 years. I certainly hope that there will be other people who are better than me doing it. Mm-hmm. But who are those people, and, and what are, what are we doing to ensure that- Mm ...

you know, they're safely entering the job market? Kevin: So there's a piece of that, as you say, that's, that's a policy question, but a- again, parallel to what we were talking about before, now companies need to make their decisions about what to do in this environment, and they're, they're under pressure to cut costs and they've got these technologies. So how, how should these, these enterprises sh- think about what to do in this environment, whereas, as you suggest, if they just don't hire young people, then ultimately they're shooting themselves in the foot?

Rumman: Yeah. Yeah, and, and it's almost this, like, age-old conversation about short-term benefit and, and- Mm long-term impact. Uh, and- You know, for the past few years, a lot of the AI in the workforce conversations started off with like AI literacy, which- Mm-hmm ... I don't think we ever very well defined.

And it's like, "Oh, everyone needs to know, know how to use AI." Again, I'm not quite sure what that means, but sure. And then it became, hey, I'm a, you know, my... I have an AI policy, but my employees are still using AI for things, you know, and became a little bit more of a harder governance thing.

Like we should not be taking sensitive company documents and dumping it into your version of ChatGPT and getting email outputs or whatever it is, right? And, and we did see shadow AI use. Mm-hmm. Now the big question is in workforce integration.

This is actually like very... I think it's a very exciting. It's exciting for a few reasons. One is I think the job of good leadership is to mitigate any fears about AI replacing everybody, and that is not just your job as a leader, it is the smart, sensible, strategic decision to make.

Because you should know how you want to roll out AI in your organization, and part of that is workforce integration. What do I mean by that? Mm-hmm. There's this great paper that just came out of MIT, Sinan Aral and, and his grad student, and I loved it.

I was, I was there at the Big@MIT conference talking to Sinan, and I saw his grad student, you know, uh, Michael present. It's called the Augmentation Trap. I love that framing, 'cause basically what they're saying is the short-term efficiency gains of putting an AI system to do work actually has the long-term impact of cognitive decline for the worker because they're not incentivized. Mm-hmm.

Like, they're sort of doing the cognitive offloading. They're just dumping it on the AI system. Yeah. So why is that bad?

One is you're gonna have a workforce that actually doesn't have discernment. People avoid that muscle will atrophy- Mm-hmm ... 'cause you have to use your discernment skills to retain your discernment skills. Uh, second is for new employees, they're call- it's cognitive death.

It's a muscle they've never built. Mm-hmm. Like for us it would atrophy, a- atrophy, but for them it never existed, so that's an even worse case. So maybe in the short term you're getting quote unquote cost efficiencies.

In the long term, you have a workforce that does not know how to make a decision, or worse, does not know the difference between a good and a bad decision, right? And that's really when we talk of all the language about human in the loop- Kevin: Mm-hmm ... Rumman: back to discernment, we need people who are able to tell if something was a good or a bad decision, 'cause an AI actually can give you input into it, but ultimately it's humans that best understand the system and the environment.

Kevin: Mm-hmm. No, it's a, it's a fascinating problem, and it, it strikes me similar to what we're seeing in other areas. People like, uh, my colleague, Hansun Bastani, have done research on education where- Mm-hmm ... similar phenomenon.

You, you give people an L- students an LLM, they use it, they get higher scores, then you take the LLM away, and they get lower scores- And they struggle than people who never used it, right? So- Yep. Rumman: Yep ... Kevin: so how do we escape from that trap?

Rumman: Yeah. I fell down - I'm doing some work on AI and education as well. Mm-hmm. Last year I presented to the California School Boards Association and, you know, one thing I had forgotten and was pleasantly reminded is that we've been studying education for decades and decades and decades, right?

Mm-hmm. Lots of brilliant people, um, have been statistically analyzing education. We have some of the best longitudinal data. So it's, it was a field that was ready, right?

They had the people, they had the means- Mm-hmm ... they've done this before to test AI. One of the most important things that I learned, and this is about the pedagogy of AI, and not just for kids, but I think it's for all of us. Mm-hmm.

We're all learning how to use AI. The difference is in your framing of what AI is for. So AI is often presented as a tool of production, and that is the shorthand, right? Use AI to write an essay, use AI to draft an email, use AI to make the slide look pretty.

It is producing a thing. Instead, a- and that's actually the bad place, and this is where, you know, like you said, you stick kids in front of an LLM, and then they're just like, "Write me a 300-word essay about Romeo and Juliet." They got it, you know, but then if you ask them who is Romeo, they could not possibly tell you. A- and not just that.

There is this deeper meaning of why we read these stories and books. It's not just to torture us and, you know, make us read old English and write stupid essays. It's to ex- to learn how to communicate and explain things, and we've lost all of that when we think of AI as a tool of production. Now instead, if we think of AI as a tool of mastery, so in education, it's often called the zone of proximal learning.

So when kids are learning, they're going to the next stage. If we integrate, and this is about human AI configuration, right? Mm-hmm. This is like the HCI crowd should be extremely excited about this, right?

Mm-hmm. 'Cause this is about how do we integrate this tool and how are people using it, and how are they taught how to use it? Mm-hmm. Again, this goes back to this idea of workforce integration, AI literacy, but like in the classroom.

So AI as a tool of mastery means what are the ways in which they're using an AI to achieve the next level of achievement? What that requires from the teacher and, and from the person designing the curriculum is that, again, the point of writing a 300-word essay about Romeo and Juliet was not to write it. It was to, you know, read this complex te- te- test, uh, sorry, this, uh, complex text. Mm-hmm.

Um, you know, atomize themes and viewpoints, string them all together, make a coherent argument. All of these things AI can help you do, but should not do for you. Kevin: Mm-hmm There seems to be a, a challenge though in, in, in both of these domains, and I, I'm, I'm thinking back to the, the enterprise governance kind of situation that we've been talking about, that, you know, the, the processes that are about, as you say, assurance. You've got this AI tool, you want it to do what it's supposed to do.

We don't want it to fail in these different ways. We need to govern it and restrain it. That seems to me to be a, a different kind of process and typically a different group of people than the ones who are going to look at these questions about integration, about how do we have that discernment about what are we using these tools for? So is there a way to bridge those gaps?

Am I, am I identifying a real problem? Rumman: Yeah. No, you, you are, and I think it, it goes back to what we've been talking about with the use of really any technology is it have to- Yeah ... has to be solving a real problem.

Uh, a lot of the AI hype was just excitement around the tool and its existence- Mm-hmm and not so much an assessment of whether or not it's useful. I do like to think that the hype is dying down a bit in kind of a positive way. I think everyone, a lot of people have gotten very smart about AI. It's actually quite easy to learn, you know, how to integrate with large language models, how to use them, and specifically agentic AI, like every user journey becomes more and more natural.

Kevin: Mm-hmm. Rumman: Um, so my, you know, and I'll give you a slightly spicy take. I was not particularly... Not that I wasn't excited.

I mean, I actually wasn't particularly excited about LLMs. I mean, they were- Mm-hmm. Gen AI was like a cool- Yeah ... achievement.

Kevin: Yeah. Rumman: I didn't really see, I'm like, "This is not revolutionizing the workforce because it's this like call and response thing." Mm-hmm. Uh, agentic AI however, I- Mm-hmm ...

for me has been, like I did not use gen AI stuff for anything. I didn't use LLMs for anything. With agentic AI- Yeah ... I'm doing lots of stuff.

Yeah. And again, it's because I didn't want to replace my discernment. I'm like, no, the output of this is very, very superficial and nothing I would ever say. But the things that agentic AI can do, and this is where it becomes very powerful, like all the different connectors, you can ask it to do certain tasks for you.

And, you know, again, I still have to be the one at the end of it with the discernment and, you know- Yeah ... giving it feedback. Like for anybody who's a builder, it's a very, very exciting time. I kind of feel like I did as, you know, in high school when I first started learning about the internet and all the things you could do- Yeah ...

with it. It kind of has that same vibe, you know? So, so to your point, you know, in terms of solving a problem, you have much more tools at your disposal, and I hope that there's a lot less hype around the capabilities of these tools, and we just see them as technological artifacts- Mm-hmm and things we piece together. Kevin: But let's get real.

There's a lot of hype around these tools, right? And, and, and more capabilities of agentic AI even more. I, I guess what I'm wondering is- Yeah ... you're right.

It's, it's surprisingly easy to most people to figure out how to use it, at least for some things. But as you say- Mm-hmm the discernment is not can I build a website? 'Cause now I can. The discernment is should I build a website or when, when can my, you know, simple website do what I need, or when do I still need- Rumman: Mm-hmm ...

Kevin: the human? So, so I guess I, I, I come back to how, how do we get to that point in this era? Rumman: Yeah, you're right. There is a lot of hype, and I think the hype kind of goes both ways.

There's a lot of AI safety hype or, like, crit- crita hype is one term I, I had, I had heard. Yeah. And then there's, like, you know, sort of the, the accelerationist hype as well. And like, what is the way around it?

I don't know. Yeah. Um, because I think for anybody who's new to the technology, you're being bombarded with frankly misinformation often- Mm-hmm um, of what these systems can do. A lot of it's very orchestrated.

It's very staged. Even capabilities testing is designed in these ideal world scenarios. Mm-hmm. So back to that open to debate thing that we were doing, my, i- in my, like, intro I, I said something like, "Anybody who thinks AI is coming tomorrow to replace our jobs has never tried sharing a SharePoint file outside of their organization."

And like half the, like m- like the room just like broke out. Yeah. 'Cause... And it's like, and it's dumb, but it's true.

Yeah. Like this is what AI integration means. It means your databases have to be connected. It's all the work that, like, no one says they wanna do.

Everyone wants to do the cool, exciting work of, like, the prototype building, but wow, is scaling and integrating and having all the hardness and security done right and done well the ultimate beast. So, you know, there is a lot of hype, yes, but hype doesn't translate to anything once you're trying to do something at scale. And, and, you know, and again, this is another split between Silicon Valley and everybody else. Mm-hmm.

And in Silicon Valley there's lots of bold proclamation about AI replacing w- you know, workers. Mm-hmm. And in literally every case it has either been rolled back quietly or not so quietly or, you know, in the case of let's say like Block Inject or I see a lot of people pointed out that, "I don't think this was AI. I think this was a- Yeah ...

a series of bad business decisions plus a Bitcoin tumbling, sir." Mm-hmm. And not really AI. You know, i- it's, i- it's easier to hide behind that and sound like a good business leader than we, we made some bad decisions on hiring and, and product design last quarter of- Mm-hmm ...

you know, last year. But yeah, the, I, I'm not seeing as much of that hype in the non-tech sector. Mm-hmm. There is enthusiasm, there is interest- Yeah ...

but I think people are, are quite wise to it. Kevin: And lastly, I wanna ak you, ask you about, I know that you've, uh, you're working on a book and a podcast, uh, even more broadly on, on the nature of thinking, so it'd be great if you could talk a little bit about that. Rumman: Yeah. So my podcast launches next week if I get my act together and, and get everything posted, et cetera.

And I did use an AI agent to make the website for that. Um, it did a really good job, I think, from my naive perspective. That's why I was using that example. It did- Thinking about thinking.

There's also a companion substack, and it, it's about intelligence as a social, political, scientific, and economic construct. So, you know, just taking a step back, a lot of the fear about AI taking jobs or making us dumb or whatev- whatever it is, right? A lot of the cognition and economic-based fears are actually rooted in this idea of intelligence, the measurement of intelligence. Mm.

So I went back and said, "Okay, well, where does this idea of intelligence even come from?" Mm. So, you know, we use these terms like superhuman AI. These benchmarks of intelligence were something that we as humans fabricated.

We fabricated them in the first Industrial Revolution, because at that time, we had to assess people for their, their ability to tally numbers- Mm ... and remember lists and, you know, have correspondence and act in a workforce, wake up on time, get to You know, things that today we au- will automate with AI. We don't need to tally numbers. We have, you know, our calculators and our AI systems for that.

We don't need to do a lot of these tasks. So, uh, I argue that notion of intelligence is actually antiquated 'cause it's built- Mm ... for the Industrial Revolution that we have now automated away, or we're increasingly automating away. Mm-hmm.

Great, so then what ... So that does not give me fear about humankind. What that tells me is that we have an imperfect measurement. Mm.

So how do we then arrive at a better and different measurement? And one of the, the main things I, I talk about are other fields of science and how they eval- how they evaluate intelligence, right? Mycelial intelligence, insect intelligence, even extraterrestrial intelligence. Mm.

The biggest mistake we made was we indexed artificial intelligence on humans. There was really no need to do that, and I'm not necessarily arguing for a new form of machine intelligence or whatever. What I'm saying is our measurement of human intelligence just needs to be recalibrated. So the book explores that.

Mm. The podcast is a little bit different in its setup. It's, you know, it's really not an interview-based podcast. It was actually just a recorded seminar- Mm ...

that I did at Harvard, and I purposely gathered, you know, from the Harvard community, people in bioethics, education, law, computer science, policy. So a wide range of really, really smart people and, you know, we had just weekly discussions and we just- Mm ... live recorded them. Um, so it's kind of a, it's a little, it's a bit of a master class and kind of a very different idea for a podcast, and maybe people will love it or maybe they'll hate it, and they'll think it's very confusing when 10 people are all talking.

I personally think it's really fascinating because I learned a lot from the students in the class, right? Mm. You know, I'm not a bioethicist, so learning about how, you know, they think about animal intelligence- Mm ... and the rights that animals get as a function of that intelligence from a bioethical perspective is very, very different from any world that I would ever be in, right?

Um, so, you know, it's, like I said, coming out next week if I get my act together. Kevin: For people who are listening to this who are interested in these topics more broadly, any other advice to, you know, where they should look, either, either to better understand the, you know, the more immediate practical issues about in organizations, how to appropriately manage AI, or to think more coherently about these broader questions of workforce integration and so forth? Rumman: Yeah.

Um, I mean, first I'm always happy to talk to folks. Yeah. My website is just my name, so that's easy enough, and I'm on, really my only social media at this point kind of is LinkedIn, but I'm not really that great at that. You know, I, I would very closely follow a lot of the legislation that's going on, and, and while I do agree that at the federal level there's a lot of, um, pushback on it, on regulation on around AI, it, it is interesting to kind of follow the narratives in certain states, Illinois, California, Virginia, and then to follow certain fields in which we're seeing it pop up most and, and a lot of it's, you know, like education and thinking about A- and it's sort of corollary is sort of online child safety or, you know, like safety for kids with AI systems.

Those are the things that I personally would keep an eye on, you know, a- and, and I don't think that we're gonna have any hard and heavy regulation in, in the very, very near future, but I am optimistic about, you know, independent bodies, verification becoming very, very important, uh, and, you know, and what does that But, uh, the nuts and bolts of what that means and what that entails. Kevin: Absolutely. Rumman, thanks so much Rumman: for your time. Yeah, thank you.

Kevin: This has been The Road to Rumman: Accountable AI. If you like what you're hearing, please give us a good review and check out my Substack for more insights on AI accountability. Thank you for listening. Kevin: If you wanna go deeper on AI governance, trust, and responsibility with me and other distinguished faculty of the world's top business school, sign up for the next cohort of Wharton's Strategies for Accountable AI online executive education program, featuring live interaction with faculty, expert interviews, and custom designed asynchronous content.

Join fellow business leaders to learn valuable skills you can put to work in your organization. Visit execed.wharton.upen.

edu/acai for full details. I hope to see you there.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Hermes Agent: Agents that grow with youPractical AI · on agentic systems83 / 100
  • Anthropic Code Leak: A Rare Look Inside Frontier AI | EP.52Hidden Layers · on agentic systems82 / 100
  • We Audited 40+ 8-Figure DTC Ad Accounts in 2026. Here's What's Broken and How to Fix ItD2C Diaries · on Anthropic Claude model69 / 100
  • IBM’s UK&I general manager: ‘Don’t treat AI like a science experiment’Management Today's Leadership Lessons · on agentic systems66 / 100
  • Investing in The Post Web with Morgan CreekThe Outlier Ventures Podcast · on agentic systems66 / 100
  • The AI Maximalist Playbook: James Raybould on Building at the Speed of IdeasAI for Business Leaders · on Enterprise AI deployment65 / 100

More from The Road to Accountable AI

All episodes →
  • Harish Peri (Okta): When the Thing Accessing Your Systems Has a Brain77 / 100
  • Logan Kelly (Waxell): The Accidental Agent Governance Company82 / 100
  • Nadav Cornberg (Eve Security): Interrogating Agents Before They Act83 / 100
  • Venkat Siva (Compfly): Governing Agents at the Execution Boundary95 / 100
  • Munmun De Choudhury (Georgia Tech): Conversational AI and Mental Health83 / 100
Explore the best B2B AI & Data podcasts →
All The Road to Accountable AI episodes →