
The Talent Blueprint · 2025-03-17 · 27 min
Jeff Pole, founder of Warden AI, explains how his platform tackles one of HR technology's most critical challenges: detecting and mitigating AI bias in hiring systems. Warden operates as an independent AI assurance provider, embedding continuous auditing into high-risk HR use cases like resume screening, interview intelligence tools, and automated interview systems. The core innovation addresses a fundamental regulatory paradox: GDPR and data protection laws restrict collecting demographic data, yet bias auditing requires testing against protected characteristics like gender, race, age, and disability status. Warden solves this by maintaining consent-based datasets of resumes and interview recordings that third parties can use for testing without ever training AI models on them. Sultan Sydov from Bimari and Jeff discuss how regulations like New York City's bias audit law and the EU AI Act are shifting from self-policing to independent verification. They explore why transparency - allowing candidates to see real-time audit results and opt in or out of AI-driven processes - matters as much as bias detection itself. The conversation highlights the gap between minimum compliance and genuinely fair AI, noting that human hiring processes contain unconscious bias that well-designed AI could actually improve upon.
Warden collects sensitive demographic data from volunteers with explicit consent, but only for auditing and testing purposes - not for training AI models. This allows independent third-party testing of a company's AI hiring system against those protected characteristics without the company itself having to collect or store the data.
Beyond gender and race (required by NYC's bias audit law), Warden tests against age, disability status, religion, veteran status, and others protected under civil rights law. The EU AI Act and state-level regulations are introducing additional characteristics that Warden is integrating into its platform.
Building fair AI (like Bimari does with skills-based matching instead of keywords) is different from independently verifying it works fairly across all demographics and contexts. Third-party auditing like Warden's addresses the verification gap, testing whether a well-designed system actually delivers fair outcomes for a particular company's job market and applicant pool.
Regulators focus on high-risk use cases like resume screening systems, interview intelligence tools, and automated interview scoring that can disadvantage protected groups. The NY AI Act and EU AI Act require transparency and bias auditing because these systems influence or fully replace human decisions about employment opportunities.
Transparency about AI use, audit results, and decision reasoning allows candidates and employers to make informed choices and opt out if they prefer human review. While not yet required in most regulations, transparency is seen as critical to building trust and will likely become a regulatory requirement.
Computed from the transcript - who did the talking, and the words that came up most.
Jeff Pole runs Warden, an AI assurance provider that evaluates and communicates the trustworthiness of AI systems, with a particular focus on HR technology. Their platform does continuous auditing and testing for AI risks, such as bias, and helps support businesses to be compliant - particularly around new AI regulations - and build trust through transparency. In this conversation, Jeff talks to Sultan about: The changing nature of AI and automation in HR, and the associated risks The process of continuous auditing , and testing for bias in AI models Transparency as a core theme in relevant regulations and guidelines Building trust in models and through user experiences The importance of explainability in AI-powered systems Using AI to actually reduce bias in hiring and other HR processes “There’s also a win-win opportunity here: Where the people-based processes the AI might be complementing are not necessarily what we should be emulating … We know there’s a huge amount of unconscious bias in human processes; for example, in hiring. And the AI can, if carefully designed and properly used, actually be an improvement on that status quo.” - Jeff Pole, Co-founder & CEO, Warden AI
Transcribed and scored by The B2B Podcast Index.
Speaker A: You won the best, you got the best. Absolute best of the best. The best, the best talent and mountains of it. Mountains of it. Mountains on it. Everybody's got talent. I got talent. I got talent. I got talent.
Speaker B: Hello and welcome to the Talent Blueprint, your guide to building talent first skills based organizations. The Talent Blueprint comes to you from Bimari, the AI platform powering enterprise scale workforce transformation.
Speaker C: Hello everyone. Welcome to this week's installment of the Talent Blueprint. I'm your host, Sultan Sydov and I'm super excited to have with me here today Jeff from, from Warden. I've known Jeff, uh, for some time. We've worked together on some very interesting AI projects and topics which we'll get into today. But Jeff, thank you for joining us. How are you?
Speaker D: Thank you, Sultan. Good to be here. I'm well, thanks.
Speaker C: Well, tell us a little bit about why you started Warden and what you do.
Speaker D: Sure. So background's in AI generally, but also in regulation technology. So worked for many years at a company called Onfido that uses AI to help with um, regulation kyc and what we found there is that we were using this advanced AI to help people be compliant with regulations and kind of overcome trust issues. But nobody was asking us questions about how the AI worked or whether it was trustworthy and so on to use. That is the kind of background to working on Warden. And what we do at Warden is we are an AI assurance provider, which means that we evaluate and communicate the trustworthiness of AI systems with a particular focus on HR technology. So our platform does continuously auditing and testing for AI risks like bias and uh, help support compliance and of new AI regulations and bring transparency and trust to these systems.
Speaker C: So that element of continuous uh, monitoring is what led to us first working with you guys and a bit of background for our listeners over the last couple of years. Obviously a lot's happened in AI. We've all uh, seen the ChatGPT of the world. But in the context of HR and the decisions that people make around candidates, employees, there's been very justified fears around bias and those fears aren't new. We've had actual examples of bias in HR for a lot of the last decade. But some of the things that are new apart from new AI models is regulations. There's a, uh, big movement from New York with bias and audit laws that was an extension of the way that bias was looked at in HR in general rather than specific to AI. But that started to apply to how people think about decision making and where AI comes in and because it's obviously such a new field, the question of, well, with all of these AI models moving, who's doing the auditing, how do you do it? A lot of the last decade has been about companies self policing and internal audits have been much more about process rather than AI. And uh, we were looking for a solution that would allow this to actually run in the background on all of our models, all of our tech, continuously and there were none out there until we came across you guys. I want to zoom in a little bit on what does that actually mean? What does continuous bias monitoring look like? How do you actually set it up and look at models? Tell uh, us a little bit more about uh, how you landed on this and how you developed it.
Speaker D: Sure, sure, sure, yeah. So it can get quite technical quite fast, but try and keep it somewhat high level. What we do is we work with use case specific, right? So we're looking at HR technology, we're looking at certain high risk use cases of AI, for example in resume based scoring systems, in interview intelligence notetakers or interview automated like AI interview. As you're touching on, the opportunities of using AI to bring efficiency are very great, but the risks of influencing key life decisions essentially about people, about employment opportunities are also very great and need to be taken with great care. So what we do is we become embedded into those key AI systems but still as a third party product. So we're embedded in the systems and we're doing continuous um, or at least regular testing and technical auditing of how those AI systems actually behave. And that uses in uh, part that uses our own real data sets that we've collected and ante, uh, didn't have permission to use for this purpose because obviously we're talking about sensitive data that largely most people don't want to handle, but we need to handle in order to actually be able to test for these risks, for example of bias across different demographic groups. So we have these data sets and we use them to very regularly, daily, weekly or monthly, ah, run tests against the AI system, see how it performs across these different demographic groups, male, female, different races, et cetera, and then package the results as a third party to any stakeholders that are interacting with that ah, AI system.
Speaker C: If we sort of zoom in on some human examples. So one of the early examples of bias in AI and HR systems was when companies were using resume screeners to look for certain keywords and then if a candidate didn't have that keyword in their resume, they were automatically rejected. And the question of, well do keywords and resumes have something to do with people's demographics? And are those systems that are being constructed being designed to screen out words that are more likely to come from somebody from a certain demographic group or who's male or female? But then it also applied to the broader question of should these models be making decisions by themselves? Should you ever allow somebody to be rejected without a human in the loop reviewing it? And how do you police where this works? Now? For us, before we met you guys, and we've been building Bimu for over 10 years, we've been very deliberate about building technologies that specifically reduce bias. Our mission is to create equal access to work. So there's an interesting boundary between building AI for the purpose of reducing bias, like we do, for example, by looking at people's skills and not looking at keywords. But then how do you actually prove that in practice? Because it's one thing to say, well, we built a model and we've tested it on our own data and we can see that with everything we've tested, people are, ah, more likely to end up being recommended for jobs or finding roles if they are, for example, women or if they're from underrepresented groups. But how do you know that that works for a particular company? You know what if a company has a job application process that only targeted people from a certain university who are white males and so on, you're still going to end up with a biased outcome. Even if the system was designed in a certain way and you talked about the data piece and one of the really challenging things, uh, when we think about both testing this stuff and how the regulations work, is that in order to build models that aren't biased, you actually have to look at a model not using data like gender as part of the model. You don't want to say, we're going to find out if you are a woman or from a certain background and then run the model. You actually don't capture that data. And in fact, a lot of regulations, including gdpr, involve protecting sensitive data. So you end up with this bit of a contradiction where regulation and good practice simultaneously says don't ask for gender data and racial data and so forth as nasty data, but then test against it? So how do you not capture it to avoid bias, but then have to test against it to confirm lack of bias? And that's, uh, a really big part of what you've been zooming in on. So tell us a little bit about how you approach that.
Speaker D: Yeah, totally. It's such a good Question it really touches on. We've spent the last, I guess, couple of decades as a society and in tech world getting very good and mature about the risks of data. Well, data protection, data privacy and making sure that that's in good order for many good reasons. But now that we're looking at these kind of new AI risks like AI fairness and AI decision making, we've almost shot ourselves in the foot a little bit by making it hard to use and get this all important data to actually evaluate that. So, you know, we embrace that. Like what we do is this AI auditing insurance. So we explicitly collect crowdsource essentially, uh, sensitive AI data, for example resumes, conversational interviews and so on with permission from the person, openness about why we're doing it. Permission from them to use that data just for the purpose of auditing, not for training AI systems. It's quite different to test and audit rather than train. And then we have access to this large and balanced data set that we can then use as a third party to independently test and validate the performance of these high risk systems. Whereas Sultan's mentioning it's data that most people don't want to handle anyway. So what we do helps overcome that access issue, I suppose.
Speaker C: So to use that as an example, if a company builds an AI model to for example school candidates that are applying for a job, they wouldn't necessarily know the gender or ethnicity of those candidates. But the model can then be run against your data set where with consent from people, you have collected enough information to test the model so you can independently verify whether that model actually performs without bias. In an environment where all of that data with consent can be captured and tested. And also in scenarios that are much broader than a model being applied to just one job or one set of candidates, what type of sensitive data have you ended up zooming in on? And I'll contextualize that question because one of the big challenges for people when they've been looking at sort of new regulations, and we'll come back to some of the emerging ones, like the EU has been cross testing. So you don't just say in one place we're going to test for gender and another place we're going to test for race, you have to sort of look at different ways of testing the models against all of these data sets.
Speaker D: Bias is a big topic and there's so many different forms of bias. Of course, on a tangent, there's forms of bias that are not actually undesirable, it's just how AI models work. Right. They're leaning towards certain weights. What we're talking about here is kind of unwanted or harmful bias where certain protected groups are being disadvantaged. So notable examples are, of course, sex or gender and race. So those are two probably the most commonly tested protected characteristics. The New York City regulation that we'll touch on explicitly requires auditing on those two groups. But what's interesting when you think about AI and regulation is that AI is also governed by essentially all human regulations that have existed for a long time. And so all characteristics that are protected under the Civil Rights act or the Equality act in the UK or whatever equivalent legislation in any jurisdiction also apply to AI without new AI regulations. So that includes, depending on exactly which jurisdiction you're in, religion, age, disability status, veteran status, and all of these, which become quite difficult to collect that data for. But it's something that we focus on. So we do the majority of those, let's say there's about eight, I think, characteristics that are protected, in particular the US and similarly the uk. So we've got almost all of those. And then the new AI, uh, regulations, like, uh, the EUA act and Colorado has one as well, are introducing additional ones explicitly for AI that they can't be biased against or discriminated against. And so we're in the process of adding those characteristics to our system as well.
Speaker C: One of the things that's interesting about this space is, uh, there's a big misunderstanding of both bias and regulations.
Speaker A: Right.
Speaker C: The actual laws that have come in, like the AI act, are, uh, actually not targeting the vendors or AI builders as much as they are targeting bias. They're targeting something that already exists because bias is something that happens all the time. Humans are biased. And one of the interesting things about, uh, your approach, in my mind, is you've been building a technology that can enable companies to think about auditing themselves, which is actually what we want companies to do, not just with AI, but in general. You want to look at whether there are, as the Commissioner of the eeoc, which was in charge of the regulations, uh, for the New York AI act, said recently at a Beamer event, we aren't trying to regulate AI, we're trying to regulate bias. You know, it exists with tap on the shoulder, it exists with people promoting people without going through due process. And if anything, taking technology approaches with AI and with transparency about these things should make it easier to monitor and test it and so forth. How have you thought about, uh, building an awareness of what people should actually be looking at and what the purpose of this is? And where have been some of the sort of disconnects that you've seen from vendors or from companies between what they should be looking at and auditing and what they're trying to focus on?
Speaker D: That's a good question. I think we see a range, right? We see some, whether it's vendors or organizations that are deploying AI who are at the forefront of this, who want to, partly for ethical reasons, to kind of go beyond the minimum I would put in that category, but also because they realize that there's also a kind of win win opportunity here as well, right. Where the human or the people based processes that AI might be complementing are not necessarily the be all end all that we should be emulating, right? Like we know there's a huge amount of unconscious bias in human processes, for example in hiring and the AI, ah, can if carefully designed and properly used, can actually be an improvement in on that status quo. So we see some people kind of embracing that and going above and beyond and we see others who are just doing the minimum to meet with any regulation that's coming out. Like the MSC law is the main one that's in place today. And we also see some people who are even in denial of the regulations and use the kind of gray area that they've opened up to not do too much on this front yet. But we're seeing that there's a lot more. There's a growing amount of pressure from both the regulators, but also of buyers to ask more questions about these issues. And that's putting more pressure on everyone to build AI more responsibly.
Speaker C: Well, to the point of responsibly. Part of the reason that we go beyond the minimum is because we're trying to set the bar for the industry and also independently verify that we're not just not biased, that we're doing what we plan to, which is reduce it, minimize it. And that's something, it's a mission to our companies, it's important to us. But I also think it's um, interesting disconnect between the sort of debate around what should we be worried about? Because a lot of what we should be worried about, there's a lot of things with AI isn't just about the AI models but what is it doing to people. You know, we have processes, including the UK government recently where AI has been used end to end to make hiring decisions without any people. I recently heard about a fast food chain where um, AI was so automated that candidates who got interviewed and hired by the AI didn't turn up to the job after they got hired because they thought the process was fake. And so there's much deeper questions and issues around how do you create a thoughtful experience for a company and for people? And a lot of this comes down to things like transparency. Are you being transparent with candidates about what's happening to them and what they can opt in or out of? And you guys have thought a lot about transparency. It's part of the continuous monitoring you provide in the case of what you did with us. There's a real time page that everyone can see every model and see the live recent tests of how it's performing against biases. How do you think about the role of transparency in AI and in the auditing you do and in the future of this?
Speaker D: Yeah, that's a great question. And what's. And transparency is probably one of the most common themes throughout all the regulations that we're seeing or majority of AI regulations, whether it's in the EU or the different ones in the States, is that transparency is possibly the most important thing here to overcome. Where the last few years AI has kind of gradually crept into more and more products and use cases like largely hidden from the end consumer. It may be used in kind of sales material from a vendor to a buyer, but largely end consumer has not been made to aware of that. And that's the main change that we see happening from these regulations where notices, transparency notices about even the fact that AI is used in this process, in this for example, job application and also depending on the regulation requirements to be able to see publicly on a website of that company what measures they have in place to mitigate AI risk, including bias, uh, at least in some cases like Boomerang's results, uh, of those ongoing bias audits that we publish. So I think the important thing with the transparency is that it's not just a tick box about having certain processes done. And then once it's transparent it allows people to make their own decisions. They know AI is being used, whether it's a candidate or know, uh, a business who wants to engage and they can get information into, hopefully sufficient information into, into how that works, how the AI is working and may make their own decision, for example, whether to opt in or opt out of the AI process and maybe opt out and go for a human review instead of the AI process.
Speaker C: And that opt in, opt out is not yet a regulatory requirement. But as you and I have spoken about, we both think it eventually will be. And you know, we're a company that started out in Europe and Thought a lot about consent from the day we started, even before it was a GDPR issue. But I do think this element of transparency and preferences is perhaps under discussed relative to topics of just bias. I think there's a lot of areas where perhaps people should be zooming in more on how do I actually trust what I'm using. Most of our AI models, in fact almost all of them aren't even to do with candidates. We're helping provide more insights about a job, identifying the right skills, right, uh, role descriptions. So most of the models we do don't even touch something that could bias the decision making. And then whenever a decision is involved, like assisting a search for recruiters or helping candidates find jobs, we built it to involve decisions. But in practice, people fear what they don't know while comfortably using tools like ChatGPT to write job descriptions that are likely to be very biased because they're built on the Internet and there's a lot of bias on the Internet. So how do you think about the sort of boundaries of what does it take to build trust? You've touched on that a little bit with the transparency, but creating trust through how you're building your products and what you're doing for technology providers like us and how to then use that to create the right sort of mindfulness is a big part of your philosophy. What uh, do you think about next and what's the role of trust in that?
Speaker D: Yeah, so trust is, it's the holy grail, I think of this next, I suppose, decade or decades of AI really, really burgeoning and becoming kind of embedded in our lives in the way that we're starting to see more recently. And I think it's partly because until now most AI I've worked in AI last decade of my career. Most AI systems are quite hidden and behind the scenes and doing relatively low risk things. But now the capabilities are getting so much more advanced. The opportunity to do much more high risk is use cases whereby particularly referring to anything that involves decision making, whether influencing or fully replacing a human decision ultimately for things like employment opportunity is obviously high risk. And now something that's happening in a way that wasn't so much in the past and there's many different aspects to it. The three that are focused for us are the bias, obviously, and that's where we're particularly focused and working with a number of companies on that today, whether it's fair across different demographic groups or whether it's actually systemically disadvantaging certain groups is obviously a major risk. Transparency we've touched upon where everyone can have insight and make their own judgment about what's going on. And then the kind of holy grail, I suppose of this issue is the explainability part. How is this AI system? How does it work and why? If a particular candidate, for example, gets a particular outcome from a process, from an AI, uh, model, why did they get that? What would it have taken for them to have, say, passed the acceptance rather than rejection threshold of this resume screening process or whatever it is. And that is a very technically difficult thing to truly understand and answer. But that's something that we're working on to kind of audit from a third party point of view how the model actually works using real data to test it, and then communicate how that is to stakeholders in the rest of the world to see that process and how the model uh, actually performs. I think that is the biggest piece of this. Once people can understand which attributes and the way the system is working, they can then get trust that it's working in fair and meritocratic way.
Speaker C: I uh, find this element of what creates trust in the way that these AI, uh products are uh, built. Not just the sort of underlying large language models or AI, but actually the sort of user experience. A really fascinating topic. I recently uh, tried perplexity. I think it's one of the more interesting AI search engines, kind of like Google. But that's my first question. I said what are the best ice cream flavors? Just to sort of see how it would respond to an unusual question. But as part of logging into the search engine, you obviously can say who you are as a user and it knows my location because I shared it. And so it didn't just answer the question, it um, also said, and by the way, if you want ice cream right now, here's some really good shops around you, here's the ratings they have and so forth, and here's the sort of references to where I got those. And I think it's a, um, great example of actually where the way that you build an AI interface has the opportunity to create more transparency and trust than traditional results. Right. If I search Google, not only would the results be more random because it hasn't asked me for more context transparently, but it's only doing it sort of in the background through whatever cookie tracking, but it's using the results based on number of links and those sort of traditional tools. Whereas this is more like a conversation. It's saying based on how I understand your question, here's what I suggest and then gives you Sort of references and click throughs. And I think this is maybe sort of an unexplored area for many, uh, technology companies and people yet like how do you actually create an experience of using technologies that is more consensual, interactive, um, and it's a big part of what you've sort of touched on in the HR speed consent. How do you envisage the ideal state, future of where the world you're focusing on in hr, in terms of candidates, employees, experiences can start to take into account of some of these more thoughtful ways of giving people control and more thoughtful experience.
Speaker D: You know, the kind of chatbot or natural language interface that is now becoming more common does allow for more explanation, uh, more deep explanation.
Speaker A: Right.
Speaker D: When you've just got a uh, user interface with point and click and key on screen elements, it's harder to communicate in some cases the nuance or the reasons behind something. Whereas if you have a chat based interface that's, you know, that's the core interface, right, it's more conversational. And just as we, when we give judgments as people, we tend to uh, at least do our best to explain why that is and back it up. And that is inherent in the way that a natural language interface works. So I think we're relevant when we move towards chatbot interfaces, then that will naturally come out as almost like continuous explainability, if you will, rather than just like, here's a couple of key points on the UI that explain this.
Speaker C: And this is one of the interesting things. People are more afraid of AI now because it's so powerful and becoming more powerful. But actually let's say in HR, AI 10 years ago used to be scarier because it was automation. It was like keyword tracking, rejecting. Now very few companies, and certainly no companies doing all this would do any automation really for these sort of high risk things. But instead a lot of the AI, so for example in our case is aimed at giving people choices rather than uh, eliminating choices through automation. For instance, we have as a candid experience firstly a way of saying, would you like any AI? And then if you do, it gives you a way of not just looking at jobs and ranking you, but the opposite, giving you a way of saying, what are you trying to do? Are you looking to work in a certain environment, make an impact here some teams in the company that are working on it. So it gives you a way of browsing where you want to apply and what you want to do with more context. And I think in the sort of public domain a lot of people assume that actually AI is there to eliminate humans and some technologies are doing that but, but others, ours included are doing the opposite. And I think that sort of nuance of where does the boundary of AI for efficiency, which is usually a scary concept and rightly so, versus AI for choice and experience, they shouldn't be in the same bucket, but currently because of uh, the newness of this, I think they often are. How do you think about uh, the sort of nuance You've worked with a lot of technology providers, you're now looking at spaces like interviews and others. How are you starting to think about where there's maybe some categories of how these things fall into maybe different risks or different approaches?
Speaker D: Yeah, I guess it's a bit like a ladder of different types of AI use case where you've got ones that are either not even influencing, not high risk in our book and not influencing a decision or anything like that. They're just maybe you're generating some text to be used and depending how you use it maybe there's some risk there. And then you've got ones that are partially automating or supporting decision making. So inputs, optional inputs to a person to then make a judgment. So that's obviously still very human in the loop and control but the AI is providing data points that are influencing it and then you just maybe a couple of rungs up and then you've got. Eventually the AI is maybe fully making that decision or at least for that part of the process. Maybe this is a first or second stage in a hiring process you can imagine maybe one day without too much human intervention, an AI gets you to that stage and then maybe you get the final stage within a human person interview or something. And so we see a bit of a range of different companies on those different rungs of the ladder. You tend to get smaller startups maybe taking more risks and going more advanced with how much AI is being used. Larger companies probably a bit more risk averse. I think big picture though, it's just that ladder is going to be climbed inevitably because the benefits both for efficiency as well as for some of the risks and consumer issues like bias which can be improved are too great. Right. The opportunities are too great. That I think as we mature, both on the technology itself, but on our ability to govern it and regulate it and get confidence and to trust it, we will continue moving up that ladder and I'm going to put a year based prediction right now. But in some years time I can imagine AI is used a lot more extensively to more fully if not fully, fully automate key decision making processes in things like employment and elsewhere, but it'll take some time to get there.
Speaker C: What are you most optimistic about as we look at the next chapter of both what people are building and how people are using it? What gets you excited?
Speaker D: Well, uh, certainly close to home with what we focus on right now, which is AI fairness and bias. Our optimism is that when this is unwell, as we mentioned, when this is carefully designed and properly used, we're seeing that it's less biased than typical human process. So we really think AI technology, when done correctly, can in fact improve some of the societal discrimination that we've been facing and butting against. I optimistically think that AI used in the right way will be an enabler that allows us to get to a higher level of gender fairness, sex fairness, all these other types of categories than we've been able to get to. But obviously it's a doubling story. We need to approach it with great caution. But I do think in the big picture that will help us and we'll look back on society and see that in this period, AI actually supported us to achieve that.
Speaker C: Uh, I couldn't agree more. And as a company that's been on the journey to enable fairer, uh, decisions in people choices and more equal access to work through AI for over a decade now, I'm very grateful that technologies like yours give that much more transparency and proof points and uh, allow us to be better at governance. So thank you, thank you Jeff, for that and thank you for joining today. It's been a pleasure.
Speaker D: Thanks for having me.
Speaker B: The Talent Blueprint is brought to you by Bimary, the leading AI talent platform. Learn more at bimari. Com.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.