The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/AI & Data/Absolute AI
Absolute AI artwork

Michael Nguyen | The Inside Scoop on Ground Truth Data for ML

Absolute AI · 2022-09-13 · 34 min

0:00--:--

Key moments - from our scoring

Substance score

27 / 100

Five dimensions, 20 points each

Insight Density6 / 20
Originality5 / 20
Guest Caliber8 / 20
Specificity & Evidence4 / 20
Conversational Craft4 / 20

Ground truth data forms the foundation of effective AI systems, yet its collection remains one of the most complex and human-centric challenges in machine learning. Michael Nguyen walks through what ground truth actually is - data that must be captured from real-world scenarios rather than synthesized or scraped from the web - and distinguishes it from synthetic data, which works well for documents but fails for human-centered applications like facial recognition or driver drowsiness detection. The conversation covers three critical components: understanding what data you actually need, ensuring accuracy to eliminate bias, and managing privacy and ethics throughout collection. Nguyen shares specific use cases from car manufacturers simulating driver drowsiness scenarios to healthcare applications requiring body composition data, highlighting how messaging and participant consent are as important as the technical infrastructure. He's candid about the reality that facial recognition bias can be mitigated but never fully eliminated, requiring continuous expansion of datasets across demographic groups. For B2B operators building AI products, this episode clarifies why data quality and diversity directly impact model performance and why the human elements - recruitment, messaging, and ethical frameworks - often determine project success more than the technology itself.

Key takeaways

  • →Ground truth data must be collected from real-world scenarios and cannot be fully synthesized, making human-centered datasets (faces, speech, behavior) significantly more challenging and expensive than object or environmental data.
  • →Bias in AI systems cannot be eliminated entirely but can be mitigated by collecting increasingly diverse datasets across demographic characteristics, skin tones, accents, and body types until coverage is sufficient to reduce misidentification rates.
  • →Privacy and ethics in data collection require explicit participant consent, clear messaging about data usage without disclosing specific clients, anonymization of personally identifiable information, and compliance with frameworks like GDPR and ISO data privacy standards.
  • →The three critical components of ground truth datasets are: identifying what data type you actually need, ensuring data accuracy and diversity, and establishing ethical collection processes with proper waivers and participant agreements.
  • →Scenario-based data collection (simulating real-world conditions like driver drowsiness) requires iterative refinement with detailed scripts and multiple camera angles to capture the exact conditions the ML model needs to learn.

Guests

Michael Nguyen

Topics in this episode

synthetic dataAutonomous vehiclesFacial recognitionGround truth dataBias in AIPrivacy regulations (GDPR)Data anonymizationDriver drowsiness detectionHealthcare AI applicationsISO compliance

Questions this episode answers

What is ground truth data and how is it different from synthetic data?

Ground truth data is real-world data collected from actual scenarios that cannot be readily obtained through web scraping - like street signs, facial images, or driver behavior. Unlike synthetic data (which works well for documents), ground truth for human-focused applications must be captured in real conditions because everyone looks, sounds, and behaves differently, making synthetic replication ineffective.

How do you collect data like facial recognition or driver drowsiness detection when you can't manufacture realistic scenarios?

You develop detailed scripts and controlled environments with the necessary technical setup - for example, putting someone in a car with multiple cameras to simulate falling asleep - then iterate until the client confirms the captured data matches exactly what the model needs to learn.

What privacy and ethical considerations apply when collecting human-focused ground truth data?

You must obtain explicit informed consent with signed waivers explaining what data is being collected and why (without revealing the specific client), anonymize personally identifiable information (keeping only demographics, height, weight, etc.), and ensure compliance with regulations like GDPR and ISO data privacy standards.

Can bias in facial recognition and other AI models be completely eliminated?

No, bias cannot be fully eliminated, but it can be mitigated by collecting increasingly large and diverse datasets across all demographic groups, skin tones, and facial variations - though you would theoretically need to scan every individual in a population to approach zero bias.

What types of ground truth data are easiest versus hardest to collect?

Objects and environments are easiest (images of food, street scenes) requiring no consent. Speech is moderately difficult due to recruiting speakers of specific accents and languages. Human data involving faces or body information is hardest because of privacy concerns, hesitancy among minority groups due to targeting fears, and the need for extensive demographic diversity.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

6 / 20

The episode offers a few genuinely useful distinctions (ground truth vs. synthetic data for human-interaction data, the scenario-setup methodology for capturing hard-to-get states like drowsiness) but is padded heavily with AI 101 explanations, anecdotes about learning from YouTube, and a sci-fi Jetsons tangent that consumes meaningful airtime.

ground truth is basically a collection of data, um, from real world scenarios
with synthetic data. With synthetic data you can't do that synthetic. When we talk about synthetic, what synthetic is great for is say documents

Originality

5 / 20

The episode recycles standard AI-optimism talking points (healthcare, autonomous vehicles, Amazon warehouses) with no contrarian or first-principles perspective; the one honest admission - that bias cannot be fully eliminated - is stated plainly but is not developed into any novel framework or argument.

I don't think you can eliminate biases in AI um, but what you can do is get as much data as possible
in about maybe 15, 20 years or so that uh, the cars will be totally autonomous

Guest Caliber

8 / 20

Michael has genuine hands-on experience sourcing and managing ground truth data collection projects, giving him practitioner credibility; however, his background is explicitly business development and sales rather than ML engineering or data science, and the episode is sponsored by his own employer, limiting independent depth.

my background, um, I've always been involved in technology, um, particularly starting my career in engineering actually before transitioning uh, over to sales and business development
for the past five years, I've been primarily focused on, uh, AI and ground, uh, truth data

Specificity & Evidence

4 / 20

The episode is almost entirely illustrative with no named clients, no project outcomes, no dollar figures, and no performance metrics; the most concrete figure offered is an illustrative thought experiment about scanning 300 million faces, which is rhetorical rather than evidential.

if you want to eliminate that, you may have to go out there and scan 300 million faces
like a car manufacturer, Right. Automatic, they want to detect, let's say if someone is falling asleep on the wheel

Conversational Craft

4 / 20

The host opens with a question about the guest's best learning resource (answered: YouTube), closes with a sci-fi novel prompt, and issues no substantive pushback throughout; questions are leading or purely definitional, and the sponsored-by-employer format produces an unchallenged product conversation rather than an interrogative interview.

What, what was one of the best resources um, that you came across when you were first getting into the subject matter?
if you were to write a sci fi novel about the year 2042, what would the world look like?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Speaker B72%
  • Speaker A28%

Most-used words

data85ground16truth15human11important11help11synthetic10facial10face10eliminate10different9capture9possible9part8certain8instance8

Episode notes

In this episode, Melody welcomes Michael Nguyen, VP of Global Data Practice and Partnerships at Innodata, Inc. Michael shares insights into all that AI is (and isn’t), the advances that have been made in ground truth data collection, and what it will take to overcome the biases that are still present in machine learning. GROUND TRUTH DATA IN MACHINE LEARNING Ground truth data is a collection of data that can’t be manufactured, it has to be captured. Primarily used by companies who develop products for AI/ML, there are several critical components to consider when building a ground truth data set, including identifying what type of data you actually need. Michael points out data that isn’t usually readily available, such as human interaction data, and how his team overcomes the hurdles to secure it. THE HUMAN SIDE OF AI Some data are very easy to come by, while some are significantly more difficult to secure. Any data collection that has to do with humans is more sensitive, and Michael highlights the ways that they address concerns and protect privacy for anyone who is sharing data.

Full transcript

34 min

Transcribed and scored by The B2B Podcast Index.

Speaker A: You're listening to Absolute AI Conversations with the Humans behind Artificial Intelligence where data scientists, ML researchers, startup founders and enterprise execs talk about cutting edge innovations and unique challenges posed by this new technological frontier. Tune in for interviews with leading experts to anticipate trends before they emerge. Hi, thanks for joining us on Absolute AI Conversations with the Humans Behind Artificial Intelligence. I'm your host, Melody Travers and today I'm speaking with Michael Wen, inadata's VP of Global Data Practice and Partnerships. Michael is an action driven and technology focused business development professional with over 20 years of experience delivering state of the art products and services to global enterprises. For the past five years, Michael has been focused on ground truth data for artificial intelligence and machine learning. Welcome to Absolute AI Michael.

Speaker B: Well, thank you very much.

Speaker A: Let's just dive right in. Tell uh, me a little bit about yourself and your background.

Speaker B: Yes, so as you stated, my background, um, I've always been involved in technology, um, particularly starting my career in engineering actually before transitioning uh, over to sales and business development. But uh, yeah, the last five years has been primarily focused on AI. Uh, to be honest when I first started in AI I really didn't know much about AI, so I had to read a lot about it and it's quite so interesting. But, but now after five years, you know, a lot of folks have come to me and uh, consider me as uh, some sort of an AI expert. But it takes a lot of uh, reading and understanding what it is and uh, how it's been implemented. So yeah, AI is definitely very interesting um, area and um, yeah, let's, let's get into it.

Speaker A: What, what was one of the best resources um, that you came across when you were first getting into the subject matter?

Speaker B: Uh, YouTube.

Speaker A: All right.

Speaker B: Yes. At the beginning, you know you can pretty much find anything you can in YouTube. So I was just going to. Okay, um, I'm involved in AI. What is this? So I went and look at YouTube, listen to some of the uh, speeches and people talking about that and uh, yeah that's to me, um, a lot of times reading I tend to forget. But if I'm just listening, whether I'm walking, running, ah, uh, in the gym, working out, I just listen to videos and people talking. That's how I uh, able to retain a lot of this stuff.

Speaker A: Oh, that's great. What were some of the things that uh, you thought you knew before you entered AI and you discovered were totally different once you were in the industry?

Speaker B: So AI. My, my thought of AI. Of course, you know, everything has to do with movies, right? We see movies.

Speaker A: Mhm.

Speaker B: AI, the movie. Um, you know, how AI is going to take over the world and robot robots going to dominate a lot of stuff. But as I get more involved in deep AI, it's much more than that. It's just not about that. It's about everyday life, how it actually can improve your life. You know, whether using AI in healthcare, you know, patients care, helping doctors, uh, you know, focusing on the patients, um, um, manufacturing. Right. Uh, you take a look at, look at Amazon, I mean, look at their, uh, warehouses. They are using AI and robotics to basically, you know, pick your packages and that's how you're able to get your packages the same day sometimes, right?

Speaker A: Yeah.

Speaker B: So yeah, there's just so many applications, AI anywhere, basically anything you can think of, AI, uh, can be a part of it. As we transition, you know, into this digital world. I mean, AI is it, um, like I said, everything you do, you come home, you can say, turn on the lights and the lights come on. Right. So yeah, it's just much more than robots taking over the earth and.

Speaker A: Uh-huh.

Speaker B: You know, and things like that. But that's, that was my first insight in AI was wow, that's dangerous. But no, it's actually. But of course, you know, with everything in life, um, you know, technology wise, certain aspect of it can be used in a certain order that, in a certain way that we don't like. But there's also many aspects of AI that's helped us improve our lives tremendously.

Speaker A: M. Yeah, absolutely. Tell me about your turn into AI and specifically you've been working with ground, uh, truth data. So how did you end up making that switch? And what about that particular area drew you in?

Speaker B: So primarily, um, for ground truth, uh, basically, I guess I should give a background of what ground truth is, right?

Speaker A: Absolutely, yes.

Speaker B: So ground truth is basically a collection of data, um, from real world scenarios. Right. For instance, let's just say there are data that you can, um, go to the web and collect people, images, whatever, that you can scrub the web. Um, but there are certain data that you can't. There's data that are not readily available. You actually have to go out there to get it. You have to go to the ground and get it. For instance, if somebody wants to go out there and capture, uh, street signs or because, uh, of the neighborhood, the neighborhood changes all the time. New houses are coming up, you know, um, cars, in a way. So there's always changes. Right. But you do have to go out there and actually get the Data. So that's ground truth data. It's data that can't be manufactured. Uh, it's not synthetic data either. You know, it's data. You have to go out there and find a way to capture it. Whether it's, you know, through, um, certain type of scenario that you set up to get that data, or someone just go out with the camera and just start taking pictures of, uh, street signs, cars, or whatever's out there. So that's ground truth.

Speaker A: Ground truth is particularly for training, uh, data. Or can you use it for other purposes as well?

Speaker B: Ground truth is primarily used by companies that develop products for AI ML. So it is primarily, it's a human, um, interaction that is associated with that data. For instance, facial recognition. You open your phone, well, you have a device and you have a software that detects your face. That's a form of ground truth. And to get it as accurate as possible, you actually have to capture people's faces. And then they will use that phase to train the ML to identify that's you.

Speaker A: So what are some of the critical components to consider when you're building, uh, ground truth data sets for AI?

Speaker B: So there are three components of it. Um, the first thing you want to know is what type of data that you actually need. That's the critical part. Recently I had a conversation with, um, one of the data scientists. These are the engineers that always works on the latest, greatest, uh, products out there. And I asked them, um, what do you struggle with? What is that you need the most? And every one of them always tells me that they need data. It's just not any data. It's data that they can use that is readily available to train their ML. And the data has to be as accurate as possible to eliminate certain things like biases, even AI. There's so much bias in AI that people don't understand. Go back to facial recognition. The only way really to eliminate bias is to capture everybody's faces. But that's not possible. But, but that's not possible. So the idea is to get as much data as possible to the point where you think, okay, this is good enough, right? So I would say getting that data and having the right data is the most important part of it.

Speaker A: So let's talk about some of these use cases, because you brought a couple of them up, but you talked about facial recognition healthcare. Can you give me some specific, uh, examples of types of data that again, like you said, was it readily available? Um, and how you and your team overcame some of those hurdles?

Speaker B: So being in this field I get a lot of requests for really, sometimes some really odd data. Um, and the challenge has always been, okay, um, we understand what the client needs, but then how do we get it? Maybe we have to set up some sort of environment to get that data. We have to set up some sort of scenario to get that data. Right. For instance, um, uh, I'll give an example. If a manufacturer wants to, like a car manufacturer, Right. Automatic, they want to detect, let's say if someone is falling asleep on the wheel M. Right. Using AI. Right. So how do you simulate that to capture that data, somebody falling asleep? Well, you would need to, you would have to uh, develop a scenario. You would have uh, someone sitting in a car simulating that they're falling asleep or the blink in the eyes, whatever it is, we'll write up a script, a scenario for that um, person, Person to follow. And in the meantime, you have cameras. You have cameras in the front, you got cameras on the side, camera back, capturing the interaction. So that's how they use, and that data is how they use to train, um, and detect, hey, if someone's falling asleep a lot. So that's just a good example of that right there, Right?

Speaker A: Yeah. So that, that seemed to me though, um, to get into kind of synthetic data. Right. Um, I don't know how you would.

Speaker B: No, you can't do that synthetic. With synthetic data.

Speaker A: So it's real data that you're collecting. Just the scenario is synthetic, basically. Okay, okay, I understand.

Speaker B: Well, yeah, when we talk about synthetic, what synthetic is great for is say documents. Right. So you want to uh, uh, take a look at your cell phone bill. Right. And you need, I don't know, maybe thousands of those and it's hard to get real uh, uh, builds. So what you do is you just collect maybe a few samples and then you can create duplicate uh, what that bill looks like. And that's one area where synthetic data is really good. But when it comes to human interaction, it's really hard to create a synthetic data because everyone is different, right? Everyone looks different, everyone may fall uh, asleep differently, everyone drives differently, everyone looks differently. So it's hard to create a synthetic data.

Speaker A: Um, how do you guys deal with privacy concerns and privacy regulations when collecting uh, ground truth data?

Speaker B: So that's one of the areas where I think, um, a lot of folks, um, I would say that there's a lot of companies out there that are collecting data. Mhm. And I would say that some may not be as truthful as what they're doing when it comes to collecting Data. I think you have to collect data ethically, meaning that you have to let the person know, you know, the process, what, what data's being collected. Right. And what is it being used for and let them know. You don't have to be specific as to who's, who the client is or what is the product. Uh, but you, you just have to explain them what that data is used for and why their data is important. Right. And then of course they have to sign a waiver to understand that. Uh, yes, I allow my data to be collected and it's going to be used in certain products and things like that. So if they don't sign, then, you know, you don't collect the data from them. But I know the instance where companies, uh, you know, go out there and collect data without, you know, having uh, the other person know. And I don't think that's ethical to do.

Speaker A: Yeah. And I mean, I feel like Europe is really leading the charge with this. But um, you know, now there are banners or pop ups on every website and, and I think this is great because it, it gives people who are not, you know, in this world more insight into how their movements are being tracked or their information is being pulled. Um, so that I, I think the, the public and I think the regulators are, are starting to catch up. I mean that ethical component is very important.

Speaker B: Well, that too. And then the other part of that is the privacy part. Once you collect that data, um, you're not supposed to share that data with anyone. Right. It's just uh, for that project specific, whether it's for your own product development or for the client, that data stays there and that's how uh, you would handle that. And of course all of that has to be done through again, signing a waiver and all that and then being uh, ISO, you know, compliance with data privacy. Um, and that would be it. Yeah.

Speaker A: And do you guys. And uh, anonymize, you make it anonymous? I don't know how to say that.

Speaker B: Um.

Speaker A: Yeah, um, anonymize. Is that the term? Um, so, so you may take somebody's, say it's healthcare data, you have their way, their a, their age, but you don't have the things associ it in a way that um, that somebody could determine. Oh, it's that, it's, you know, it's Melody.

Speaker B: No, no. Yes. So what, what, what, what you're describing, it is the uh, metadata of that person. Right. So do we need the name, address, person, information? No, we just need maybe the demographics or the white, Asian Hispanic. Right. The height, the weight, um, that's really. That for healthcare reasons, that's it. I mean if you're developing say um, um, I don't know, uh, a software that can capture say your BMI or body fat or things like that, you would need as much data as possible because when you're taking a picture of someone, what they do is then they just use that and then verify against all the data that's been collected and find a match. Right. Try to match it as close as possible to that data set that's already in the system. Right.

Speaker A: So you just brought up a really interesting example of something that is pretty private, right? Like somebody's body weight, their body type. I'm um, sure that's something that is, you know, ah, important to associate or find correlation or causation within healthcare. But that's something that most people, I mean talk about private data, you know, know. Um, so, so just for, for something like that, how would you approach, um, trying to get, you know, I, maybe not this example or another example, but, but data that um, that is, you know, as, as private as it gets, I guess.

Speaker B: Right. So the key to that is, is in the messaging. Right. Um, without going into much detail, you know, details about it and how we actually do it, but you really have to just sit down and explain to that person why uh, their data is important, right? Whether it's to improve a certain product, services, and at the same time it's for their own benefits too. When was the last time somebody gets a body scan for example? None. Right. So wouldn't you want to know, you know, what your current uh, BMI or body fat or heart rate or what is your current stat looks like. And part of the, say, participating something like that is that you get a free scan. So at the same time you're helping out, uh, you're helping to uh, improve technology or developing new technology that would uh, detect uh, body compensation maybe uh, a lot better than what it is today, but again it's a messaging and again they'll have to sign a waiver that they totally understand. If they don't, then they don't.

Speaker A: Absolute AI is sponsored by Innidata, a leading data engineering company. From startups to enterprise, Innidata delivers ground truth training Data and customized AI services and platforms at scale. Learn more at innidata.com. That's so interesting. So um, you know, we're, we're so focused on technology. You know, in this podcast especially we, we often are talking about algorithms and we're talking about data from a very, um, I don't know, impersonal perspective. But what you just talked about is very, ah, based on human interaction. Right. You are, you're seeking something and like you said, you have to, you have to message it in a way where people understand what they're doing, why they're doing it, why it's important. Um, and, and those are things that um, you know, are, are really soft, soft human skills making their way into um, into this very scientific area. That's so interesting.

Speaker B: Yes. I mean the data itself, it's uh, it's important to get. But you know, a lot of the challenge is especially human data, right? It's getting people to participate in the study or the survey or whatever it is that we're trying to uh, collect out of them. And again, you know, it comes down to messaging. If you have the right message and have them, uh, you know, understand why the data is important, hopefully they'll participate in the survey. You know, and there may be, you know, a, uh, gratitude or some pricing that is available to them for participating. But again, you know, that's uh, you know, I, in every project that I work with, um, you know, we take that approach very, very truthfully, uh, and let them know that what we're doing without, you know, um, having a disclosure of the client or any other personal stuff, but just more of, hey, this is why we're doing it, this is why your information is helpful. And if you want to participate, great. Please sign a paper.

Speaker A: So what kinds of data? Um, you said human focused. But, um, can you give me some more examples of um, the types of data that are really easy to come by and then the ones that are the hardest to come by that have been the most challenging but still important.

Speaker B: Okay. I would say the uh, easiest of objects, like everyday objects. Right. For instance, the companies started developing say, food. Right. People want to capture thousands of images of food, uh, set up in a dining room table with different variations of what a dinner plate looks like, what a breakfast plate looks like, and what a lunch plate looks like. That's easy. Objects. There's no human involved. There's no need to sign any waiver or anything like that. Um, so object. Quite easy environment. Quite easy. You can just go out there and with the camera and go out and take pictures. If it's in a public area, hey, you can take pictures, right? So none of that stuff. Um, the hard part is, I would say get into a uh, little bit more difficult challenges is like speech, right?

Speaker A: Okay.

Speaker B: So speech itself, because there's so many different languages, dialects, regions, accents, you know, so recruiting the right people for the right accent or the right region is a challenge enough and then get them to sign an agreement. Even though we're not using the face, we're not using at us, and we're just using the voice, it's still a challenge to some people that, hey, I don't want to do it. So getting to that. So that's challenging. Um, the other one is, uh, I mentioned earlier is scenario, uh, where you have to set up an environment. For instance, trying to simulate, um, you know, somebody falling asleep in a car. That's, that's tough to do because you have to have the right scenario, right script to do it. And you have to do it over and over and over again until the client say, yeah, that is the data. That, that's exactly what we need. Right. So that, that's challenging. Uh, and the last one, of course, we just talked about is human. Human data.

Speaker A: That's.

Speaker B: Human data is the most challenging, especially when it involves, you know, capturing the. Any part of the body is, uh, especially the face. Right. Where a, uh, lot of people hesitate to do that. Um, a lot of that has to do with the fact that especially with minority groups, um, it's a challenge because, you know, there's this tendency that, hey, are you going to target me? Are you using this for targeting m. You know, so that's, that's the challenge. Human data. That is the toughest step.

Speaker A: Yeah, I, um, was reading with facial recognition that now people, or it has been developed that you basically have like a facial fingerprint, right? So you've got your actual fingerprint. Um, but that is something that, you know, unless you've been booked for some crime, nobody has access to but your face, which is, um, each face is totally unique. Um, and the fact that there's, you know, facial, facial fingerprints. Now, like I was at the airport and. And they said, oh, we don't, we don't need your information. Just, just sit there in front of the camera. And they were using facial recognition at the airport. And honestly, I felt a little taken aback. I thought, oh, I, I guess this is, you know, I show my id, I show my face, but usually to a person. And um, yeah, I wasn't sure how I felt about that.

Speaker B: Well, I think anytime you step out of your house in a public environment, there are cameras everywhere capturing your face. So the idea is that really, in many ways, the only way really to improve that is to have, uh, as much data as possible. That's really the only way to eliminate, um, bias when it comes to facial recognition. Because right now, uh, the biases is 10 towards more darker skin type of people because it's harder to detect. Um, but to eliminate that, yes, you need as much. But really there's always gonna be some biases when it comes to facial recognition.

Speaker A: Yeah. You mentioned some of these different pitfalls for training data. What are some of those biases that can slip in there and what are ways that we can overcome those?

Speaker B: Well, certainly you cannot eliminate, I don't think you can eliminate biases in AI um, but what you can do is get as much data as possible. So when there is a comparison, if it's comparing your face to somebody else's face that it can recognize, okay, you are not that person. And that's really the only way to do that. Like I said, the only way really to eliminate biases is let's just say how many people in the U.S. uh, 300 million people in the U.S. if you want to eliminate that, you may have to go out there and scan 300 million faces. But then again, but then you have people with dark skin, light skin, brown skin, long hair, short hair, no hair. So you have to get all that, um, to ensure that, uh, the software, the AI is not mistakenly you for someone else and that really, that's the only way to eliminate that or at least, um, control that, the bias. But I don't know if there's a way you can actually eliminate 100%, to be honest with you.

Speaker A: Yeah, I think that that's, um. I like that. That's an honest answer. Um, there's a lot about mitigating bias, but you're absolutely right. Um, all of the systems are, are going to have some, some imperfections. It's just important that we know what those are. And especially in those cases where it is something that, um, you know, will have an impact on an individual's life.

Speaker B: Yeah. So, I mean, if, if it's actually mistaken, say, you know, an Asian man for an African man or vice versa, that just means that we need to go out there and get more data to get into the system to better train the ML to recognize it.

Speaker A: I, uh, want to move into some more general insights about AI. Um, I was wondering, how do you see the evolution of AI um, as we are hopefully coming out of the pandemic? Um, where are we today and what's the next crest that you're seeing in front of us?

Speaker B: The evolution. AI is continually evolving. Um, I think, uh, we see it every day already. But I think the next really in healthcare is where it's going to be taken to um, uh, the next phase. We've seen already robots doing surgery with doctors, you know, that uh, located uh, at a different uh, facility or things like that. Right. And that's going to continuous improve. And then now you're talking about uh, so driving cars, right? It's happening right now, but it's not quite there yet. I think I'm where that in about maybe 15, 20 years or so that uh, the cars will be totally autonomous and there wouldn't be any need of a person actually driving a car, that the car can drive itself, right?

Speaker A: Well that's fine by me. I don't like driving. I do hope that there's like autonomous public transport transportation though, because the focus on just you know, one person in that ride, I, I hope it expands further out. And I know there's been some big developments in trucking, you know, to, to help with some of the strain on like you were talking about uh, drivers falling asleep, which is so dange.

Speaker B: Yeah, yeah. I mean it's just not uh, for a personal uh, vehicle, you know, we're talking, you know, public transportation, buses. Right. We're talking uh, uh, I don't know, Uber or Lyft that could come pick you up and there's nobody in the car. Right. And things like. And then delivery, um, it's already happening.

Speaker A: Yeah.

Speaker B: So that's going to continue to improve. And of course uh, you know we uh, uh, drones, right? I mean AI and drones that are dropping packages in your front doorsteps and things like that, all that stuff, it's going to continue to evolve. And again that's where I think AI actually help us in many aspects in terms of our daily lives. It can help us do a lot of things that we can't do ourselves. Uh, and again I think especially in healthcare, uh, where uh, AI is going to take off and um, for instance, I think, uh, I just read that one of the companies is working on uh, AI, um, say um, you know, like for home in care. Right. Uh, if you. In the past, let's just say that you um, you fall down, right. And you can't get up. In the past, uh, you have to press a button or something like that for uh, to get uh, you know, dial 911 now through AI, maybe do Alexa to some voice command. You can say help, help or something like that. Right. And ah, it will call uh, 914 you. So things like that, um, that continues to evolve.

Speaker A: Yeah, that's great. All right, um, I'm coming around to my last question. This is a little bit of a goofy one. Take it in whatever direction you want. But, uh, if you were to write a sci fi novel about the year 2042, what would the world look like? And have the robots taken over like you thought back before you were in artificial intelligence.

Speaker B: So growing up, I watched the Jetsons a lot. Right. So there was, uh, Was that the flying. The flying, uh, spaceship.

Speaker A: Yeah. And it folded up in his, um, in his briefcase. I thought that was so cool.

Speaker B: Something like that. Yeah, I thought it was a. Yeah. So in 20 years, that's. I envision that is that we're not gonna be driving. I mean, wherever you go, you know, driverless cars, definitely. For sure. Uh, you don't have to carry your wallet anywhere. You know, everything's gonna be, uh, you know, either recognized by your face, your eyes or your palm or your fingerprints or wherever. Um, yeah, um, that's, that's, that's where the future is headed.

Speaker A: Um,

Speaker B: it's kind of scary if you really think about it. But at the same time, um, it's something to look forward to. And I know that maybe not our generation, but the next generation, they're going to take a look back at where we are and says, wow, you guys use a phone where they may not even have to use. Completely different through AI I told my

Speaker A: nieces and nephews that our phones used to be attached to the wall corded, and they're like, no way.

Speaker B: Ye. Yeah. Yeah. So I'm definitely excited for AI where it's going. And uh, like I said, there's so much things that I can do for us to, uh, help our, uh, everyday lives.

Speaker A: Well, I think the work that you're doing is really important to, again, getting that, that data, especially for groups that are underrepresented, um, which has a, uh, as you said, it's becoming so integrated and ubiquitous in our lives and it needs to be representative of, you know, all the different types of people out there.

Speaker B: Yeah, yeah, I, uh, I would agree with that. And, uh, I really enjoyed this conversation.

Speaker A: Me too. Um, let's wrap up with, uh, some call to action. How can people get a hold of you? Um, what can you help with? Anything, Anything else like that?

Speaker B: Sure. So, uh, you know, as we mentioned for the beginning, you know, for the past five years, I've been primarily focused on, uh, AI and ground, uh, truth data and how to get data for. Not just for perhaps, you know, our own use, but for client as well. Um, so if you need, if anyone needs help, uh, if you're developing an AI product or you need help with data in terms of say, um, you know, humans, uh, objects, maybe spaces or speech. Well, maybe you have a scenario that you want to, uh, um, you know, script to get the data. I can certainly help with that. Uh, you can find me on LinkedIn. If you search for Michael T. Win, I'm like, probably one of the top that pops up every time. So please do connect with me on LinkedIn. I, uh, love to connect and uh, find out more and if anything I can help. I surely would. Um, like I said, AI is very, very exciting field and the types of data that we need will continue to grow as new AI products are being introduced. That's the one thing about ground truth data. It's like the data that we capture today will not be the same as the data for tomorrow because things change, people change, environment change, everything changes. So we'll have to continuously capture that data.

Speaker A: Well, thank you so much and we will put, uh, your LinkedIn profile in the show notes and um, links to innidata. And thanks so much for joining me, Michael. This was a delight.

Speaker B: Right, well, thank you so much, Melody.

Speaker A: Thanks for tuning in. We make this program for listeners like you, so if you enjoyed this episode, share it with your community, write a review or drop us five stars. Every little bit helps spread the word. See you next. Um, time.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • How Fortune 500s Use Procurement to Manage Vendor AI Training Data RightsEnterprise Tech with Fexingo · on synthetic data90 / 100
  • How Enterprise Software Buyers Now Demand a Vendor AI Training Data AuditB2B SaaS Talks with Fexingo · on synthetic data90 / 100
  • The Evolution of Crash Test Dummies: Ensuring Road Safety with Chris O’ConnorAVL's Reimagine Mobility Podcast · on Autonomous vehicles86 / 100
  • Digital Twins and the Limits of Synthetic Behavior with Olivier Toubia of Columbia Business SchoolData Gurus Podcast · on synthetic data83 / 100
  • Your AI Agent Doesn't Sleep. Are You Ready for That? NVIDIA Answers.AI Proving Ground Podcast · on synthetic data80 / 100
  • Democratizing AI? Not Without Data Intelligence & Synthetic Data - Ari Kaplan on the Biggest Roadblocks to Scaling AIAI & Data Democratization Podcast · on synthetic data79 / 100

More from Absolute AI

All episodes →
  • Florin Tufan | Big Data on Small Businesses
  • Howie Altman | Solving Optimization Problems with Genetic Programming
  • Azmath Pasha | Evolve Your Business Processes With AI
  • Mark Kujawski | How AI Will Change The Future Of Healthcare
  • Tim Huckaby | Predicting the Future of AI
Explore the best B2B AI & Data podcasts →
All Absolute AI episodes →