The B2B Podcast Index
Index
All categories
MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
MethodologySubmit
Best of:MarketingSalesSaaSFinanceHROpsLeadershipCustomer SuccessAI & DataProductStartups & FoundersRevOpsEngineering & DevTools
An independent project byFame
SearchBest episodesGuestsInsightsMethodologySubmit a podcast
Index/Finance/Data Gurus Podcast
Data Gurus Podcast artwork

Digital Twins and the Limits of Synthetic Behavior with Olivier Toubia of Columbia Business School

Data Gurus Podcast · 2026-05-26 · 29 min

0:00--:--

Key moments - from our scoring

Substance score

63 / 100

Five dimensions, 20 points each

Insight Density12 / 20
Originality13 / 20
Guest Caliber15 / 20
Specificity & Evidence14 / 20
Conversational Craft9 / 20

Dr. Olivier Toubia of Columbia Business School leads a major research initiative examining whether digital twins - AI-generated personas trained on extensive behavioral data - can reliably predict human responses to questions they've never encountered. Working with 23 collaborators, Toubia's team created a publicly available dataset of 2,000+ respondents answering 500+ questions across four waves, then built digital twins and tested their predictive accuracy across 19 studies spanning luxury consumption, misinformation, privacy preferences, and political decisions. The findings are sobering: adding detailed behavioral data to twins provides minimal improvement over demographics alone, with correlations between human and twin answers hovering around 0.2 - equivalent to the correlation between height and IQ. While digital twins show promise for out-of-distribution predictions and accessing hard-to-reach audiences, they struggle with emotional and affect-based decisions, tend toward hyperrationality, and exhibit systematic biases (stronger pro-technology sentiment than humans). Toubia cautions against the hype cycle, drawing parallels to neuromarketing's failed promises, and argues the field must move beyond accuracy percentages to understand what's actually being lost in the other 20% of predictions.

Key takeaways

  • →Digital twins achieve only 0.2 correlation with actual human behavior - comparable to the relationship between height and IQ - indicating that human behavior is far more difficult to predict than commonly assumed.
  • →Adding 500 behavioral questions to train digital twins provides minimal predictive improvement over using demographics alone, suggesting diminishing returns from additional data collection.
  • →Digital twins exhibit systematic biases including being more pro-technology and hyper-rational than humans, which can skew business decisions even when accuracy metrics appear acceptable.
  • →Digital twins perform better on cognitive, text-based decisions and worse on emotionally-driven choices, political decisions, and video/creative content evaluation.
  • →Out-of-distribution prediction - where digital twins answer questions never seen before - represents the most promising use case, though still with limited accuracy.

In this episode

  1. 1Introduction to Digital Twins and Synthetic Data in Marketing
  2. 2Evolution from Synthetic Data to Digital Twins Technology
  3. 3The Columbia Business School Digital Twins Research Project
  4. 4Key Findings: Accuracy, Correlation, and Performance Limitations
  5. 5Out-of-Distribution Prediction and Domain-Specific Performance
  6. 6Business Applications and the Promise of Better, Faster, Cheaper Solutions
  7. 7Systematic Biases in Digital Twins and Lessons from Neuromarketing Hype
  8. 8Future Research Directions and Evidence-Based Approach to Synthetic Data

Mentioned

Olivier ToubiaColumbia Business SchoolSima VasaParadigm SampleProlificHugging FaceCornell Business School

Guests

Olivier Toubia

Topics in this episode

synthetic dataLarge language modelsgenerative AINeuromarketingDigital twinsConjoint AnalysisAdaptive Conjoint Survey DesignBehavioral PredictionColumbia Business SchoolProlific research panel

Questions this episode answers

What is the difference between synthetic data and digital twins in market research?

Synthetic data typically uses demographics and a short persona backstory to ask LLMs to simulate behavior, while digital twins are trained on extensive real data from specific individuals (demographics, psychology scales, behavioral games, etc.) to create more nuanced, heterogeneous predictions that attempt to mirror individual cognitive styles and decision-making patterns.

How well do digital twins predict individual human behavior according to Toubia's research?

Digital twins achieve only a 0.2 correlation with actual human responses - roughly equivalent to the correlation between height and IQ. Adding 500 behavioral questions provides very little improvement over using demographics alone for individual accuracy predictions.

In what scenarios do digital twins work better than in others?

Early evidence suggests digital twins perform worse on political decisions and better on human-tech interactions; they tend to perform better on text-based cognitive claims than on emotional, affect-based responses like reactions to videos or visual content, because LLMs struggle to replicate the emotional and sensory shortcuts that drive human behavior.

What are the main appeals of using synthetic data over traditional survey respondents?

Synthetic data promises to be faster, cheaper, and enables testing A/B variations on the same persona without contamination, allows questioning on sensitive topics, produces verbose detailed answers, and can provide access to hard-to-reach niche audiences like CTOs or Chief Counsels.

What systematic biases did Toubia find in digital twins compared to human respondents?

Digital twins tend to be more pro-technology and more hyperrational than humans, exhibiting stronger positive sentiment toward technology products than actual respondents would, which could systematically bias feedback on technology launches.

What our scoring noted

Our reviewer’s read on each dimension, with quotes from the episode.

Insight Density

12 / 20

The episode contains several genuinely useful research findings - the negligible benefit of deep profiling over demographics, the 0.2 correlation benchmark, and the systematic pro-technology bias - but these are spread thin across 29 minutes of conversational padding and soft transitions.

we find that adding the 500 questions that you answer doesn't really help versus just your demographics
we found a correlation of only 0.2 between the human answers and the twin answers, which is about the same as the correlation between height and iq

Originality

13 / 20

The episode draws directly from primary research rather than recycled frameworks, and the hyper-rationality finding and within-person A/B test insight are genuinely non-obvious; the neuromarketing hype-cycle comparison is a familiar trope but applied with relevant specificity.

Digital Twins tend to be a bit more hyper rational than humans because they're basically, it's an LLM that has information about you, but the ALM knows almost everything and it has unlimited reasoning capabilities
I can assign A and B to the same Persona because there would be different API calls, there would be no memory, no contamination. Like I can do true a B test within a person, which is fantastic

Guest Caliber

15 / 20

Toubia is a genuine practitioner-researcher with 25 years of relevant work, a multi-institution study of real scale, and publicly released data - not a thought-leader but someone who has actually built and tested the thing being discussed.

I got into marketing from operations research and in, uh, 1999 I started working on how do we optimize online surveys
we interviewed over 2,000 people, uh, on the Prolific...We asked them, basically we put together all the possible scales we could find from psychology, business, economics, uh, IQ questions, scales, Big five, lots of different economic games

Specificity & Evidence

14 / 20

The episode is anchored in concrete numbers - sample sizes, question counts, correlation coefficients, co-author counts, and named platforms - which is unusually grounded for a market research podcast, though some claims about domain-level performance differences remain tentative and unquantified.

Over 2,000 people, each answering over 500 questions over four different waves
we found a correlation of only 0.2 between the human answers and the twin answers, which is about the same as the correlation between height and iq

Conversational Craft

9 / 20

The host has genuine domain familiarity and lands one sharp clarifying challenge, but most questions are validating or open-ended softballs, and the guest's tentative claims about domain-specific performance are never pressed for harder evidence or practical thresholds.

It does help or it does not help?
I'm curious from your perspective, I know you're still researching everything, but do you feel like a lot of the excitement around it is really based on m. The commercial aspect?

Conversation analysis

Computed from the transcript - who did the talking, and the words that came up most.

Share of words spoken

  • Olivier Toubiaguest77%
  • Sima Vasahost19%
  • Narrator4%

Most-used words

data55human31twins23synthetic20digital17better17research15behavior15twin15predict14humans14together13exciting11based11questions11question10

Episode notes

Dr. Olivier Toubia , Glaubinger Professor of Business at Columbia Business School , joins Sima Vasa to discuss his landmark study building digital twins from over 2,000 real participants - and what the results reveal about the genuine limits of synthetic data in market research. Olivier explains why digital twins skew hyper-rational, why a 0.2 correlation with real human behavior is the honest benchmark, and why the "better, faster, cheaper" promise of synthetic data still has a question mark on "better." Olivier also covers the hybrid panel model for keeping digital twins calibrated over time, the structural advantage of within-person A/B testing with synthetic respondents, and what the neuromarketing hype cycle can teach the industry about moving faster toward evidence-based answers. KEY TAKEAWAYS 00:00 Introduction. 02:07 From operations research to marketing, Conjoint analysis and capturing human preferences with math. 03:54 The adoption cycle repeats: every new technology prompts replication before reimagination. 05:44 How synthetic data evolved from basic LLM personas to data-rich digital twins with real heterogeneity.

Full transcript

29 min

Transcribed and scored by The B2B Podcast Index.

Olivier Toubia: What's really exciting, I think, about Digital Twins is their ability to predict what we call out of distribution. So which means that they're able to predict your answers to a question that you've never seen before, that the researcher has never seen before. And so that's really exciting and we find that there is some ability to do that.

Narrator: Guided by over 25 years in the data and research industry and assisting innovators with investment banking and advisory services, Sima Vasa brings you Data Gurus, a leading market research podcast that offers actionable insights for business acceleration and value creation. Join her as she speaks with key innovators in the space to bring you up to speed with the current state and the future of data analytics and data ecosystems. This is Data Gurus need support on your market research projects. Paradigm Sample is a full service market research solutions provider. Whether you need help with questionnaire design, survey programming or online data collection, we are ready to assist. Paradigm can do as little or as much as you need, saving you time so that you can focus on insights. Learn more@paradigmsample.com welcome to another episode of Data Gurus.

Sima Vasa: I'm really excited to welcome Dr. Olivier to VA, who is a professor at Columbia Business School and has done a lot of work on AI, its impact on research, and so many other things. And I'm really excited to have this neutral voice that's based on real research to help us understand what's going on with AI and synthetic data. Welcome.

Olivier Toubia: Thanks. Thank you for having me.

Sima Vasa: Thank you. First of all, can you just talk? As you've been working on this topic for a while, it's become mainstream over the last few years, but this has been kind of a lot of your purpose over the last several years in terms of trying to understand and unpack how potentially AI can help understand human behavior. Sentiment understanding, if I understand that correctly.

Olivier Toubia: Yeah. So I got into marketing from operations research and in, uh, 1999 I started working on how do we optimize online surveys, Particular conjoint, uh, adaptive conjoint. How do you construct each question to get as much information as possible based on the previous questions and the person answered. So that was my, uh, entry point to marketing. And so at the time as, ah, someone with a math background, I was really intrigued by using math and data and models to try to capture human preferences and human behavior. And so that's how I got into marketing. And so since then I've been doing more work on trying to use AI and machine learning to really try to quantify these things that seem like they're hard to quantify. How do we choose, how do we behave? Also the creative process, how do we create and consume creative content? So that has been really my. Exciting for me. And then, uh, when Genai came about, I immediately was intrigued.

Sima Vasa: Because you were drawn to it.

Olivier Toubia: Yeah. Two of the first use cases were creativity and synthetic data, which are exactly measuring preferences and the creative process. So I was uh, drawn to it immediately. And so I started doing work more than three years ago, I guess in the, in the space and then got into synthetic data more specifically a uh, couple of years ago. It's been very interesting, very exciting area. A lot of interest these days for sure.

Sima Vasa: By the way, when you did, uh, the conjoy, I know that it's such an easy way to capture consumer response because they're not thinking about it. It's just choice of figuring out and deriving the utility in terms of what's important. So I have a keen love for conjoin and just free choice because it's just, it doesn't burden the respondent in any way. Unless the questionnaire is too long, which is a whole.

Olivier Toubia: Yeah. And there's many questions. Sometimes it gets a bit much. Yeah. So then with the adaptive, Ideally, could we ask fewer questions and get the same information?

Sima Vasa: Right.

Olivier Toubia: And so I think also what's interesting is that when the Internet came up in the 90s, the first thing we did was just let's take paper and pencil surveys and put them online. Uh, and then we said, oh, but hold on, where are we online? Maybe we can do more. We can actually do some computations in real time. So similarly here with Genai, the first instinct is let's take what we know and then just put it on Genai. But actually now we realize that we can do more than that actually. And how do we actually adapt the whole workflow based on these new technology as opposed to just copying and pasting to this new context. So I think it's a similar cycle,

Sima Vasa: I would say, isn't it interesting, like how human behavior is like that, like even from phone to paper. I understand. It's like just taking what you know and putting it in another data collection and then going online and now this, our human behaviors to like just say, can I replicate the same thing the same way but using a different tool.

Olivier Toubia: Yeah. So the same icon. You see the floppy drive.

Sima Vasa: Yeah. Yes. Yes.

Olivier Toubia: Like the attachment is the paper clip.

Sima Vasa: That's right. That's right. And I, I know I had looked you up, I found that you had done this pretty Vast research project with, I think you said, 23 other academics to, to really understand where does synthetic data where digital. And we should clarify synthetic and digital twins. Like what are we actually talking about? And I'll give you a moment to talk about that a little bit. But that's a lot of brilliant minds looking at this and trying to figure out where the use case is. So talk to me a little bit about number one, how did that come about? But number two, just a little bit of definition before we go into the findings.

Olivier Toubia: So the idea of synthetic data, I think started in political science, to my knowledge, in 2023 and initial, um, studies saying, hey, maybe we could use alms to actually simulate people's opinions and political questions. And initially it was more telling the lm, hey, behave as a Republican, you know, this age, this region, then here's a backstory and then tell me how you would vote on this question. So fascinating, uh, idea of being able to simulate human behavior. Very quickly this got picked up in marketing and so we started seeing, I think probably in 2023 or so early papers saying that maybe we can do conjoint, for example, using respondents or uh, other like maybe perceptual maps and other types of market research. So at the time it was mostly what I would call synthetic data, which is uh, usually demographics and then maybe a short backstory like a Persona and asking the LLM to behave like this synthetic, you know, person. And so that's exciting. But we quickly found that the, there's not maybe a lot of variation in the answers and maybe sometimes there could be a bit of bias. The model may be, tends to be systematically Republican or Democrat and then also some systematic biases. Uh, so then the next evolution was, well, maybe if we can get more information on a given person. So I'm going to maybe get a lot of data on SEMA and then I'm going to be able to create a twin, a digital twin of Seema that is really going to mirror her behavior. If we can do that for her and 2,000 other people, maybe actually would have a panel of digital twins. So it will still be a silicon sample, but because it will be trained based on these real humans, it will have more heterogeneity, it will capture the nuance of their behaviors to replicate your own cognitive style. And so that's when the uh, digital twins came up as a way to do synthetic data, maybe in a more accurate manner. So it's a bit more costly because you have to put together all the data. But this is the promise. And so I think the idea started floating around, uh, I would say in 2024. And so at that time we saw a lot of excitement on this and we were just curious whether it would work and how it would work. We decided that it might be a good idea to actually put together a data set that would actually allow us to create twins of people to be able to predict their behavior. Because there was no publicly available data, it would be useful for the field, for industry and academia to have resources, data, uh, to be able to learn faster together. So we don't. Instead of relying on proprietary data and I can't replicate what you're doing, what do you have a common benchmark, common data. And you probably remember neuromarketing in the mid 2000s. Similarly, like some initial studies say, oh, you can put people in an FMRI machine, look into their brains and then predict how they're going to behave. And so everyone got so excited. Hundreds of millions of dollars being spent, gurus and this. And 15 years later, well, it doesn't really work. So I think we should not replicate this. If it's going to work, we should find out quickly. If it's not going to work, we should not spend 15 years. And there's probably some way to make it work better. But for that we need to have common resources to be able to move together and learn together, as you said. So we started actually with a small group of colleagues at Columbia Business School and computer Science. And so about a year ago, a little bit more than a year ago, 2025, we put together these data sets. Uh, we interviewed over 2,000 people, uh, on the Prolific, which is, um, uh, consumer panel representative on demographics. We asked them, basically we put together all the possible scales we could find from psychology, business, economics, uh, IQ questions, scales, Big five, lots of different economic games. Uh, so we basically, let's throw in everything we know about human behavior. Let's see, uh, probably we don't need all of these, but let's just throw it anyway. And so this way we don't miss anything. So first of all, we were very, very lucky to get very generous funding from the, from the school, from Cornell Business School. They say, yeah, let's do this. Here's enough money to do this, which is quite a lot, and go for it. So we're very grateful to the school. So we're able to put together the data set. Over 2,000 people, each answering over 500 questions over four different waves. Uh, so that was the initial thing with it. So putting basically these Blueprints. And the data set is publicly available, it's on hugging face, it's available to anyone. And it's kind of unique in its breadth actually. Even if you don't care about synthetic data, there's not a lot of data sets with that many people answering such a wide range of questions. So it's pretty interesting, we think, for social science in general. And uh, so that was kind of. And so we started looking at creating digital twins and, and doing some initial tests on this data set. It looked pretty promising. And then that's when a lot of colleagues was, oh, that's really cool, you know, I want to try this in luxury consumption. How about, you know, use consumption and how about misinformation and privacy preferences? And so we said, well, let's try. So we basically ask, uh, anyone who has a study that they want to run that come within Columbia Business School. So we ended up with 23 co authors. Together we designed 19 new studies, wide range of studies. And then we went back to the same 2,000 people. We ran the studies on these 2,000 people and on their twins. So each question that a human answered, we asked exactly the same question to that digital twin. So we're able to compare across 164 different outcomes, the answer from the humans to the answer from the twins, uh, apple to apple. So really to be able to see what domains do they work better, like political domain versus personality, uh, scales. Uh, that's what the study we did. It's very exciting. But to be honest, the performance was a bit lower than we were hoping when we thought. If you look at the accuracy, so whether SEMA's answer is close to CMA's twins answer, it looks pretty good on the surface, but actually we find that adding the 500 questions that you answer doesn't really help versus just your demographics. So it was a bit disappointing.

Sima Vasa: It does help or it does not help?

Olivier Toubia: Very little, basically. No really in terms of the accuracy in our study, again, and maybe there's better ways to treat your twin. And so the data are publicly available. We hope that people will find better ways to get the twins to behave. We just use what we know and didn't read. We do find that it works a little bit better in terms of ranking people. So if I think SIMA is going to answer, you know, more than a little on this, uh, then actually we do find that we're slightly better able to predict the correlation between twins and humans, but the correlation still remains pretty low. Basically, we found a correlation of only 0.2 between the human answers and the twin answers, which is about the same as the correlation between height and iq. So it's not really very high, but here it does a little bit better when you have the full information. So that's the state of where we are. But we do find actually I, uh, think a few things from that. One of them is that actually human behavior is very hard to predict. I think, uh, we tend to sometimes be a bit arrogant and think that we can predict behavior very well because we have all these wonderful scales and models. But actually it's not.

Sima Vasa: We're complicated, we're messy and depending on the day and your interaction with somebody.

Olivier Toubia: Yeah, and definitely. And so we find that it's maybe we need to have the right expectations that we should not explain a 0.9 correlation between the twin and the humans. Because you can see these are very hard things to predict. What's really exciting, I think about digital twins is their ability to predict what we got out of distribution. So which means that they're able to predict your answers to a question that you've never seen before, that the researcher has never seen before. So I'm going to have these new concepts, these new products. I'm going to ask you for feedback even though there's no data from you. And so that's really exciting. And we find that there is some ability to do that. But again, there's only so much we can predict about human behavior based on what we find.

Sima Vasa: And there's a thesis here that says you can predict, not that you said it, but you can predict high turn, high velocity item choice because it doesn't take into emotion, values, culture, which I have an opinion on this. But it's good enough information to get from a digital twin versus when you get the highly invested decisions, uh, of a house or a loan, you know, whatever those really big decisions are. That's really harder to use digital twins for because all the complexity of human beings comes into play. But I've also heard people say no. Also those fast turn items still, even though you're not in your conscious subconsciously, that still weighs in the decisions. I'm curious what you think.

Olivier Toubia: So one of the things we're trying to understand is when does it work better. The study that I mentioned, we do have early evidence that for example, it doesn't do very well on political decisions, does a bit better on, um, human tech interactions. So we find some domains and we're doing some new tests now. And so one of the things we find also is that digital Twins tend to be a bit more hyper rational than humans because they're basically, it's an LLM that has information about you, but the ALM knows almost everything and it has unlimited reasoning capabilities. And so we find that questions that involve rational behavior or knowledge typically the twin is not able to mimic the behavior of the human. It tends to be hyper rational, uh, hyper knowledgeable. Uh, and then we have a recent test that we're running in which we compare for example the reaction to a video or commercial visual material versus uh, claims that are more text based. And we have some early evidence that actually they tend to do better on the text based claim which are a bit more cognitive because you said it's hard for the airline to capture, uh, emotional affect base. And so for example we, so we're still analyzing the data but it's still tentative. But for example, maybe if you show a video, uh, maybe the human will say oh, I don't like the voice of this person or I don't like this color. But the twin is going to look more at the rational. It's hard to replicate the thought process that makes us human with all the affect decisions, the shortcuts we take and so on and so forth. So we still don't know exactly yet what questions work better. There's indications, but I don't have a full answer. And even if I tell you this works better in this versus that, it's not going to be a huge range. It's not like you get perfect prediction in 1 and 0 in the other. But there is some variation.

Sima Vasa: As I mentioned before we got on, I was at a conference last week and everything was about AI. And you know, obviously AI for workflow makes a lot of sense. Synthetic data was another area. But a lot of the conversation was around saying synthetic is better than response data, uh, questionnaire, you know, long form feedback from a questionnaire because of the fraud and the time and everything else. What do you see as synthetic being? What's the driver here in terms of why people are so curious and excited about it?

Olivier Toubia: Yeah, it promises to give you the uh, the holy grail of better, faster, cheaper, so fast service is pretty, pretty clear. Although you still have to uh, construct the instruments, you have to buy, you know, know what you're going to ask about, you see, to analyze the data. But in terms of the collection, you know, you don't have to wait for, I mean frankly, you know, if you post something on Prolific, you can get hundreds of people in an hour. So, so it's, it is Faster, cheaper, yes. Also, it's not free. I mean there's a, uh, misconception. It's free. It's not free. It costs money to access the API and if you want to create twins, it costs money to get the intake questionnaires. But typically it is cheaper, definitely. It could be an order of magnitude cheaper. So let's check and check. The better is the question mark on the positive side, maybe indeed your synthetic respondents will give you detailed answers. They tend to be more verbose. They, uh, don't get tired. You can ask them questions that you may not be able to ask humans. That may be sensitive, ethical. One thing also very attractive is that let's say you want to do an A B test with humans. I have to assign some people to A and other people to B because if I show A and B to the same people, it will be contaminated with synthetic data. Actually I can assign A and B to the same Persona because there would be different API calls, there would be no memory, no contamination. Like I can do true a B test within a person, which is fantastic. They're able to give you their thought process, allegedly. So that's all possibly better. Now the question is whether the answers are authentic. Do they match human behavior? Can you replicate the complexity of human, uh, cognition? Can you capture the affect based responses? And so that's one where the jury is still out and it seems like there's going to need to be quite a bit of work to get it to level that. I mean, at least I would be comfortable saying. Yeah, yeah. And also I think one thing is that I often see people looking at the only the numbers. Oh, 80% accurate, good enough. Thumbs up, let's go. I think it's important to go beyond the 80% or the whatever it is first. Understand what does it mean? So 80% as opposed to what? Well, what's the benchmark? But also what happens in the other 20% because we do find that there are some systematic biases that could actually influence business decisions. We find the twins to be more pro technology than humans. And so you are launching a technology product and you're getting feedback. Oh, technology is great, we love technology. And so it's not just about the accuracy number, it's also about the content. And I think we need to go beyond just the number. But yeah, but definitely a, uh, huge promise. I think also more practically, a lot of companies are under a lot of pressure to do something with AI to show that AI forward that they're not, uh, missing out. Maybe an easy way for companies, uh, to show, hey, we're doing something with AI which are in the box. There's some cynical views on that we can save for later. Uh, but yeah, it's also attractive, it's cool, it's exciting and it's very promising. And also, by the way, another one that we haven't discussed yet is just like we discussed earlier, the first instinct is to do what we're already doing and just doing it differently. But there's also incremental things, like if you want to have a focus group with a group of CTOs or of Chief counsels. It's very hard to do this with real humans just to schedule with synthetic data. Possibly you can talk to very niche people that are very hard to reach in the real world. So it augments also the type of data you can get. So it just opens up even new opportunities.

Sima Vasa: And, uh, that might be good enough, quote unquote. Right. Just to like, message test within a hard to reach audience and get feedback. I could see that. I'm curious from your perspective, I know you're still researching everything, but do you feel like a lot of the excitement around it is really based on m. The commercial aspect? Like, there's a lot of money coming in and there's a lot of hype. Like you talked about neuromarketing is, is it similar from your perspective? And you know, what did you learn from the other part of the hype of neuromarketing to where we are today?

Olivier Toubia: Yeah, I think there's definitely. It captures our imagination when we can get into your brain. Like, it's exciting to think that maybe, um, we can use technology to control humans, to have deep insight into humans. I think that's a fantasy we have. So that's definitely exciting. And then there is also there's people who make bold claims that capture the imagination of people and that look very promising on the surface. And so that's definitely also something we see here. And also just the potential is huge. So even if there's a risk, it might not really work. It's a bet that you probably want to take. Uh, and then there's the fear of missing out that kicks in. Then people, everyone else is doing this, everyone is doing it better than we are. So we have to catch up. So that's definitely, I think, similar. I think social, uh, and psychological processes at play here, and I see it will have the same fate as neuromarketing. But I'm just saying that we should learn from the past and just take a step back and Think about, let's try, yeah, let's try to be evidence based and try to look at the data and move together, learn together and experiment together and exchange knowledge to be able to uh, get to some truth faster.

Sima Vasa: Yeah. What I really, I find interesting is that at least in the industry, so now this is not academia, but an industry. Faster, cheaper for sure, but it's also quality. Right. There's a lot of people that say there's so much fraud in survey data that you can't count on it. But it feeds so many of the models in terms of capturing and educating the models. So I find that dichotomy a little, it's interesting.

Olivier Toubia: Um, yeah, so I guess fraud means that people are using AI to answer surveys. So you're basically ah, it's going to be AI anyway so might as well get AI for 1 10th of the price. So some of the panels have put uh, in place some AI checks and in data that we collected, I think, you know, it seems like the data uh, are not just noise and garbage. I think we're seeing some patterns that suggest there is quality. I think also there's some panels where the respondents actually take pride in their work and they want to share their voice with researchers. It's hard to verify either way.

Sima Vasa: Yeah, so the other part of it is the word hydration. Like how often do you have to hydrate a model in terms of, if you follow with consent with a parent of a teen through their lifetime, the world changes. I mean we're living it now, the world changes every day. And so how often do you have to continue to get a pulse of understanding from humans and whether that's social media listening or whatever it might be to be able to predict how the digital twin is going to, going to react?

Olivier Toubia: Yeah, yeah. So the, yeah so the social media listening is definitely one alternative way to get data to train the digital twins. Um, I haven't tested it myself but I've seen some presentations about it. It's uh, very interesting. So we have to remember that the digital twin is powered by an LLM M. And uh, let's say GPT or Gemini themselves also evolve. Uh, and part of the social media data actually is going to be actually fed into the model. So if there's micro trends, some of the action will be captured by the base LLM. Now there's maybe you seema, maybe uh, you evolve as a person based on lifestyle and so on and so forth. And so that would be, we need to capture that from you directly. And so there I Mean, definitely keeping your finger on the pulse of social media. Some copy also had behavioral data, so maybe I can observe transactions over time. Your behavior allows me to refresh. I think also very promising model that's emerging is one in which we have panels of humans that actually agree to have their twins made, and they agree to also take surveys periodically. So maybe there's an intake questionnaire that, you know, to create my twin. And then once my twin has been created, maybe people will run surveys. So each Survey may be 90% will come from synthetic response, but 10% will come from humans. And so that's a way to also refresh the human data. And it also has the other benefit of then giving you some human data to calibrate the synthetic data. So you have a small sample of human data that you can use to debias maybe the model, and then use synthetic data to augment the human data. So that might be a compromise. So you're still getting the scale from synthetic data, but at the same time, you run a few human subjects, so it doesn't give you as much of a cost reduction, but maybe you're keeping the twins fresh and grounded in actual human behavior.

Sima Vasa: Yeah. So really, personal question for you. Not that personal, but, like, when you look at the digital twin data, do you feel like you're getting human responses? I'm curious.

Olivier Toubia: Yeah, I mean, if you look at the verbatim, I mean, they tend to be, um, pretty good. I think there's a bit of craft also in how you instruct the LLM to sound more human. So we're doing some research now with one of my colleagues in which we have people talk to their own twins. And so we wanted to make sure that actually the twin sounds like them, or at least like a human. So we had to tweak the system prompt quite a bit to remove some of the LLM speak, you know, and to make sure that the twins sounded human. Again, it may not, uh, have known many personal things about you, but it might sound pretty human. You can make it sound more or less human.

Sima Vasa: Well, I so appreciate you joining me today, Olivier. It's fascinating. I'm glad I stumbled upon your research, and I, uh, look forward to keeping in touch as we continue to navigate and understand the application of, uh, AI research.

Olivier Toubia: Thank you for having me.

Sima Vasa: And if people want to learn more about your research.

Olivier Toubia: Yeah. So actually, we created a digital twins lab, actually, at Columbia Business School. So I'm going to put in the chat the link here. So all the resources are there.

Sima Vasa: Great.

Olivier Toubia: And actually we are actually, um, we want to democratize this. We want to make it easy for people to test, to experiment, uh, with little cost and friction. So we're actually putting together a platform where people will be able to run studies on our twins. You can create a quantric survey, upload to the platform, ask any question. So we're still in beta phase, but we're hoping soon to be able to offer this. We won't be able to do it for free because there's API costs involved, but it would be minimal fee and you can very quickly with, for a few dollars, basically try it out yourself and then see for yourself, you know, if they sound human, if the results look good. Yeah, so, yeah. So trying to basically make it easy to learn. So that's also coming soon. Hopefully everything will be available from this website.

Sima Vasa: Awesome. Uh, I'll just put that in the show notes when we publish the.

Olivier Toubia: Yeah. Thank you. Thank you.

Sima Vasa: Ah, thank you so much for your time. I appreciate it.

Olivier Toubia: All right, well, have a good day then.

Narrator: Thank you for listening to the Data Gurus podcast brought to you by Infinity Square. If you enjoyed this episode, please leave a five star review and be sure to subscribe so you never miss an episode. Tired of market research solutions that put your project in a box? At Paradigm Sample, we approach market research solutions support with customized and consultative solutions. Whether you need help with questionnaire design, survey programming or online data collection, we're ready to assist. Let us know your needs and we can customize a solution just for you. Learn more@paradigmsample.com.

Related episodes across the Index

Other episodes covering the same guests and topics, from across The B2B Podcast Index.

  • Utilizing AI internally to iterate faster and empower smaller teams to upskill w/ Vivek Raghunathan #263The Engineering Leadership Podcast · on generative AI96 / 100
  • How Fortune 500s Use Procurement to Manage Vendor AI Training Data RightsEnterprise Tech with Fexingo · on synthetic data90 / 100
  • How Enterprise Software Buyers Now Demand a Vendor AI Training Data AuditB2B SaaS Talks with Fexingo · on synthetic data90 / 100
  • Ep 90: AI Pioneer Jürgen Schmidhuber on the State of AI TodayUnsupervised Learning with Jacob Effron · on Large language models85 / 100
  • AI Is Ready for Government. Is Government Ready?The So What from BCG · on generative AI84 / 100
  • How B2B Marketers Use AI to Personalize at Scale for EnterpriseB2B Marketing with Fexingo · on Large language models82 / 100

More from Data Gurus Podcast

All episodes →
  • #278: The Leadership Playbook Has Changed with Kelly Monahan of Beyond the Desk72 / 100
  • Why 85% of Thought Leadership Fails with Mike Nash of KS&R62 / 100
  • Diversity as a Business Imperative with Dr. Poornima Luthra75 / 100
  • From Data Silos to Insights Intelligence with Thor Olof Philogène of Stravito
  • From Gut Decisions to Causal Scenarios in Research with Jason Cohen of Simulacra Synthetic Data Studio
Explore the best B2B Finance podcasts →
All Data Gurus Podcast episodes →