
The Silicon Roundabout Podcast · 2022-01-13 · 25 min
Key moments - from our scoring
Substance score
58 / 100
Five dimensions, 20 points each
Responsible gambling technology is far more complex than detecting simple betting patterns. Karim, a technical lead, and Maris, the domain expert driving the initiative, walk through their evolution from a basic three-to-four-feature system to a sophisticated 24-feature model at Kindred. The conversation reveals why seemingly simple concepts become technically intricate: a "night play" feature must account for time zones, working night shifts, live sports events across different regions, and varying population densities. Their approach uses proportion-based thresholds adjusted per time zone rather than rigid cutoffs. Karim's team, including data scientist Abed who uses C++/AVX assembly optimization to process millions of logins in seconds, handles identity linkage across devices, IP addresses, and geolocation - working around issues like 4G relay clustering that creates false location matches. The system now monitors players continuously to build a unique dataset for supervised learning, while Maris plans version three improvements: real-time interventions for a larger customer base (0.5% to 2%), and specialized detection for "binge" gamblers who lose quickly. This is meaningful work for technical hires seeking impact in a company genuinely committed to harm prevention within gaming.
Rather than using rigid time cutoffs, Kindred analyzes the distribution of active players per time zone in 15-minute slots and identifies what constitutes "anti-social time" as the proportion of players below a specific threshold (adjusted per time zone). This captures abnormal patterns while ignoring legitimate events like US sports, which create observable spikes in activity.
The lack of a clear ground truth dataset: problematic gambling fluctuates over time, and self-exclusion data is unreliable since players may have been gambling with competitors or at land-based casinos before excluding. Kindred is collecting a year-long dataset with manual RG team review to eventually enable supervised learning.
Kindred uses two levels of linkage: a strict version based on financial patterns and behavior, and an advanced version using device fingerprints, IP addresses, and geolocation. Geolocation data is cleaned to account for 4G relay clustering, and locations are compared efficiently using grid-mapping techniques to avoid exponential computational costs.
Version three will expand real-time automated interventions from 0.5% to 2% of customers using pop-ups and other triggers, develop specialized detection for binge gamblers who lose quickly and self-exclude rapidly, and move the system entirely to real-time processing for faster identification.
Data scientist Abed used C++/AVX (assembly-level language) instead of higher-level languages like C, and implemented grid-mapping to spatially index geolocation data, so only players in the same geographic square are compared rather than all-to-all comparisons.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains genuinely substantive technical methodology - the population-baseline approach to defining antisocial play time, geolocation grid-mapping, and the C&AVX optimisation story are non-obvious insights a data practitioner would learn from. However, these are heavily diluted by host cheerleading, praise sequences, and off-topic chemistry commentary that eat significant airtime.
what we do is basically take all the players per time zone...you look at the distribution per time slot, like 15 minutes time slot, for example
the level of problematic gambling of a player over time fluctuates. So when someone self excludes the high risk of activity might not be just before the self exclusion
The reframing of 'night play' into a population-relative 'anti-social time' using per-timezone activity distributions rather than fixed clock thresholds is a genuinely non-obvious, first-principles approach. The point about self-exclusion being a lagging and noisy label for supervised learning - and the strategy of building a proprietary human-reviewed dataset to overcome it - is also fresh and specific to the domain.
instead of nighttime we call it anti social time
if the first version you had like um, 0.5 of our customers approach, we're looking into increasing this to 2% so even lowering the risk, so accepting more people, but in an automated intervention
Karim and Maris are genuine practitioners who built and own this system at Kindred - Maris drives the domain specification and Karim leads the technical implementation. They are credible mid-level operators with real depth, though neither is a senior industry leader, and the episode is partly a recruitment pitch rather than a pure knowledge-sharing exercise.
we went from what, three, four features in a model into 24
we collaborate with different data teams. Uh, we are dedicated to leco. We have also, um, a global data science team based in London
The episode offers concrete numbers - 15-minute time slots, 5% activity thresholds, movement from 0.5% to 2% customer intervention rate, 20 - 25 minutes down to under 10 seconds for a million logins, 24 features vs original 3 - 4, and approximately one year of proprietary labelled data. Business-impact outcomes (e.g. harm reduction rates, model accuracy metrics) are notably absent, capping the score.
for a million logins, it was taking less than 10 seconds
if the first version you had like um, 0.5 of our customers approach, we're looking into increasing this to 2%
The host occasionally asks good follow-up questions that unlock real technical detail, such as pushing for a walkthrough of the night-play scenario. However, the conversation is repeatedly derailed by effusive praise of the guests, team-chemistry commentary, and recruitment messaging, leaving many technical threads underpursued and no meaningful pushback on any claim.
Give us a walkthrough. This is very interesting. I understand how complicated this sounds. Absolutely. I can only imagine, you know, the work that's been put into it
I absolutely love the chemistry between you guys. I've always thought that the chemistry between colleagues is, you know, a huge part, if not the main part, as to why work, uh, is delivered well
Computed from the transcript - who did the talking, and the words that came up most.
In a previous episode we spoke with Marin Catania, Head of Responsible Gambling & Research at Kindred, about how, and why, a gambling company has pioneered the efforts towards responsible gambling. In this episode, the focus goes to Karim Chikh, Head of LeCo analytics at Kindred, who discusses the highly complex tech that makes up the sophisticated system used by Kindred. You can find Karim on - You can find Maris on - You can find our Mustafa on -
Transcribed and scored by The B2B Podcast Index.
Speaker A: Foreign.
Speaker B: And that's how. This is how basically. So, um, you've went from version one of your, uh, the pseds system onto version two. So version one was basically taking all of Maris's work and making it digital automation. Basically. That's, that's, that's what, that's what you had to do. And this is when you told me now your people wanted to be things more precise. Basically Maritas team wanting things to be more precise. I'm guessing this is what happened in 2019. This is when you started to work on the second version. Or am I completely off the market?
Speaker A: No, that. That's pretty much it. So basically I think Merits realized that she had like, um, the sky was the limit in terms of technical things. And uh, I Remember back in 2018 there were already some discussions around, yeah, we might need like. So Maris was preparing me psychologically, uh, we might uh, improve pseuds, blah, blah, without saying too much. And then she came in 2019 and said, Listen, I have an idea. Could uh, we look at all of that? But we just start with a few extractions and stuff like that so that she can review them manually. So, history of some given players, things like that. Then she was going through like, I don't know, hundreds of thousands of records on her own, checking, chat and stuff like that. And uh, yeah, we went from what, three, four features in a model into 24. Wow. And each of them, uh, is a bit tricky or sometimes very tricky, but there's nothing. There are maybe a couple that are easy. Like simple ones, like number of beds, for example. Yeah, that's easy, you know, but then like something like a night play. Yeah, we want to see when someone plays at night. Yeah. Okay. What do you mean by night play? Uh, do you mean like in the current time zone? The time zone of the player between 1am and 7am Maybe? So some people will tell you from 11pm It's a bit, uh, late. Okay, but then you have a lot of players playing between 11 and midnight. And then you have people working at night, uh, you know, like a security guard, for example, maybe he places a few spins at 2:00am and then most importantly, you have events in the U.S. if you have a big NBA game, NFL game, whatever, we will have quite a few people who plays a bet at 2, 3am M for, uh, our time zone, you know, because they're watching
Speaker B: the game live, they're following the game.
Speaker A: Does that mean that the guy has an addiction problem? No, he's actually following the game. Like, leave him alone. Um, so then you see like just a simple concept like that playing at night turns into a complex situation where you need to check like, okay, you need to think about it, like, how do you actually do that?
Speaker B: Give us a walkthrough. This is very interesting. I understand how complicated this sounds. Absolutely. I can only imagine, you know, the work that's been put into it. But give, give our audience a taste of, of what it's like. I like this specific scenario, as you said, somebody from, from, you know, based in Europe, um, and they're placing a bet on an NBA game. Let's just say, you know, distinguishing that from somebody who's up late at night because they're developing an addiction, addictive behavior.
Speaker A: But the thing is like, uh, at least I didn't know of any method known method to do that from a data perspective to detect that. So it's always the same thing. Like you need to take a step back. So initially you have your first idea, okay, nighttime play. Okay, um, so you check all the players who played from 11 midnight until 7am Something like that. But then you need to take a step back and think about the situation, what kind of false positives you get and stuff like that. And then you analyze. So it's uh, for each feature, uh, you need to analyze what's going on. So you extract a lot of data. You look at the distribution and all that. And then you understand that for example, on some days you see spike at a specific time and all that. So then how do you cope with that in the end? Um, the way we do it is we almost forget about the time zone. What we do is, or the country and all that. So from just simply taking the time, what we do is basically take all the players per time zone. All right, so for example, all players who are based in the CET time zone, uh, UK time zone, whatever, uh, and you look at the distribution per time slot, like 15 minutes time slot, for example. So let's say you had 100,000 players within 24 hours. What proportion of this 100,000 played was active? Uh, during this 15 minute time slot? This time slot dim timeslot, you cut it down so you have a proportion of players and at some point spikes. You reach like 30, 40% of your actives were active at this point in time. And sometimes you see a drop, which means that if you have like, if you look at it on a normal regular, uh, period, if we speak central time zone, for example, you will see that there is a massive drop between 1am and 7am, which already challenged the 11pm site. You know, uh, but also it's not like um, rigid. Like when I say 1am maybe on one night it will be half past midnight or midnight or sometimes it will be 2am it will depend. But more specifically, uh, when you have an event in the US then you will have a spike. So then the proportion active during that time slots is higher than a certain given threshold. Now another issue that you face once we've done that is that the thresholds needs to be different by time zone. Because on some time zone we don't have that many players, so it fluctuates a lot. So if you say like under uh, 5% is. So we switch the name, for example instead of nighttime we call it anti social time.
Speaker B: Okay.
Speaker A: And so for some market, let's say like if you have less than 5% of your actives in this given 15 minutes time slot, yeah then it's anti social time. But for some markets you need to increase it and for some markets you need to, for some time zones actually you need to decrease it. And I'm even speaking about time zone, but uh, like time zone, like in terms of data, you need to know in which time zone the guy is. So financial activity, you get it from the Betsy place, the deposits he made and all that. But the time zone you would get it from the login information. The login not necessarily happened when he bet. So you need to join all the people who played with the latest login, look at the um, geolocation of the login and then you need to join that to the time zone. But a country like Australia for example has four or five different time zones and stuff like that. So how do you deal with that? Because then you end up with some time zones where you have only 10 people in there. So there's a lot of little details here and there as well, you know. And it's just one feature.
Speaker B: That's just one feature. And you said you went for. And at the start when you joined, you guys put together three to four and now it's basically 24. So. Okay. Um, from an employee point of view, the challenge is just enormous. Absolutely enormous. It's very exciting at the same time because if somebody actually wants to be busy and they want to do productive work and they want to see the results of what they're doing, um, this is a great opportunity in general. Now I'd like, I want to ask a question though. This is version 2, now you're working on version 2. The system is there, has the Main challenge is already gone. Have you gone past it now? Is it more of a day to day work or are there still some challenges up ahead that uh, you know, people who are looking to get into Kindred, for example, can say to themselves, actually there is still a great opportunity for me here to challenge myself to really learn, to really have an impact,
Speaker A: at least from a technical perspective. Yes, sure. Maris would agree. So the thing is that uh, we have a version 2 in place, but there's still a lot of improvements we can make. Like it's a great model, uh, it works very well and all that, but it's, I mean it's a complex question to answer. It's a complex thing to monitor and we want to reach a very, very high level of accuracy. We want to capture people uh, very early. Um, and we want to avoid false positives and we want to avoid false negatives. The main issue is that we don't have a clear database with a flag saying this is problematic. Uh, gameplay. This is not. Okay, Maris mentioned earlier uh, m research paper using self exclusion. So players who self excluded consider them as problematic and then you look at their history. Now that being said, the um, level of problematic gambling of a player over time fluctuates. So when someone self excludes the high risk of activity might not be just before the self exclusion. It could happen one month before because a guy could have a spike in problem gambling and then he stops playing and then he just places a few bet and then he disappears. He self excludes.
Speaker B: Why?
Speaker A: Because he was probably playing with a competitor or he was playing on land, uh, based casino and things like that. So this is the biggest challenge of this story because if we had that it could be easy to just use supervised learning and do that. So one of the main step we wanted to reach with that first version of the V2 is to get something that monitors players over time, uh, give a risk level over time and then the RG team can review these and modify um, the risk level if they think it was wrong. Which means that we're basically building a massive data set. It's been running for how long? A year now? No, third of November will be one year of data, uh, with a lot of players and a lot of fluctuations and things. So there's a lot of, there's a massive toy basically coming up, a massive data set which is unique with a lot of data features already, uh, calculated and then dig into it to improve it. And there is also another side, sorry, uh, it's the real time aspect. We Want to make it more real time and all that. That's even another story.
Speaker B: So basically bottom line is the challenge isn't even close to being over. There is so much yet to be done. Look at Mar is also shaking their head that no challenge is not even close to being finished. One thing I do wanted to ask, um, we had discussed uh, before, um, was Karim, you mentioned to us how you, on the technical side you're able to, uh. Part of one of the features that you have is you're able to detect whether somebody's using multiple accounts, um, using different accounts, um, you know, um, any of your other websites, you're able to create, you know, their profile. You know that this is still this person. Now obviously we can't go into too much detail, but do you want to give us perhaps a little bit of an overview as to the challenge of actually creating, being able to connect the dots here. And uh, I don't know if that's something again on the technical side purely, or is that something in, you know, um, you know, as a combination between you and Maris. That's something you guys had to sit down and discuss and work on together.
Speaker A: But you always need the domain expertise. That's for sure. That. That's Maris.
Speaker B: This is, this is all about. You're, this, you're the, that's it that you're the head of this. It's your baby. No matter what happens, she's monitoring everything.
Speaker A: Yes. In the, in the end you have different uh, level of uh, uh, resolution of that problem.
Speaker C: Um,
Speaker A: like we basically have two, two levels. One is a simplistic approach to minimize the number of false positives. And then we have another one that is more advanced, um, that will have false positives and stuff like that, you know, so because you, you need to have like one strict version of linkage and one non strict version version of the linkage. For example, the strict version of the linkage was built by a, uh, data science team because. So we collaborate with different data teams. Uh, we are dedicated to leco. We have also, um, a global data science team based in London for example that works with any department, including us. Um, while we dug more into the one that looks at almost everything basically. Um, so for example, the thing is, it's mostly relevant with fraud to be fair. Uh, when you see players. So we link players based on what they do financially, for example. Uh, so you will have a lot of players who do exactly the same thing. All right? They create an account and they follow exactly the same pattern. Uh, or they Reactivate and they follow the same pattern and stuff like that. It's basically either the same person or group of person in collusion. Uh, we look at any kind of linkage, uh, linked by device, fingerprint, linked by, uh, geolocation. Geolocation is a crazy one as well.
Speaker B: Um,
Speaker A: because you can use that, you can look at the IP address, can look at geolocation. But for example, um, when you start looking into that data, you realize that when you use your phone, you're connecting with 4G, you're connecting with a, uh, uh, relay and everybody in the area will use the same. So you will end up with the exact same geolocation. So when you look at the geolocation pre size one, you will end up with one that has, uh, 2,000, uh, actives in a day, which is massive. And you think maybe I made a mistake. Uh, so you need to clean that up a lot. Uh, all those geolocation data. Uh, but also there's another issue is that you need, um, when you get latitude, longitude, uh, details, you need to compare, um, everyone. So let's say S3, we are logging in, we have different logins with different locations. Uh, how do you compare those locations? Efficiently? So we might end up with six logins, uh, each. Yeah. So I will have my six different locations. Maris will have, uh, six locations. So it's six times six to comparison. And then you have six. So it's times six again. Yeah, six power three. Imagine if you have a hundred thousand players. Goes crazy. Yeah. So then there's a way with something.
Speaker B: What is it that you've started? What is it that you've. What is it that you've started?
Speaker A: Well, she was pushing for research.
Speaker C: Huh.
Speaker A: That's the thing, you know. So the way we ended up doing it was, um, so we have a data scientist, Abed, uh, in the team. Like, uh, without him, like a lot of the things we've done wouldn't have been possible. And um, for this, for example, what he's done is, um, map the words into squares, basically. Agreed on top of the word, let's say. Yeah, he does cool stuff like that. Like you think you're powerful. He is more powerful than anyone. You know, he puts a grid on the word. So basically what happens is that we all end up in little squares. Yeah. And then you compare just the people falling in the same square and then
Speaker B: you start up, I don't know if you've heard of it. They have a similar, they have a similar thing called, I, uh, think what three words. It's I don't know if you've heard about it, it's here in the uk. Um, it's an app which essentially does exact same thing. And when you click on the app it gives you three words so that if you're calling the emergency services and you just, you have no idea where you are. Even if you know a remote address, but not very clear, if you click the app, they'll give you three words and those three words reflect the exact square where you are. Uh, funny enough. So yeah, it's very interesting that it's this, this you guys have also used that technique. It's, it's as accurate as it gets, basically. Yeah. Which tells our audience really how high tech, how high level the work that you're doing is.
Speaker A: But if you want to know how high level, like the guy used C&AVX. So AVX is um, a, uh, machine oriented language. So it's not like C or something. Like it's between C and binary 01101. Because it was running on RStudio server back then, it wasn't that powerful. So you had to actually go into the calculated, uh, and I remember for 100,000 players logins, to compare them, it was taking him 20 minutes, 25 minutes. And then after all the changes and using C&AVX for a million logins, it was taking less than 10 seconds like he went like crazy. I was shocked, really.
Speaker B: Well, there uh, you have it. This is the kind of team that uh, that's currently present in Kindred. This is you. You, you want to learn, you want to be surrounded by geniuses. Here's one place, here's one place to consider.
Speaker C: Wow.
Speaker B: Yeah.
Speaker A: If you said geniuses, it's more like him. He's exactly, it's a bed. Anybody who sits near him, um, like the new joiners or whatever, people from other department. Like he doesn't speak much but when, when you see what he does, just you look at his notepad, you're like, okay.
Speaker B: Mind blowing.
Speaker A: Yeah, yeah, he's quite, quite impressive. Incredible.
Speaker B: Maris, from your side of things, that's on the tech side, from your side of things, what's next? We understand obviously there's a lot more to do, tech side to get. So Karim has a. And his team, they have a lot of work to do to get you more, uh, to make this as detailed, uh, as niche, as accurate as possible. But is there anything more or do you feel. Because clearly you come, you know, you approach this from an academic background. Yes, there is the ethical background which Made you start this whole thing. But you're also very driven from a research point of view. Um, again, to not go off on a tangent. What's next? Simply building, making this more accurate or do you have more in mind?
Speaker C: I will always have more in mind. Always.
Speaker B: Look at Karim. He's like, oh God, she has more.
Speaker A: I know, I know. It will never end.
Speaker C: Well, when we launched the second version, I was already thinking of the third version. So no, it will always be.
Speaker B: Kareem is nodding. He's like, I know, I know. We submit. We just launched the second. She was already working on the third. I know.
Speaker A: Yeah, I have a Jira. We work on Jira for the tax. I have a Jira with uh, an id.
Speaker C: Yes, yes. I, I get this, these random ideas and stuff and I always go, and then I ask Karim, can all the data for this specific customer group for this. And he's not even asking anymore. He just gives it to me and just waits.
Speaker B: That's it. He, he, he's learned by now, but okay, so there's always next. Anything you can share with us or
Speaker C: perhaps in another podcast, um, I mean Karim mentioned the technical side. So to move everything to real time, to have more automated interventions. So if, um, if the first version you had like um, 0.5 of our customers approach, we're looking into increasing this to 2% so even lowering the risk, so accepting more people, but in an automated intervention so they get more pop ups and that sort of thing. And once that is launched, there's so much more data to look at because then you see like which is the best.
Speaker B: Look at the smile she has or look at the smile.
Speaker C: And uh, something that I really want to work on is about binge gamblers. So binge, uh, gamblers similar to any binge behavior. But you see some customers that might come in, lose very quickly and self exclude or stop gambling. So it's not just a matter of detecting them, but it's a matter of intervening in the very, very quick time to get them to change their behavior. So right now that is something that uh, I'm really interested in looking into how we can. Because it's a whole new ball game. Then it's a whole new different of
Speaker B: course because as you said, if it's binging, it's a short period of time.
Speaker C: Exactly.
Speaker B: And uh, wow. I can only imagine what the work will. Well look, I can see how excited you are about it and I can see Karim thinking to, I can see the look on his face. Yes, we have all that Work to do. Absolutely, yes. She's going to throw all that on us.
Speaker C: He's just complained. Now he's not even questioning it anymore.
Speaker B: I love the chemistry between you guys. I absolutely love the chemistry. I, I, I, I've always thought, I've always thought that the chemistry between colleagues is, you know, a huge part, if not the main part, as to why work, uh, is delivered. Well, a project, the success of a project basically, is to do with the, with the chemistry. Sometimes, not necessarily the skills of the individuals. It's more to do with the, with the chemistry. But clearly you guys have both here. So. Which is which, which is quite incredible. Um, Maris, Karim, I cannot say how incredible this has been. Absolutely brilliant. Very informative, uh, to even, you know, we just need to let our audience let this sink in. We're talking to a gambling company and they're ethically driven, you know, responsible gambling by a gambling company. Um, this whole approach, Maris, that actually made you want to do this, the journey that you went on, uh, the work that you started. Karim, you know, from a technical point of view, anybody who's technical, all the developers and the, you know, uh, data analysts we have and scientists we have in the community will be able to, you know, hats off to you guys as a team for the work that you have developed. And, you know, I'm sure lots and lots will be, lots of people will be, uh, very intrigued and interested to, in a way, be associated, Associated, uh, with, with, with that, um, are we gonna see you guys in the upcoming event?
Speaker A: Yes.
Speaker B: Brilliant. Brilliant. It's not gonna be too long from now. Uh, we are working on it. Um, in the meantime, again, Maris, thank you very much. Karim, thank you very much indeed. Um, it's been very informative and I will let you guys have the rest of your day. Thank you. Have a great one.
Speaker A: Mustafa, thank you.
Speaker B: Thank you very much. My pleasure indeed. My pleasure indeed. Speak soon, Maris. Karim.
Speaker A: Bye, bye, bye.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.