
The Genetics Podcast · 2026-09-10 · 41 min
Key moments - from our scoring
Substance score
77 / 100
Five dimensions, 20 points each
Drug discovery has shifted fundamentally over Dave Hallett's 30-year career from reductionist target-based screening to high-dimensional phenotypic mapping enabled by improved data generation and computational systems. AI has compressed discovery timelines by 50-70% and reduced compound numbers by 80%, but the critical bottleneck remains clinical translation - phase two efficacy still determines whether a program succeeds. Hallett emphasizes that while generative chemistry, AlphaFold-driven structural biology, and computer vision on phenotypic screening deliver measurable impact, AI struggles with out-of-domain predictions and cannot yet model emergent pathophysiology or idiosyncratic toxicity. He highlights Recursion's practical wins in clinical trial optimization: using perturbational maps and simulation to relax overly stringent inclusion-exclusion criteria, improve patient recruitment by 30-40%, and expand eligible populations by 20% - proving that much of AI's value in drug development lies in trial design and patient selection rather than fundamental target discovery.
AI has achieved 50-70% time reductions to clinical candidate stage and 80% reductions in the number of compounds needed to be synthesized, primarily through multi-parameter molecular optimization and active learning loops that compress SAR cycles.
The biggest failure driver is selecting biologically irrelevant targets; even with good target selection, about 20-30% of programs fail because the drug modality cannot adequately test the hypothesis due to off-target effects preventing adequate dosing, and trial design and patient selection account for another major portion of failures.
A perturbational map systematically perturbs cellular systems using CRISPR, compound libraries, or biological reagents, then embeds resulting cell images into high-dimensional space using deep learning to identify genes or compounds producing similar phenotypes that may share mechanistic relationships, uncovering functional connections unavailable through traditional screening.
Recursion uses clinical trial simulation and real-world patient data to systematically test inclusion-exclusion criteria, identifying which parameters actually matter for efficacy versus which are unnecessarily restrictive, enabling 30-40% recruitment improvements and 20% expansion of eligible patient pools.
AI models struggle with idiosyncratic toxicity, species differences between animal models and humans, off-target phenotypes, tissue-specific exposure, and cannot yet simulate emergent human pathophysiology at organism scale - areas where compounds may look perfect in preclinical assays but fail in clinical translation.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode delivers substantial, non-obvious claims about AI in drug discovery with strong domain specificity. Hallett articulates concrete distinctions between where AI genuinely helps (SAR optimization, clinical trial simulation, patient recruitment) versus where it remains overstated (simulating human pathophysiology, idiosyncratic toxicity). The perturbational mapping explanation and the discussion of phase-two efficacy as the true bottleneck are particularly valuable. Some throat-clearing and repetitive framing reduce density slightly.
We're still a long way from I would call simulating the human body. I think models still struggle, I think to simulate emergent human pathophysiology on an organism scale.
We've probably expanded the number of patients we can reach by about 20%. So they're some of the real, the real areas.
Hallett avoids common platitudes and instead offers contrarian specificity: the clinical efficacy cliff is not solely about biology but about trial design and patient selection; perturbational mapping anchored in human genetics is Recursion-specific and less widely discussed; the 70-80% irreproducibility point challenges the field's data quality assumptions. However, the broader frameworks (AI as tool for optimization, multi-parameter trade-offs) are industry-standard, and the call for domain expertise spanning biology-chemistry-computation is well-worn.
I think what's really interesting is that biology wouldn't have predicted from the other means.
I think it's all of those kind of criteria. I think the biggest, the number one risk and the reason the number one risk of failure is just picking a biological target that was not cause related to disease.
Hallett is genuinely exceptional for this topic: 30 years in drug development (Merck, CROs, now CSO at Recursion), deep medicinal chemistry background, and hands-on experience with clinical translation and AI application over 5-7 years. He has shipped real programs (MEK inhibitor in FAP, targets in neuroscience) and carries real operational burden (clinical trials, regulatory engagement, portfolio management). He avoids pure punditry and grounds claims in executed work. This is a practitioner with substantial scar tissue.
Yes. And say thank you for introduction. Yes. And guess like many of us, if you work in this space for long enough, whether it's applying AI ML or thinking about 20 years, a lot of scar tissue and a lot of learnings along the way.
I'm a big believer and as we said at the start of the call, have the scars that come from clinical development. I think you have to have skin in the game.
Hallett provides concrete numbers and examples throughout: 80% reduction in compounds made, 50-60-70% time savings to candidates, 30-40% recruitment increases, 20% patient population expansion, 1 trillion human iPS neurons manufactured, 17,000 genes knocked out, 30 million images captured, 3 trillion searchable gene-compound relationships, 1 in 5 trial sites fail recruitment, $40,000/day site costs. These specifics ground abstract claims. However, some assertions lack full citation (the 70-80% irreproducibility claim, efficacy failure attribution percentages), and the Roche/Genentech collaboration details are mentioned but not deeply unpacked.
We've seen a benefit in terms of 30, 40% increases in recruitment, picking the right patients.
We had to manufacture over a trillion, one trillion human IPS derived neurons just to give you some context, that's about the same number of neurons as in, as in 12 human brains.
Patrick Short asks strong, substantive follow-ups that push Hallett to clarify and defend: 'Where is AI really genuinely changing what you do and where is it overstated?' and 'Do you have intuition on how much [of efficacy failure] is about understanding biology vs. trial design?' These prompt Hallett to articulate nuance. However, the host occasionally lets claims sit without deep challenge (e.g., the 70-80% irreproducibility statistic is mentioned but not probed for source or implications). The conversation meanders slightly and misses opportunities to pressure-test specific Recursion claims or ask harder questions about execution risk and timelines.
I'd love your perspective on that. Where is AI really fundamentally changing what you do and where do you see it as? Not yet, but maybe with a few more cycles in the future it will.
Do you have any intuition on if we, if we think about this efficacy Cliff that you've described which, which I think the whole field would probably agree is, is the biggest point where if we saw AI driven drug development move phase two success rates
Computed from the transcript - who did the talking, and the words that came up most.
This week on The Genetics Podcast, Patrick is joined by Dr. David Hallett, Chief Scientific Officer at Recursion. They discuss the biggest shifts in drug discovery over Dave's three-decade career, where AI is genuinely transforming the field today versus where the hype outruns the reality, and how Recursion's perturbational maps have uncovered and validated a novel neurodegeneration target. Show Notes 0:00 Intro to The Genetics Podcast 01:00 Welcome to Dave 01:57 The biggest shifts in drug discovery over three decades 06:39 Where AI is delivering real wins across drug discovery today 12:19 How AI-assisted trial simulation reveals which eligibility criteria to relax 14:46 The three biggest reasons drug programs fail in the clinic 18:20 How Recursion's perturbational maps uncover new drug targets 24:58 A four-step framework for validating a novel drug target 28:44 How Recursion balances deep therapeutic focus with partnership breadth 30:52 Why AI can't shortcut clinical trials, and what proof of real impact looks like 34:57 The skills scientists need most in the AI era, and why trusting AI outputs starts with trusting the data 39:52 Closing remarks Find out more: Recursion ( )
Transcribed and scored by The B2B Podcast Index.
Speaker A: Hello and welcome to the Genetics Podcast. I'm your host, Patrick Short. My background is in population genomics and studying the genetic causes of rare disease. I did my PhD at the Sanger Institute and the University of Cambridge and have been in biotech since 2018, when I started Cyanogenetics. Cyanogenetics helps academic and industry researchers to run large scale genetic testing programs that speed up their clinical trials, generate data sets for the next big breakthrough, and give participants the best possible experience taking part in research. Each episode of the Genetics Podcast, we bring you insights from the leading minds in genetics and precision medicine, including household names and Nobel prize winners, as well as early career scientists and biotechs working on the next big breakthrough. Whether you are a scientist, entrepreneur, executive, patient advocate, or simply someone curious about how genetics shapes our world, you're in the right place. Thank you for listening and let's get started.
Speaker B: Hi everyone, and welcome back to the Genetics Podcast. I'm really excited to be here today with Dr. Dave Hallett, who's the Chief Scientific officer at Recursion. Dave is a medicinal chemist by training, but has more than two decades of experience across really the whole spectrum of drug development and in particular AI driven drug development in the last five to seven years. So I am really looking forward to this episode because AI is transforming the way we do drug development right now, and there are few people that come to mind besides Dave that I think have the level of scar tissue and time spent in the space and also the deep background in chemistry where I think a lot of the application of these models are going today. So, Dave, thank you so much for taking the time to join me and looking forward to the discussion.
Speaker C: Yeah, likewise. Uh, and say thank you for introduction. Yes. And I guess like many of us, if you work in this space for long enough, whether it's applying AI ML or thinking about 20 years, a lot of scar tissue and a lot of learnings along the way.
Speaker B: If you think back on your career, I think you started around 2005 at, uh, Merck. A lot has changed, not just in chemistry where, where you did your initial training and work, but drug development more broadly. What are the biggest shifts and what do you think has. Has driven the most change in that 20 or so odd years?
Speaker C: Sure. Actually, I guess I look younger than AIP. I actually started my career after a postdoc at Merck. 96, 97. So, yeah, nearly 30 years. Um, I think what's. I think what's changed the most for me, I guess, are the. Where the bottlenecks have moved to and I guess underlying philosophies as well. I started my career a few years, my professional career a few years before the human genome products I think projects. I think one of the things that has obviously radically changed is just the ability to generate data, meaningful data. Yeah, 25, 30 years ago it was like yeah, trying to get, trying to get RNA sequencing data was. Yeah, that was several postdocs of work was the day, yeah, you just send off your kind of sample and you'll get a sequencing. And yeah, you think about the cost of the human genome project. But today it's like it's just routine to do a genetic sequence of a human. So I think the ability to generate I think much more molecular data has certainly changed. I think for me I think what's some of the big differences are, I think, I think because a lot of this is due to like just the ability to actually deliver assays is that uh, our ability to generate complex assays improved. So the early discovery, I think back in the 90s when I started my crib relied really heavily on reductionist target based screens, think sort of purified enzyme assays. And that's because that's all we could do at the time is that we wanted, we knew that they were reductionist but when we wanted to model complexity but couldn't. I think you look at what we do today is that programs prioritize kind of high dimensional phenotypic kind of mapping. Multiomics is the word of the decade and we use a lot of functional genomics now to both to help to validate kind of looking at targets very much in disease context. I think yeah, the ability to generate assay systems that are disease relevant. I think as a chemist that's what I identify as. I think hit discovery used to mean brute force kind of high throughput screens on massive millions of chemical libraries and then really slow sequential SAR cycle. Clearly what that's shifted to is the ability for particularly using computational systems is true multi parameter optimization. As a human I think I'm pretty good at carrying variables in my head but I can probably think about six or seven at once. Clearly a computation system does not forget any data, does not have any recency bias. And so and then you've got this concept of active learning loops. I think computation systems are programmed in such a way that they can understand the limitations of the model that are produced. Humans tend not to do that. And so what you see now I think is true kind of iterative parameter optimization. It's not changed. I'm still thinking. I think there's a paper back in 2010 that came out in Nature that has looked at like yeah the cost of kind of failure. Like what why ah, why is. Why is the pharma industry really struggling? I think the thing that hasn't changed is that yeah clinical translation is just not negotiable. Just because yeah you can design a sub nanomolar binder in silico is that doesn't mean that that target is in any way l to a disease hypothesis. And if that's flawed so phase two efficacy still remains that ultimate truth that we're trying to look for, don't think that's changed. Despite the kind of the improvements in molecular prediction and recursion through mold gps, I think we have a really remarkable system to predict kind of general properties of molecules. I think it's fair to say that while admet modeling and toxicity predictions have improved the ability to model tough areas like idiosyncratic toxic how does an individual patient respond? Tissue specific exposure off target phenotypes. Yeah. Species differences. Yeah. A rat and a dog and a human are not the same. They're related clearly. So you still lose compounds that look flawless in pre clinical assays. So that's still an area that I think that we can work on. And then last but by no means least I think we think we need to remember like a molecule for me is a, it's like a, like a race car but it's always a multivariable compromise. There's no such thing as perfection. I think every drug is uh, always going to be an intricate trade off between potency, selectivity. Does it get to its site of action? Is it cleared? And optimizing one parameter always places a stress on another and I think that's the thing that's not changed but I think what kind of modern AI ML methods have helped is to kind of like yeah it's a more rapidly find the best, not the perfect.
Speaker B: Yeah, yeah. One thing I mentioned to you before we started recording is something I'm really excited to understand is if we get into the detail of where AI is really genuinely changing what we do and where it's overstated but maybe there is promise for the future but it's just not happening today. I'd love your perspective on that. Where is AI really fundamentally changing what you do and where do you see it as? Not yet, but maybe with a few more cycles in the future it will.
Speaker C: Well that's a fabulous question I think is the uh, let's not dismiss the day to day work. So I'm talking outside of the kind of the scientific environment thinking about the ability to, to just search data. The, the ability to summarize, the ability to summarize documents, the ability to create things. It's like. And if I'm not the most artistic person in the world. So for me creating PowerPoint decks or visualization is not great. That's why I have people like Britain working, working with me. But the tools today are just phenomenal. Like they're so it's very easy for me to go in because our recursions internal systems are so interlinked. I can go into a very natural kind of chat prompt either one of our own or yeah using kind of something from anthropic and I can ask a question, say tell me everything about this particular piece of information. I still need to verify that. So yeah just productivity tools of. Yeah my postdoc and my PhD would have been so much easier 25, 30 years ago than, than they were at the time. I think the kind of, I think the reality. I think AI has massively compressed the time and the cost that's required to get from an idea to a candidate. You see that in the numbers that we put out in terms of significant kind of 80% reductions in terms of the number of compounds that are made and shaving sort of 50, 60, 70% off the time to get to candidates. But I think what still hasn't really changed is that so we still, I think we're getting to the clinical trial line faster than we were but we still haven't fundamentally solved that historic kind of that attrition rate in the clinic. Yeah but coming back to like the wins then the multi parameter optimization I spoke about modern generative chemistry and active learning. Optimizing SAR is so much faster than manual cycles. Uh that's dramatically changed in my 25, 30 years kind of as a practicing scientist. I think. Yeah. In some of the local solutions and what I mean by that, uh, I think structural biology and the virtual screening. So think bolts, think alphafold, think those, those Nobel prize winning kind of movements. So I think protein structure prediction and yeah large scale structure models have certainly transformed what we do and they've turned what was an expensive bottleneck I think into like an accessible starting point for chemistry. But I think we have to also admit that those, those models still struggle out of domain. They're great on their training sets and one should never forget that binding and function are not necessarily directly correlated. But I'm glad we, I'm glad we have those tools for sure. And then both in discovery and development, think computer vision, think. And uh, we use it obviously in terms of like yeah high content phenotypic screening but yeah computer vision on, on massive kind of microscopy or data sets so that because the human eye just can't, can't deal with that amount of data, it's not, it won't pick up the subtleties and so we can pick up say disease signatures that a human eye can't. But clearly anybody that's had much if you're aware but if you've got an X ray these days or an eeg, it's very likely that those systems, that raw data will be initially pre processed and analyzed by an AI system. And then the kind of the things that look maybe out of place are then flagged to a human expert. So I think they're the areas where we're seeing the wins, maybe just down a couple of ones. I think from a recursion perspective is the again quite, quite practical this is in development now is that there's no point trying to disrupt kind of drug discovery if you don't then understand yeah kind of what the strengths and weaknesses are just of clinical development. And so using, yeah first of all pulling kind of relevant data sets together and then putting them into, and getting them into a shape where you can build models from them. We've been able to use that uh, to do clinical trial simulation and to do. And the benefit there has been like real benefit in terms of accelerating recruitment into a clinical study, finding out where the patients are kind of this the thing. This sounds really trivial but I'm not sure your listeners would know is that apparently like about 1 in 5 clinical trial sites fails to recruit a patient during a clinical study. And yet on average I think it costs something like $40,000 a day to run these trial sites. So you go okay, wouldn't it be great if it's the. I didn't use those sites in the past and that data is there and we spoke about this and my counterparts, my colleagues have spoken about our ability now to really pinpoint where potential patients are around the globe, not just in a particular country. And so we've seen a benefit in terms of 30, 40% increases in recruitment, picking the right patients. So I think yeah, the clinical trial simulation around our recent pi3, pi3 kinase molecule that's going to the clinic maybe you're the kind of the less development experienced people. What tends to happen with a small company is that often they'll go to clinicaltrials.gov and they'll just copy a protocol. But what you really want to do, and there's nothing wrong with that, but what you really want to do is like, you really want to, you want to pressure test the inclusion exclusion criteria, Are they fit for purpose? Talk to Kols and then run the simulations, because each one of those criteria will have an impact on who and who cannot come into your trial. Clearly, you need to make sure your drug is safe and do no harm. But we've applied that kind of like that, that kind of, that process to that particular clinical study. Yeah, we've probably expanded the number of patients we can reach by about 20%. So they're some of the real, the real areas.
Speaker B: So is the big thing you're finding there that there are ways to relax the protocol in the sense of. Make it more permissible for patients without undermining the, the scientific integrity? Because what I see a lot on protocols is they're, they're very narrowly scoped and, and that fits my intuition, right? There are probably a lot of those parameters. You can relax and, and still get the, still get the high quality trial, but not turn as many patients away. Is that the key left?
Speaker C: Yeah, it's exactly that. It's like, it's like looking at, if you, if you go into a detailed protocol, there's often like tens, if not hundreds of like inclusion and exclusion criteria. There's like certain, like, yeah, uh, if you do a blood test kind of what, what are their metabolic kind of parameters? And so for a lot of those are perhaps set maybe 20 years ago in some cases or. Ah. And so you go, okay, well, okay, but then you can actually do an analysis, say, okay, which one of these actually has the biggest impact? And so what you can then focus on with Kols, because clearly the whole point of the clinical trial is not to do harm. And obviously then to then to show efficacy, you can then ask kind of questions going, okay, is this, is this inclusion or exclusion kind of parameter really fit for purpose in the modern environment? Is there data that suggests that this is too stringent or too lenient? But you can do that very quickly through a clinical trial simulation before it would take teams weeks or months to manually do this. You can do this in a matter of hours now and go, okay, I need to focus on that particular parameter. Uh, and that translates into much more accelerated clinical trials.
Speaker B: I wanted to. Yeah, go ahead, please.
Speaker C: Just to finish off. I think the, what do I think is overstated because, I mean, that's an Important part of your question. I think I've spent a bit of time hopefully convincing you that where things are working. I think we're still a long way from I would call simulating the human body. I think models still struggle, I think to simulate emergent human pathophysiology on an organism scale. I think as an industry we're getting better at both cellular and tissue context. But I think there's still a lot of work to do around. Yeah, immune crosstalk for example or predicting long term organ talk from scratch. Not saying we aren't making headway there but I think it'd be remiss of me to pretend that all those kind of problems are fixed and. Yeah, and complex toxicology and off target biology. I think phenotypic systems like the ones we operate at recursion allow us to look more holistically uh, at a drug's fingerprint. But yeah, the kind of classic areas of videosyncratic talks are still very difficult for any system to kind of model.
Speaker B: Do you have any intuition on if we, if we think about this efficacy Cliff that you've described which, which I think the whole field would probably agree is, is the biggest point where if we saw AI driven drug development move phase two success rates that would be a, a sigh, a sigh of relief at a moment of excitement. I was, I'm wondering if you have a, intuition or, or data around how much of that is about understanding the biology. So the, the program is to, to some extent doomed from the beginning because you haven't picked the right target and then how much of it is the trial design, whether it's enrollment or endpoints or something like that where actually you probably were on the right track but you didn't pick the right endpoint or didn't pick the right population or something like that. Do you have a sense of where the, the biggest driver is?
Speaker C: I think it's all of those kind of criteria. I think the biggest, the number one risk and the reason the number one risk of failure is just picking a biological target that was not cause related to disease. That's still the big challenge for the industry. I think once you step outside of monogenetic diseases where I think we've got a lot, we've got a lot better. The Human Genome Project allowed us to look at that. I think we understand now that highly penetrant monogenetic diseases, we haven't got treatments for all of them to be clear. I think some of the rare diseases but I think we understand yeah that the link between like a Huntington's disease mutation and the people getting the disease. Unfortunately most diseases are not monogenetic in nature. They're not necessarily polypharmacology, but they're pathway related. And so I think that that's one of the real bottlenecks is like when you, when you, even when you look at kind of real world evidence data like as in patient molecular data, there's one thing saying, okay, uh, this is interesting. The responder non responder population, they're differently differentially expressed genes compare and contrast but then go yeah, but is that, is that causative or is that just a reflection of disease? So I think that's still the hard thing, but I think the, I would still say in probably 20, 30% of the occasions. I think people pick a target, they find a modality. It could be a small molecule, it could be a biologic, could be something else that effectively interrogates the target. Because that's the other thing I think sometimes projects fail because the idea with the biological idea was a good one. But you take your molecule into the clinic and maybe because of off target effects is that you can never dose it high enough. So you never actually test the hypothesis. You get so far up the ladder, but you never really ask the question. I think the other side of that equation is like you tick the first two boxes, I. E. You picked the right mechanism, you had a modality that was good enough. I think it's then the practical realities of clinical trial execution. So this then comes down to the classic of did you pick the right patients? Because the last thing and the most important thing, a clinical trial is an important event, but it's purely a statistical event. You're looking for a signal like yeah, a perfect clinical trial would be I dose 100 patients, 100 patients respond, what a fantastic statistical outcome. Doesn't work that way I think. And so one of the areas is still and AI is actually having an impact here is like choosing the right patients in your clinical study because it is very rare that you'll the blockbusters like the GLP1s for example, the anti TNFs are few and far between the realities is that whether it be in oncology or neuroscience or immunologies that you need to really understand the patients that will respond to your drug. And so that, yeah, I guess they're the three areas I would look at
Speaker B: in across any of those areas. Pick whichever one naturally lends itself. I'd love to go into the details a little bit more about actually what an agent, uh, or a system that you Build does and whether they're, whether there's a huge need for new data or actually we have a lot of this data and it's really about gathering it. So m. Maybe I'll provide an example. You have a, I, uh, think a recent announcement with Roche and Genentech that you've had a milestone, uh, for the microglia map you've been building. I think this is, I'm sure, uh, a way to get at this. Do we have the right target in the first place and we really understand the biology. Be great to talk about how much of that is hoovering up all of the existing publicly available and maybe internal single cell data and other types of data sets and how much of that is actually we need to go and generate new data because there's some pocket of biology that we, we don't understand and we need to feed it into the model. How. If you could open the hood under that kind of program for a minute, I think that'd be really interesting.
Speaker C: No, it's this, It's a fabulous question, I think. And it goes to the heart, I think, of what we're trying to do at recursion is that I think ultimately the critical piece of information you need is data that's reliable. It doesn't always have to be on huge scale, but it has to be relevant to disease. Clearly we, like everybody else, we ingest and we build models on freely available public data because everybody can do that. But what we spent more than the last decade doing is building these perturbational maps. I'll maybe just walk your audience through what that actually means and I'll use the neuroscience collaboration we've got as an example because I think that's the secret here is like, doesn't matter how good your models are. And remember, the models get better. Like seemingly every hour, every day somebody brings uh, out another model that sits on top of something else and goes, this one's better than the previous one. And that probably is true, but then what really matters is the data. That's that you're asking the question of. So just a quick teaching. So the way I think about a map or an atlas of biology is like, is what we're doing is we are, uh, systematically perturbing cellular systems. Now that could be via, uh, genetic manipulation. So using crispr, it could be we could be amplifying a gene, but we can also perturb it with compounds or biological reagents. And then what we do is we then take a snapshot, a photo, because that's what it literally is of those cells and it could be a bright field image or cell paint where you stain them, but basically both work. And then what you do is you then ask using then deep learning methods, you embed those, the resulting images into what we call high dimensional space. Because as a human I could show you two pictures, you go with that, that, that looks similar to that and, and that's, that's partly the underpinnings of it. But what computer, computer vision will do is that can look at many more features than a human eye can see, but it's still the same kind of context. It then just integrates those kind of perturbations into an embedding space. And so what you then see is like, okay, is that, can I, you can ask questions of if I knock down gene A, are ah, there other genes that actually look visually, that look visually similar? Uh, so then the following question from that is, well, if they're phenotypically similar, are they mechanistically similar, you still need to go ahead, go away and prove that, that they're the same. So that basically, so what happens is that genes or compounds for example, that produce similar phenotypes, they end up near to each other in that multidimensional space. And that allows you to uncover functional relationships that aren't always obvious from like looking say in the literature at say pathway or interactome. So for me a map is an idea generator, but grounded in human genetics. And I think that's the power of the perturbational maps that are unique to recursion. They're anchored in human genetics, they're unbiased. That's the critical piece. You can ask a question across a whole genome. By definition you uncover biology, you expect that makes you feel great. You go, okay, I expected to see that interaction that gives me some sense that this is true. But what's really interesting is that biology wouldn't have predicted from the other means. And then you can test that in a disease context and in the neuroscience setting, link the. Well, actually holistically we've fully mapped something like 3 trillion searchable gene and compound relationships across our data sets. And then if we step into the neuro space, because I think this is, I think people need to understand how difficult this is. This is, even today is not a trivial undertaking that in order to build the. Let's, let's take the first one, the NGM2 map, never mind the microglia one, is that we had to manufacture over a trillion, one trillion human IPS derived neurons just to give you some context, that's about the same number of neurons as in, as in 12 human brains. We had to manufacture those. And then once we've done that, uh, and built stable protocols, we then knocked out one at a time, 17,000 genes. I think we captured something like more than 30 million images. Like just, just the scale of this, um, experimental, this is not predicting, this is the experimental data on which we predict. And I think the, the two maps that we've built, so one in neurons and microglia, which are the brain's resident kind of macrophages, they're the, they what drive probably a lot of the neuroinflammation is like we've, we've had, we've had to, we spent 10 years building a capability across different maps to even get to even be able to think about doing this. And I think what's more exciting is like, so the maps are the foundation, they're the foundation that allow us to ask questions, to surface biological insights. And we've, yeah, we've produced six for our partners and produced kind of like over a dozen kind of for ourselves and for others. And I mean the critical thing here was the, the recent announcement, literally in the last, in the last month where we've leveraged those, those maps in a neuroscience setting to uncover, um, and then to at least in an in vitro setting, validate a novel target in a neuro generation setting. I think anybody on this call on the old podcast that works in neuroscience will understand how difficult that is. Like, I come from a neuroscience background where I started my career and uh, most of the time we're all, we're all working on the same 30 or 40 targets. And I think, I think. So the strength here is the depth of the data is that the, it's unbiased and it's allowed us to kind of like it's reusable as well. So we don't, it's not just one target we're identifying, we're identifying many. And so I'm um, quite optimistic over time we are surfacing, yeah, novel, generally novel insights that will ultimately turn into, into novel medicines for patients.
Speaker B: In one of our most recent episodes with, with Brent Richards from five Prime Sciences, he had this really great analogy described picking, nominating, uh, a candidate, starting a clinical program as like you're walking up to the edge of the cliff and you're about to jump out. And, and the new target is, is one of the hardest versions of that jump because there's nobody there with you almost always by Definition. And I would love to just get your perspective on what you need to see to have the conviction to walk up to one of these novel targets and, and, and make the jump with your, with your parachute, your hang glider or whatever it may be, because they're, they're harder for a reason. Right. But the opportunity is much bigger. So I'd love to get your sense of what you're looking for today before you go after one of these novel targets.
Speaker C: I think the critical thing is that you, you need it, you need a lab, you need it, you need an experimental validate. So by definition what we're all trying to do is we're trying to predict more and test less. We really are. But at some point, like any lawyer will tell you, you need to trust but verify. And so having the experimental infrastructure we've created, recursion is something I'm deeply proud of. So because every kind of map based hypothesis that we generate has got to be validated in a laboratory setting, either on our own or with our partners. But I think about it as four. Yeah, four things. So I think the first one is that when we surface these novel insights is that, well, okay, the map's producing a really interesting idea, an interesting relationship, but uh, we then need to validate it in an orthogonal setting. So basically the relationship that we see between a gene and a function in a map, we need to confirm that in other experiments. So this could be a different perturbation, they could be different cell donors or backgrounds to basically to give us that first level of confidence that this is actually not a modeling or an experimental artifact that is potentially real, we need to show its disease relevance. So sometimes we build systems that we would argue that uh, are directly disease relevant. They, they essentially collect a classic double perturbation where you engineer in or you use a patient derived system. So you know that cellular context represents disease. But sometimes you, what you're looking at is you're perturbing kind of a wild type system. And so you, the next thing you must do is basically show that the signal that you've just identified is actually connected and associated with disease. So one of the things we did as part of the neuroscience effort is then is then to demonstrate that the perturbation we identified would actually, if you like pheno, correct an engineered system. Because the, so we test, basically we're testing in kind of multiple models. That's the key thing. Having surfaced an idea, shown that it's kind of, that it's relevant, uh, uh, in orthogonal systems, you then go, okay, but how do I turn this into a drug? So then you have to look at kind of tractability assessment. So because it could be biologically interesting but might need a small molecule or another modality to get into there. So that's where AI plays a big role. We can predict a significant amount now about how tractable or not a given protein is, how to attack that protein. Do you need to degrade it? Is it a peptide target? Is it is an antibody? Does it small molecule and then the last thing. And it's no minor thing, but I think, I always think the best validation of a technology is third party validation. Not what I say about it. And so because this is part of work, is that there's, there's a formal option here as well. It's like so, uh, highly scientifically credible and talented scientists at Ross Genentech, they look at the data, they look at it, they bring their own expertise to the table. I mean that's why part of these, these, these, these kind of partnerships, I think they're, I think they're really strong proof points of what we're doing. It's not, it's not just me saying how great we are. It's like it's our partners going yeah, this, this is long way to go. But this is, this is, this is surfacing novel targets.
Speaker B: Yeah. And that's right. I may be related to this. I'm curious how you think about depth in a therapy area versus building the infrastructure that can work across a lot. Because it's a classic challenge. Right. And there's no, there's not necessarily a right strategy but there may be better strategies at certain points in time. I'm wondering how you think about that.
Speaker C: Yeah, no, it's interesting having yeah. Worked at, yeah. Big pharma at Merck, uh, 14 great years a great CRO as a evotech and then Accenture recursion, each one of those. I was clearly very different kind of business models. I think recursion think proudly sits in that middle space. We are, we are both a platform, but it's an applied platform. It's important for us that we can prove our platform can generate novel medicines. I think I'm a big believer and as we said at the start of the call, have the scars that come from clinical development. I think you have to have skin in the game. I think you have to be trying to develop drugs to understand how to then apply methods to that. So I think your answer is that I think pure platform companies have breadth. So we like others, we can generate these hypotheses across diseases, but really deep disease specific kind of clinical and biological data, you have to build it. So Recursion we've spent effort and time building at an internal level, really largely focusing on oncology and more recently on immunology. We built the corresponding translational and clinical infrastructure to go with that. Because as a, as a platform company that's focused on really decoding biology and trying to find transformative medicines, it's like, how do we do that elsewhere? I think partnerships allow you to do that. I think we all know how difficult neuroscience is and we've got fantastic partners at Ross, uh, Genentec. And so what the partnerships allows a company like recursion to do is to play in other spaces, but leveraging the disease capabilities that exist in our partners. So I think it expands our reach and I'm proud that we do that because it allows us to expand our reach to more patients.
Speaker B: We are living in a really interesting time and I think not least because the scaled AI companies, Anthropic OpenAI are spending an increasing amount of time in life sciences and in drug development. I even think just a couple days ago Dario, the CEO of Anthropic, said something to the effect of we may or will cure most diseases in the next five to 10 years. And these are very credible people. And they may be wrong on timing sometimes, but they've been pretty right over the last few years. But probably famously, drug development has been very difficult for technology companies to crack. And there have been promises in the past, but you have credible people like Demis hassabis@ um, eddiemind and I just would love your, your take on this. Are we on the precipice of actually everything about our industry may change fundamentally in the next five years, or is it maybe something in the middle? There's going to be a significant improvement in productivity, but we're not going to see a, uh, flood of new approved molecules. I realize this is probably the billion trillion dollar question, but I can't help but ask it because it's on my mind.
Speaker C: I think from my perspective, I think technology in its various forms, including aiml, has transformed what we're doing. Perhaps not at the speed that some have promised, if we're honest. And I think that yeah, even five years might still be a big ask. If you think about like what the real test is. The real test is like, does your drug working a regulated clinical trial and unfortunately like AI cannot change the laws of physics. If, if you're running a regular, if the FDA says, okay, design your clinical trial. This is the endpoint you're looking at. Uh, and that endpoint M is going to take four years to measure. AI will not make that endpoint M one year. I think what AI will help hopefully do is like try and identify other endpoints M that are meaningful. So, so the true, the true test of any technology is always going to be but did you make a difference? Did you, did you identify a new drug? So if I, if I think about, I would say the next five to 10 years, but I think I'm, I think we'll see a lot of, a lot more proof points than the next five. Patrick. So I would go, what would I look for? I would look for a meaningful number of. Because this is not just about one and done. There's no point that, that doesn't like one for me is not proof. 2, 3, 4, 5 is so repeatability is proof. So I'd look for a meaningful number of first in class drugs whose, whose target or mechanism wasn't covered by aiml. That's still the biggest lift because that's, we and others have got kind of targets in our pipeline, but they're only just really going into clinical trials. We're not just talking about showing efficacy in A Phase 1B, like efficacy and registration studies. I think that's kind of, that's why I think we're proud of like our MEK inhibitor in the FAP studies. The first example we've shown in a, in a, in a devastating disease, kind of real clinical impact. Going back to your second point, I think this, this is, this is being more demonstrable which is like repeatable reduction in time and cost. Everybody wants novel, everybody wants something that humans have not done before. But at the same time we also want to be more efficient and I think we're seeing that kind of playing out. So really durable cycle time progression. But not just, uh, here's one project and it only took me a year. Like, no, do it again. Do it. Show it across your portfolio. Show it. Show you having impact across the portfolio. I think just, yeah, just to blow our own trumpet. I think in the neuroscience space, I think over the next five years, I think about we've just had one target accepted, five to six years. I think that's enough time to start to see, to demonstrate that we can actually find drugs for these and actually get them into early clinical testing. So we won't, in five years from now, the target that we kind of surfaced a month ago, that won't be in a phase three registration study. But I'd like to, I'd like to think that the work we're doing in the neuroscience space with our partners will start to surface, like be able to talk about generally novel insights. We're finding that. So I'm quietly optimistic. Whether it's five years or not, I'm not so sure. But I think, make no mistake here, Patrick, I think, I think we are transforming industry. I think it's just that the cycle times for what we're doing are uh, kind of longer than say, building a chip or making a new tv.
Speaker B: Yeah, yeah, absolutely. As we wrap up here, I have one more question that I love your perspective on, which is about how the actual work and the uh, the field for scientists or for, for drug hunters is changing. What skills do you think are most important right now? And what are the things that are, that are least important? That if you're, you know, transitioning or starting your career or just, just trying to do your best work right now, what's, what's most important and has changed?
Speaker C: That's a really relevant question for me. I have a son who's 21 who's doing mathematics. I have a daughter who's 18, just about to start, uh, a natural sciences degree. So uh, they're following in the kind of same domain as me. I think what I've learned and continue to learn at recursion is that I would tell any student, build a domain expertise, don't be a jack of all trades. But unfortunately you're also going to need to be fluent across the biology, chemistry, computation, access. Doesn't mean you have to be a fabulous kind of coder. I'm not. Thank goodness for kind of Claude code. But it's this, we use this word bilingual a lot and that's, that's what it means to me. It's like, yeah, you need, you, you've got to be, you need the specialist. I want, I want the kind of the chemist, the biologist, the people that kind of trained that really understand that. But at the same time, like, unfortunately you can need to have some, some breadth. There's uh, so think about, let's say if you're a lab scientist, so yeah, be, be computationally literate. I said I don't expect. And we don't expect those people to be like ML engineers. We hire those domain experts. But you need to have a good background in, I think in statistical methods and computational fluency. Because ultimately what I want those scientists to do in tandem with their uh, kind of ML partners is like, but is that, is that predictive relationship even biologically plausible? Like that's, that's what domain expertise teas comes in and I think separates pure platform plays. Conversely, if you're a computational scientist, so working in that machine learning space. Yeah. Unfortunately learn a bit of biology, please learn a bit of kind of chemistry again to try and meet your counterparts in the middle. I think. Yeah. Understand assay noise. Understand. Yeah. When a model is basically just is good enough in the uh, in that kind of. In an area where they've held out data. But what, but where is its domain of applicability? You need to understand the biological system because they're noisy. And I think last but by no means least is like I think, yeah. Optimistically skeptics, I think that's uh, I think we said this before. I mean lawyers always say was it trust but verify? I think if I look back at my training, I think the, the skill that kind of like, that really good scientists have is like, is really knowing when to demand kind of confirmation.
Speaker B: Yeah.
Speaker C: That yeah. Supplement your instinct with the kind of tools and so that ultimately yeah try and find, yeah be that domain expert but really, really leverage and extract the kind of, the enormously kind of powerful kind of sway the tools we have today. So I won't lie, it's a tough space to come into as a scientist, but at the same time I think it's like supremely exciting because like, I just wish I had access to some of the technologies that uh, I have today. Say 20, 25 years ago.
Speaker B: Yeah, it is amazing. That last point you made actually, uh, that landed really clearly with me. It's something I've been thinking about lately because when the models, the large language models were a little bit less capable, you kind of knew they hallucinated regularly and, and you could only trust them really with, with heavy scrutiny. As they've gotten better, I think it's easy to. And, and they're so much better. It's easy to not have that level of curiosity or to push to really, really go deep and understand. And uh, I found that that becomes much more important that you don't take the first answer and you push and push and push because the first answer is probably 80% right, but it's definitely not 100% right. And the, and the devil really is details. So having that curiosity to go five levels deep and not just take the first answer that comes back and run with it is so important.
Speaker C: You have to, it's the way that you optimize agents because an Agent's essentially a learning tool. And I think one final point I'll leave you with is that we often forget this is that I remember maybe a decade ago, a publication came out from. I think it was a large pharma group that essentially said that in about 70 to 80% of cases, the kind of key biological result in a paper was. Could not be reproduced across labs. Now, I'm not for one moment saying that there's kind of any kind of mal. Mal intent there. I think it just shows you how noisy biology is. So I think we should never forget that. Remember, we're all. We're all building models on data and in some cases building on. On public data. So I guess the, uh. Was it buyer beware? I think. I guess that would be the word I would use is that as you. I think you. You nailed it. It's like, go into the detail kind of like, yeah, use these tools. I use them all the time. Use these kind of these LLMs. But then ask the question you get. But is that true? It's like, where's that data come from? Can I trust it? Because ultimately, I think the big undoing of AI is not. Is not AI itself or the ML. It's like it's just how good is the data or how bad it is?
Speaker B: Yeah.
Speaker C: And I think that's. That's the. That's the real differentiating.
Speaker B: Great. Well, David, this was so much fun.
Speaker A: I learned a lot.
Speaker B: I got exactly what I was hoping for from this and more. So thank you so much for sharing your time with us now.
Speaker C: You're welcome. Thank you, Patrick.
Speaker B: Great. And everyone, thank you as always for listening and we'll see you next time.
Speaker A: Thanks as always for tuning in to the genetics podcast. If you enjoyed today's conversation, the best way you can support the show is by sharing it with a friend or colleague who might find it interesting as well. We'd also really appreciate if you could subscribe, rate and review us on Apple Podcasts, Spotify, or wherever you listen to podcasts. This helps other people discover the show when they're searching for content on genetics, precision medicine or biotech. We always want to hear from you as well.
Speaker B: If you have feedback or questions or want to be featured on the show,
Speaker A: you can email us@podcastsonogenetics.com you can find me on LinkedIn. Patrick Shore, you can find us as well Sonogenetics on LinkedIn or reach out on Instagram. Ah, the genetics podcast.
Speaker B: And finally, a special thank you to
Speaker A: the team behind the show who makes this possible? Joy Ismail produces and manages the show. And James Pierce from Selective Frequencies for his expert audio engineering.
Speaker B: I'm Patrick Short, your host.
Speaker A: Thanks again for listening, and we'll see you next time on the Genetics podcast.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.