Data Science Leaders · 2026-06-02 · 32 min
Key moments - from our scoring
Substance score
50 / 100
Five dimensions, 20 points each
When Steve Johnson first entered pharma as a PhD physicist, he was tasked with an unusual mission: salvage three decades of discovery chemistry data scattered across incompatible media - hard drives, tape drives, CDs, VAX machines, Unix and Windows systems - piled to the ceiling in what staff called 'the cave.' His six-month project, completed in two, involved mounting hardware, writing Python scripts to understand file metadata, deduplicating data across formats, and building a searchable web application to rescue data at risk of permanent decay. But Johnson argues this problem never went away. Modern pharma companies still struggle with 'data swamps' - data generated for single use cases without foresight for future analytics, metadata captured poorly if at all, and AI pilots that fail because they lack clear ROI or end-to-end process integration. At Dash Bio, his company focuses on bioanalysis and drug development automation, Johnson champions operational KPIs (rare in R&D), end-to-end ownership, and treating data as a productionizable asset rather than an afterthought. The core thesis: most pharma AI failures stem not from weak models but from treating the model as separate from existing workflows, poor upstream data quality, and absence of measurable business outcomes that justify transformation.
The cave was a dark, dank room at a pharma company storing 30 years of discovery chemistry data across every conceivable format - hard drives, tape drives, CDs, VAX and Unix machines - left behind by a retiring data manager with no documentation. The data was decaying, potentially duplicated, legally risky, and inaccessible; Johnson's job was to recover, curate, deduplicate, and make it searchable before media failure destroyed it permanently.
Most fail because they lack clear ROI, assume the model itself is the solution rather than part of a larger workflow, and skip reproducible data pipelines; data scientists clean data in ad-hoc ways that cannot be operationalized, and the resulting model has no integration path into existing pharma software and processes.
Metadata capture is critically weak; while animal studies and clinical trials produce relatively small datasets, the contextual information about subjects (species, strain, genes, dosing, biology) is poorly recorded, making aggregated data useless for future AI or secondary analysis without this context.
Without measurable KPIs, executives cannot make data-driven trade-offs between data science, clinical development, and other teams; projects proceed on vibes and organizational momentum rather than measurable impact, and companies cannot identify which bottlenecks or efficiency gains would most benefit from investment.
Dash Bio owns bioanalysis and drug development end-to-end, focusing on industrializing traditionally bespoke consulting work through technology, automation, strict KPI discipline, and proper data management with embedded AI models - aiming to reduce the cost and time to bring drugs to market.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains a handful of genuine practitioner insights - particularly around non-productionizable data pipelines and the metadata gap in R&D - but these are buried under extended storytelling and conversational filler. The insight-per-minute rate is low for a 32-minute runtime.
all of that data aggregation, data cleaning is not reproducible, it's not productionizable because you didn't do it in that way
how much data lives In Excel, how PowerPoints are used as essentially databases like this same problem persists and I'd argue hasn't gotten any better
The reframing of poor pharma data stewardship as economically rational behaviour is a genuinely non-obvious angle, but most other observations - pilots fail without clear ROI, models alone aren't enough, China is a competitive threat - are well-circulated takes in the industry.
the behavior of how most biotech companies treat data is a rational choice in some sense because their objective is to prove out a particular scientific hypothesis, a target, a uh, potential drug. Fast
You might walk away with a process that's just different. Not better, not worse, just different
The guest is a credible hands-on practitioner with a PhD in theoretical physics, real technical experience solving legacy data problems, and a leadership role at Moderna before founding Dash Bio; he is not a career podcaster or pure thought leader, though Dash is early-stage and his scale of impact remains limited.
in my time at Moderna, for example, like, hey, here's a problem, can you solve it? The first question we'd have to ask is, can we even find the data?
I can make these trade offs and say, okay, I can invest my data science time in this project here which has this perspective roi because I measure that
The cave anecdote is pleasingly concrete - specific media types, room dimensions, a compressed six-to-two-month timeline - but broader strategic claims such as the $5B-to-$50M cost reduction target and the China clinical-trial assertion are asserted without data or sourcing.
it was slated for about six months to get this thing done. That was the length of the project. I had done all of the work to collect the data in about Two months
if we can change the cost to, to get a drug to market from 5 billion to you know, 50 million, right, 500 million even
The host structures the conversation well and asks a few substantive questions about AI pilot failure and KPIs, but consistently echoes or flatters the guest rather than probing claims, and there is no meaningful pushback or productive disagreement across the full episode.
There are a lot of pharma companies that have dozens of pilots running that actually never make it to production. Why do you think that happens?
I really love what you just mentioned. There's a lot of humility in, in talking about the arrogance of the, uh, of the founder
Computed from the transcript - who did the talking, and the words that came up most.
Dave Johnson ’s first pharma job wasn't in a lab. It was in a room, dark and cramped, stacked floor to ceiling with 30 years of decaying discovery chemistry data on every storage medium imaginable. They called it “the cave.” A physicist turned data leader, Dave has spent two decades solving the same problem: scientific data is generated for one purpose, poorly captured, and left abandoned. Now co-founding Dash Bio after years leading data science at Moderna, Dave joins Thomas to make a case the industry would rather not hear: AI pilots keep failing because organizations can't understand their own data or tell whether anything they build improves how they work. Listen in to hear: Why "the cave" is still an accurate description of how most pharma companies manage data Why clinical trial numbers mean nothing without the metadata behind them How the US needs to treat China’s clinical trial volume as a wake-up call
Transcribed and scored by The B2B Podcast Index.
Speaker A: Before the show starts, we want to let you know that Data Science Leaders is coming back in 2026 with all new episodes and a renewed focus on the humans that are driving AI forward. We're grateful to be part of your podcast lineup and we ask only one thing. Please take a moment to rate and review our show right now on a platform you're already listening on. And please don't forget to check out this year's REV Roadshow in Philadelphia, New York and London. We're putting together these events for you, Data Science Leader, to learn how to lead with impact. Find out more at Rev Domino. AI uh, thanks for listening.
Speaker B: When I walked into the cave, my first thought was, what the hell is this? Right? I knew nothing about pharma. This is literally my first project in this industry. I had no idea. And so in some sense I didn't know better that this was, uh, you know, maybe this is the way all Discovery Chemistry sites are.
Speaker A: That's Steve Johnson, co founder and CEO at Dash Bio, a company on a mission to industrialize drug bioanalysis and development to bring therapeutics to market faster than ever. Dave's journey in the biopharma space had an unusual beginning, and one that parallels many of the data management problems he sees holding an industry back from his true potential today. I'm Thomas Bean, and this is Data Science Leaders. Beginning with an interest in science and computers from an early age, Dave later earned his PhD in theoretical physics. After transitioning away from academia, he was steered towards a career in biopharma. One of his first roles in the industry involved being brought in by a Discovery Chemistry group to solve a legacy data problem that was anything but usual. What was this problem? What were they asking you to figure out?
Speaker B: I was brought on site. They walked me into this room. It's kind of dark, it's dank. There are boxes piled to the ceiling, there are computers everywhere, lights flickering on and off. And what was in front of me in this, you know, probably 10 by 20 room was 30 years of discovery chemistry data and every possible medium you could imagine. So hard drives, tape drives, CDs, DD's, active computers, Linux drives, VAX machines, Unix machines, Windows machines, everything possible, all piled up to the, to the ceiling of the room.
Speaker A: A newcomer to pharma, a room piled to the ceiling and three decades of discovery data. The person that accompanied a young Dave referred to this room as none other
Speaker B: than the cave, because that's how it felt when you walked into that room. This, uh, was the product of, you know, the, the guy who was responsible for managing this data for the 30 years, who had just retired and left this essentially toxic waste dump, um, of data that no one knew anything about, didn't know what it was, didn't know if they had regulatory obligations around it, didn't know if we had it stored anywhere else. And they just needed someone to solve this problem, make it go away. So that was my job, was essentially collecting all this data, curating all of it, finding out what it is, cleaning it up and purging the cave from the company.
Speaker A: And, uh, did it give you more information about the scope of the problem? Uh, was it really unannounced, here's the key and figure it out?
Speaker B: Pretty much was pretty much that. I mean, the problem, the cave, was kind of hopeless in their perspective. Um, no, they just walked in the room and said, here's the problem, can you solve this? Uh, and it was, it was interesting because in some ways, and this was a kind of foreshadowing of my career, I had this really interesting background as a child of like, falling in love with computers, but falling in love with science and technology and just always dabbling, playing with that. So I build my own computers from scratch as a kid. I, you know, compiled my own Linux kernels in college, you know, as we all did, right. And, uh, so I was kind of like the perfect person to solve this particular problem. Like, I had dealt with a lot of this technology and so I said, all right, let's go. And I had them procure me, just give me the biggest, fattest machine you can with like four DVD drives and, you know, find me a tape reader and stuff. And, uh, I just attacked this problem. So I was pulling hard drives out of machines I was mounting them on, I was trying to read the data, pulling all the files from. This was, uh, really just step one of this problem, uh, because then you just have a bunch of files and no one knew what they were. What's the context of all of these? So then I had to do a lot of work, I had to write a lot of Python code at the time to look into these files, to try to understand metadata of them, to deduplicate them, because you would find the same data on some tape drives and then DVDs and stuff where he had moved them forward. Uh, and so it was just a, you know, really all encompassing project across a lot of dimensions. Uh, it was slated for about six months to get this thing done. That was the length of the project. I had done all of the work to collect the data in about Two months, all collected in a database and so on. This is how data is treated in a lot of pharma still to this day is that so much is just not properly stewarded or cared for.
Speaker A: It's almost like an Indiana Jones kind of aspect to the story. You went from physics to archeology, but this still a lot of parallels to what data scientists are still facing today. But first I want to talk about the outcome. Once you were done, you came out, out of the cave. The cave is now clean, understood, explored. What was the outcome for the scientists? Like how easy was the access to information.
Speaker B: So the access to information from this obviously was super fast. So after I collected all the data, I built a whole web application, front end with a database where you could search across all this data, you know, find this particular data set of that. So that was certainly valuable for the scientists looking back and trying to find data on these compounds that they've created over 30 years. The bigger value I think was for the organization as a whole. Essentially when this person retired, they were left with just a bomb, like a toxic waste dump. And they had, there was nothing they could do with it. Right? I mean, you have media there that is actively decaying. And so there was quite a lot of records that were corrupt. I'd have to run like, you know, routines to try to recover some of this data. But tape drives do not last forever. They're already starting to decay. And so they're left with this mess in a room that they can't touch for fear of losing data that may actually be important. So simply taking it all, boxing it up, sending it off to Iron Mountain or something was not enough. Like that data had to be rescued. So in a sense it was, it was solving this problem, making this mess go away for the management, for the company, for the entire department of people.
Speaker A: Yes, on one end they have one free room, but most importantly, a lot of data that they will keep on forever. Because I guess this data becomes very useful even years after.
Speaker B: Well, and it's the risk and the liability of having data there that you have no idea what it is and whether it's important, whether you have to keep it or not.
Speaker A: You walked into a room full of scientific data, came out with this data accessible. The challenge was not even analyzing it at the beginning. It was just simply being able to read it. What's left of these kind of problems today in life sciences companies? Do we have caves somewhere in the cloud or somewhere that still exist that need to be explored?
Speaker B: Every company seems to have its own version of the cave, whether it's virtual, whether it's on prem or not, they all have the cave. We in the data science world call them, um, data swamps and so on. It's the same problem that we all face. And I feel like every job in my entire career here has been facing this exact same problem. And it's not any better today. I mean, the way data is managed today, um, the way raw data is stored, the way it's archived, if it's archived at all, how much data lives In Excel, how PowerPoints are used as essentially databases like this same problem persists and I'd argue hasn't gotten any better. In some ways it's probably gotten worse. And so you alluded to it earlier, it sounds like a lot of what data scientists do. I mean, we would be asked, you know, in my time at Moderna, for example, like, hey, here's a problem, can you solve it? The first question we'd have to ask is, can we even find the data? Do we have enough data? Can we even read the data formats that are coming off the instruments? So it's the same problem repeating itself over and over.
Speaker A: So let's go, uh, one step later in the process, or one step higher in a stack if you want to, and talk a little bit about AI or data science. There are a lot of pharma companies that have dozens of pilots running that actually never make it to production. Why do you think that happens? Data might be a factor, but why do you think that happens?
Speaker B: I think the biggest reason is that they're not focused on the outcome, the end solution that they're after. There's so much FOMO and it's justifiable fomo. I mean a lot of the, the LLM capabilities are just truly transformational. It's fun being in a new company building one from scratch. You build kind of LLM native and you can experience the power of it directly, so it's warranted. But if you don't have an end goal in mind to say, when I implement this, you know, here's my before, here's my after, then it will ultimately fail. Like you have to have strong enough ROI that will endure beyond the fomo, will endure beyond this year's budget.
Speaker A: What breaks is it really? Ah, this, not having the vision, or is there more in between the outcome that's expected and really where the work happens and the pilot itself? Any glaring gap or pitfall that you've seen in your experience?
Speaker B: Well, I think people assume that the model is the problem. Right. And the model is just a piece of the problem. You know, we think about this at dash. It's a huge part of our thesis as a company is to own a problem end to end. And you're going to leverage a lot of different technologies, whether it's robotic automation or data science, AI models or just plain software integration and plumbing and unsexy stuff no one wants to think about. When you look at any use case that matters in a large organization like a pharma company, there's already something there. There's a process that exists, whether it's integrated or not, manual, whatever it exists, uh, and there very likely is some level of software. I mean, enterprise pharma companies have thousands of pieces of software covering countless workflows. You have to operate within that vacuum. And so just in that, in that um, system. Right. So just having a model that predicts a piece by itself is useless. It has to exist in that framework. And so, uh, that's where I see a lot of the pain and frustration is, is not understanding how it fits into the workflow, how some person's workflow is going to change, you know, before and after that. The other big challenge I see in these deployments is around the data quality, which is you'll see data scientists say, okay, here's a really valuable problem that we're going to solve and we're going to collect all this data from all these places. We're going to spend all this time cleaning it up and working on it. Got a nice pretty data set. Look, we get this great model, it's highly predictive. Then what do you do with it? You actually have no path to productionizing it because all of that data aggregation, data cleaning is not reproducible, it's not productionizable because you didn't do it in that way. The problem then exists upstream. Right. Of um, how do you actually get the data into your model in that way? Now doing such a project where you can demonstrate, hey, this model works if we can clean up our upstream systems is actually really valuable if the ROI is strong enough to then drive that upstream change to make it work.
Speaker A: So part of it we're coming back to the initial data problem you found in the cave, which is where is it coming from with the format?
Speaker B: And yeah, I mean that's why, that's why I said every company still has their caves, whether they're virtual or not. Right. I think that the challenge that people have to keep in mind is that data is generated. For one use case in mind, it's typically not thought about in terms of the subsequent use cases for that data. So the person who, the team that was generating the data that went into the cave, they were solving a problem that day, which is trying to find chemicals that have particular properties they were interested about today. They had no concept of the future use cases that might be valuable. Right? And you think about if they had only aggregated uh, that data, the kind of AI chemical models that you could have built on top of it, which again are now lost to the sands of time here. It's the same thing as we think about any enterprise process, any data generation, right? The process whereby someone is creating this is meant for one intended purpose. They're not thinking about the reuse of that. So I always encourage people as they design and think about systems to kind of think one bubble of complexity around it of how do you design the data and the systems and the processes in a way such that other people could make use of them.
Speaker A: Which I guess has incredible applications in a, in a, in a scientific context and especially in uh, in life sciences because you could have made an experiment that validates something completely you have never thought about. So it's, it's something that's kind of vital to the, to these life sciences companies. So there's probably an advantage, a huge advantage to, to deri.
Speaker B: It's a trade off. I mean it's, you know, I like to think, you know, in another life I was an economist and I like to think that markets are rational. So the behavior of how most biotech companies treat data is a rational choice in some sense because their objective is to prove out a particular scientific hypothesis, a target, a uh, potential drug. Fast like that is what the economic incentive is. It is not to collect a whole bunch of data in a pristine, perfect way for future data scientists who may or may not appear. So it is a trade off for an organization to balance their primary objectives of today with these, these prospective future objectives. Certainly large pharma companies who are generating this and have much greater potentiality for the future have more of a motivation and impetus for that.
Speaker A: On this work on the data, um, maybe from your own experience of ah, working on messy scientific data, are there some kind of decisions you can take, like can automation actually help to a point, uh, accelerate or improve this uh, process or maybe working on the uh, upstream even process of capturing the data. What's the extent of what you can control just to make sure that this data can be delivering its full promise. Delivering its full promise at the end,
Speaker B: I think the Biggest gap I've seen when it comes to kind of data collection and creation is, is actually the metadata that goes with it. So if you look at a huge amount of what happens through drug R and D, the actual numbers that come out of say, an animal study or a clinical study typically can fit in a spreadsheet. It's really not a lot of data, but each one of those data points has this, an immense amount of complexity of metadata about how the subject was treated. The biology of the subject, the genes, the species, the strains of these, these animals and humans, um, all of that metadata is often very poorly captured. So if you aggregated all of the outcomes of all these studies together, it would do you absolutely no good because you don't have all of this context that gives that data meaning. And that's the hard part, right? Building these really complex data models to accurately capture all of this in a way that's also not too painful and onerous for scientists to input. It's a real challenge.
Speaker A: I want to go back also to what you said about not a lot of these AI projects not thinking about, uh, the outcomes or maybe sometimes even lacking m real KPI's outside, maybe in life sciences of a very particular part of the business, which is manufacturing, where things get physical, so to speak. How does this lack of KPIs affect, uh, AI projects in terms of even how they're designed and how they're managed, they're operated?
Speaker B: I think it's a really great point about the lack of KPIs in this industry. And you pointed out rightly that the one place I've seen it is in manufacturing, right, where you have really defined processes, inputs and outputs. Throughout my career, I would ask people, you know, how you know, what are your KPIs like, how do you know if you're doing a good job? How do you know if you're improving? And I would get blank stares, right? It's just not the way people think about R and D processes. I think a lot of it is just due to the fact that it's not a repeatable machine. There is a lot of serendipity, there is a lot of dead ends and explorations and so on. But I do think that there are components within that where you certainly could measure. If you can factor out parts of animal study execution and kind of on time delivery of things like that, I think that would drive a lot of that, um, improvement. Right. Um, and if you think about what, what DASH is doing as a company, we are essentially that we are an operations based company. We have factored out, you know, bioanalysis, testing. And so we care very much about KPIs. And every investment decision we make is how do we improve that KPI. So you can see directly for us, I can make these trade offs and say, okay, I can invest my data science time in this project here which has this perspective roi because I measure that. Right. And so I think it has a real detriment to companies that don't have the ability to actually measure those things to know where to, to allocate those resources. And I think it's particularly hard for the folks who are, who are creating budgets on a global scale. Because it's one thing to say, hey, I have a pool of data scientists. Which problem do I point them at is very different from saying do I give more money to the data science team or do I give more money to the clinical development team? Right. That only works in the context of real business KPIs that an executive can make that kind of trade off decision globally.
Speaker A: And thinking in terms of leadership, I guess you expose also yourself to the problem of reinventing not only the wheel, but some of the problems. If you don't have the right KPIs, there might be bottlenecks or efficiency gains that uh, you might lack.
Speaker B: You might walk away with a process that's just different. Not better, not worse, just different. Right.
Speaker A: It's more, yeah, it's more of the journey you took that got you there.
Speaker B: Maybe you're high fiving around the successful delivery or something, but it's a lot of organizational pain and challenge for something that may not have actually had an impact. But you also have to remember that post deployment of a project like Post Go live, it's very common, take a very long time to truly embed a new product, a new model in a way that is having meaningful impact. In fact, you may see a regression for a period of time. It's like, okay, we've moved from system A to system B for a while. System B may be less efficient because people aren't familiar with it. People are making mistakes. It's only when you kind of fully embed it that you start to see that productivity gain. And so if you're not measuring that or worse, you have no way to measure that. How would you ever know if you're done? So if you truly have a platform M that you can use AI to discover new proteins, new small molecules and you could create a hundred of them, well, you need $500 billion to bring those drugs to market. Right. It's an inordinate amount of money. And so what I have seen happen time and time again is the actual discovery platform kind of is irrelevant ultimately because when you get to the clinic, when you get to your GLP tox studies, you basically have cash to get one of those through, right. So you pick the one that's most promising and then everyone just watches that and the entire company rides or dies on that because investors are going to watch that now, right? When you're early on, you can sell vibes like you can sell the potential of the platform. But as you get later and later in development, you switch to traditional biotech investors who care about data, you know, NHP data, GLP data, phase, uh, one clinical data, that's what they care about. So when you get there, you're not going to raise much more until you get that clinical validation of what you're doing. And the company is really valued by those assets. So the whole point of that is like, I think these, I'm a huge believer in AI drug discovery. I think there's an immense potential for it. I think the science is real. I've seen it myself. I've had teams that have developed these. But until we do something to transform the economics of development and how slow and expensive it is, we'll never realize the true potential that's there.
Speaker A: And that puts also the focus on the processes as we were, as we were discussing before, which become incredibly important. You need to measure them the right way. You need to have your, uh, you need to have faith, trust actually in this, in this process.
Speaker B: Yeah. This is where the lack of focus on operational excellence comes back to bite you. Right. Because if you're rewarded essentially for speed to data tranche, there's a whole lot of inefficiency you can have on your way to get there. Right. So as an industry we're just not doing a good job of being efficient, of reusing resources. Every biotech is starting this problem over and over again, solving the same challenges
Speaker A: and hence the, the creation of some solutions or some, uh, that's, I guess that's also where technology can help.
Speaker B: That's, that's exactly right. And this is why you, uh, know myself and my co founders founded dash, was the idea that the cost to bring drugs to market is just ignorantly expensive. And we see this huge potential, this huge wave of drugs and discovery that we have insufficient capital, insufficient capacity and development to do anything about it. And so the whole mission of DASH is how do we transform development moving it from this kind of highly inefficient cottage industry of kind of bespoke consulting type work to an industrialized product centric tech first way. So how do we use technology to radically accelerate and radically reduce the cost to bring drugs to market? Right? So if we can change the cost to, to get a drug to market from 5 billion to you know, 50 million, right, 500 million even, how many more drugs could we take? It? You could actually take a whole bunch of shots on goal, right. And then promising technologies where they just pick the, the wrong kind of indication for their phase one study and they're out. Like you could actually survive some failures there and actually progress a platform for it. So that's why we found a dash with this belief that we can reinvent this CRO market and bioanalysis specifically in a way where we focus on these KPIs, we kind of own it end to end. We have proper data management, we build AI models into this end to end.
Speaker A: And then you open uh, the aperture to a uh, very interesting world where indeed more drugs can come to market with the same safety that we have today. But just the whole system becomes more efficient and we all benefit from it. We just spoke about the cost of taking your drug to market and have uh, this process. If we look at the world market itself, while preparing, we were discussing that uh, China now runs more clinical trial than the U.S. what does that also mean for the future of biotech in terms of what are the potential impacts? Is it kind of um, sorry for using this expression, kind of an arms race or um, what do you think is going to happen?
Speaker B: I think we should view it as an arms race in that way. I think we absolutely should. We are uh, dramatically less efficient and there's financial reasons, cost of labor reasons, there's kind of structural reasons about how the health care system is in China. I think we blame a lot of regulatory compliance and there's some truth to that in the U.S. but I think it's mostly overblown. Uh, I, I think if we are serious about retaining our leadership as you know, the, you know, drug developers of the world, like the most innovative companies, the biotech center of the world here in Boston, Cambridge, then we do need to view it as an arms race because that is how China is viewing it. Right. They're investing huge amount of capital, a huge amount of focus to pull that. I see some resignation from people like oh yeah, it's fine, like we'll just, well just in license from them or so on. But I mean we can look at the kind of national security elements of that, of essentially not having control of our own, you know, drug supply system and so on, and the risk that come with that. So I, I think we need to take it very seriously. I'm very concerned about it. I know plenty of other people are very concerned about it, uh, including the, the government. Um, the only way we can solve that is with technology though. Like, we cannot compete with China on labor costs. We can't, uh, but we can compete on technology. And we've time and time shown America's ability to do that in a tech first way. And I think we have a real opportunity to do that here. But we have to have the right focus and the right incentive structures to
Speaker A: do that, knowing they also have some advantages in terms of scale. They have access to a population that is absolutely massive. So a trial, uh, they can validate even by our standards, uh, in a much better way.
Speaker B: It's not just on the operation side of like running these trials, but it's on the actual patient recruitment and it's how their hospital systems are structured and the ability for them to get patients in very quickly. Absolutely.
Speaker A: We started with physics, then we spoke a little bit about archaeology. Now with Dash, you're adding another element, which is the, uh, entrepreneur. Ah, and the founder element. Uh, what do you bring from the person who opened the door of the cave or the kid who was, uh, like working on computers? Um, what's, uh, what would also drive. What do you take from all of these experience into, into creating Dash and, and having. Creating such a, uh, strong potential with Dash.
Speaker B: You know, I landed at Moderna after that, the stint in consulting in large pharma companies. And I loved it for so many years. It, it was like riding a rocket ship. It was incredibly difficult growing myself every year, growing my scope of what I'm doing, challenging myself. And year over year, my job was never the same. And I could really feel the level of impact I was having. So in some ways, that same kind of impact and ownership and drive that you get as an entrepreneur, I had that for so many years at Moderna. And so it wasn't until the company was very large that I felt like I was back in large firm again. And my efforts felt diluted. And so I just then reached that point. I was like, okay, it's time like this. This arc has been amazing, but it's time to move on. I'll always keep it in my, my genome, my, my rna. Um, but then I decided to, to found a company kind of, again, not intentionally. I, I did a very methodical exploration of what do I, what do I want to do. I have these talents, these gifts for science and technology and this hybrid of these two. I love driving, leading teams. Uh, I looked at a bunch of different companies and just came to the conclusion that there's just the combination of tech and bio is just incredibly limited in pharma unless you do it yourself. And uh, so it was like, all right, I'm going to do it myself, I'm going to start something. You know, I picked up my, my co founders. We kind of analyzed like, given our background, what is the problem that we can have the biggest impact on. And we settled on tackling the CRO space, settled on, you know, bioanalysis as a really great starting point for this. Uh, and it was incredibly difficult to convince investors that this was a good idea. Right. Um, which, you know, it's, it's how every, um, you know, every scalable, new disruptive play comes is, you know, being contrarian, but, you know, ultimately. Right. And so the fact that people didn't think it was the right idea is not a problem. Like you have to have that kind of thick skin as an entrepreneur. Not saying it's easy, uh, because the difference between an idea that will ultimately work out to be great and an idea that's a bad idea is impossible to tell at that stage. It just is. Uh, but we persisted. We found some truly incredible investors who got the vision of what we're doing, could understand it. And it's been a fun year and a half, two year ride.
Speaker A: Now, last question for me and just going full circle. How do you make sure that at dash there is no cave?
Speaker B: It's really hard. It's really hard. I think there's some arrogance that anyone who starts a company has to have that you can do different. Right. You have to, to believe that the company you build is different in terms of how you organize people, how you organize processes, systems, data. Right. Um, but it's really hard. I mean there is the, the kind of, you know, entropy is what it is. You know, gravity is what it is. Like I see it all the time of, you know, the trade offs you have to make between business priorities to keep the company functioning versus the platonic ideal of data management. Right. And it's a trade off that you'll never get. Right. You know, like so many different things, like there is no perfect organization. There are just varying degrees of less optimal. Right. And so it's just something, I think you have to obsess uh, about how you structure the company, how you structure your processes, how you think about data. And, you know, I got comfortable early in my career, you know, at Moderna, working at a company that was just really growing tremendously quickly with pursuing perfection but never practically being able to achieve it, you know, and so that's a philosophy that, that I take to heart.
Speaker A: Well, thanks a lot for this. Thanks a lot also for the idea, the entrepreneur perspective. I, uh, really love what you just mentioned. There's a lot of humility in, in talking about the arrogance of the, uh, of the founder. Yes. We will do different things differently. And uh, and I really appreciated the, uh, the perspective of someone who's always been at the frontier of two worlds. I think your PhD thesis was about entropics applied to, uh, quantum physics, which is another way of bringing two worlds together. And so it's no surprise that at DASH, you bring all of the best that AI technology and data management has to provide in certain way to lead to the world, uh, of life sciences. So Dave, it was a honor to have you on the, uh, Data Science We Respond podcast. Thank you so much.
Speaker B: Thank you so much for having me, Thomas.
Speaker A: The cave, the swamp, whatever you want to call it, is an overwhelmingly large problem. Its prevalent makes meaningful assertions, not only difficult, but problematic by the day because of the sheer velocity at which data is being collected. Just as Dave knew that day early in his career, standing in a musty storage room, room loaded with aging obsolete hardware. The value was there, but it needed context, interpretation and clarity. What Dave shared with, uh, us that really resonated with me is his curiosity that took him from being this kid interested in computer and science to being the young man standing at the entrance of this room, stepping in, taming 30 years of data, getting out with understanding that's going to impact the business. And now today, he's the founder of a company who does discovery. We're using AI and data science that we all will benefit from. If you're here with us at the end of the show, please take a moment right now to leave us a rating and review. Thank you for helping us to grow the show. This is Data Science Leaders where we learn from real experiences of those building and governing the intelligent systems shaping our world. A show by AI leaders for AI leaders. This show is brought to you by Domino Data Lab, trusted by the world's most advanced enterprises to operationalize AI responsibly and at scale. I'm Thomas Bean. Thanks for joining us.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.