
DataFramed · 2026-06-29 · 49 min
Key moments - from our scoring
Substance score
60 / 100
Five dimensions, 20 points each
James Zou, a Stanford professor and head of Frontier Agents at Together AI, explores how AI agents are reshaping scientific research rather than simply automating routine tasks. The conversation centers on two key platforms: the Virtual Lab, an open-source system of specialized AI agents (professor agents, data science agents, biology agents) that collaborate in research meetings to design experiments and generate hypotheses, and DS Gym, a virtual environment for training and evaluating data science agents. Zou describes concrete successes like AI agents designing new COVID-binding proteins that outperformed human expert-designed candidates - achieving a 5-10% success rate that's actually exceptional in drug discovery where thousands of candidates typically yield only one viable option. The core insight is that science requires a fundamentally different approach to AI than customer service or autonomous vehicles: while those domains demand 99.9% reliability, scientific discovery thrives on exploration, failure, and novelty. Zou explains how retraining models away from pure imitation toward discovery-focused objectives - rewarding rare breakthroughs rather than average performance - enables AI to generate diverse hypotheses through analogical reasoning across domains. The discussion covers implementation details like having humans provide high-level guidance (participating only ~1% of the time in virtual lab meetings) while agents handle execution, and how self-improving agents use automated feedback signals to optimize their own prompts and parameters.
Yes, in specific computational tasks - AI agents designed new COVID-binding proteins that outperformed previous human expert-designed nanobodies when tested experimentally, though this required testing 92 candidate designs where 2-4 showed promise.
The Virtual Lab is an open-source system of specialized AI agents (professor, data scientist, biologist, etc.) that collaborate in meetings to design experiments and conduct research; scientists provide high-level project descriptions and constraints, then let agents execute with minimal human intervention.
DS Gym is a virtual training environment with diverse data science tasks, evaluation harnesses, and synthetic data pipelines that enable agents to receive automatic feedback signals and self-improve through either supervised learning or reinforcement learning approaches.
Science benefits from a high failure rate where 1-in-10 ideas succeeding is excellent, so AI training rewards rare breakthroughs and exploration rather than average performance, whereas customer service and autonomous vehicles require 99.9% reliability with zero tolerance for failure.
Changing the training objective from imitation-based next-token prediction to discovery-focused learning that rewards novel solutions, combined with teaching agents to find analogies across diverse domains to borrow ideas for new hypotheses.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains several genuinely interesting concepts - mode collapse in hypothesis generation, the 'learning to discover' training paradigm rewarding peak rather than average performance, and Paper-to-Agent MCPs - but the conversation moves slowly, repeats itself, and is padded with enthusiastic affirmations that dilute the information-per-minute ratio.
instead of rewarding it for average performance, which are rewarded for like in some sense like the best performance and that's actually quite a different way of training these models but that's lead to more let's say innovative behaviors
we found to be very useful is to explicitly teach these models to look for analogies because Often a lot of the best ideas in data science and science in general come from taking ideas from adjacent domains
The episode surfaces a handful of genuinely novel framings - rewarding best-case rather than average agent output, a reverse CAPTCHA to prove you are AI, and converting static papers into agent-native MCPs - but several stretches default to well-worn AI-hype discourse about agents replacing workflows and simulating societies.
to interact with the platform, each time, the agent would have to solve, like a numerical puzzle, which will be very easy for AI to do, but it'll be very tedious for humans to do
the MCP then itself contains the different tools, the insights and know hows from that research project. And the MCP is actually optimized in a way that enables the agents to be able to reproduce results
James Zou is a working Stanford professor who published the Virtual Lab work in Nature, leads Frontier Agents at Together AI, and speaks from hands-on experimental results rather than speculation - a genuine practitioner-researcher rather than a circuit-riding thought leader.
we actually published a paper, uh, so a paper published in Nature a few months ago that introduced the platform of the virtual lab
we actually tested all of these 92 candidates, and from these 92, I would say about three or four showed quite promising results and two in particular worked better than previous nanobodies designed by human experts
The episode is anchored by concrete numbers - 92 candidate proteins tested, ~1% human participation rate in virtual lab meetings, 8B/4B parameter models benchmarked against named frontier models, 12 well-known problems solved on Einstein Arena - though some claims remain vague ('a few days,' 'quite promising results') and third-party data is largely absent.
the humans, we don't talk very much and Maybe only about 1% of the time do we actually participate and speak in these virtual lab discussions
we actually tested all of these 92 candidates, and from these 92, I would say about three or four showed quite promising results
The host occasionally asks useful probing questions (success rate, specialization trade-offs, human involvement frequency) but routinely responds to answers with affirmations like 'that's absolutely fascinating' and 'I love that idea' without challenging any claim or pursuing meaningful follow-up; the result is a pleasant but largely uncritical PR-adjacent conversation.
That's absolutely fascinating that the idea of like agents stabbing meetings
I love this idea of a virtual lab
Computed from the transcript - who did the talking, and the words that came up most.
AI agents are no longer limited to automating routine tasks like customer support or report generation. Research labs and pharmaceutical companies are beginning to deploy teams of specialist AI agents capable of designing experiments, analyzing data, and proposing new hypotheses - in some cases producing results that outperform human experts. For data scientists and researchers, this raises urgent questions: Where do AI agents excel in scientific workflows today, and where do they fall short? How do you build an agent that can genuinely innovate rather than just replicate what's already been done? And what does it take to scale a single model into a fully functioning virtual research team? James Zou is an Associate Professor of Biomedical Data Science, and by courtesy of Computer Science and Electrical Engineering, at Stanford University. He leads the Stanford AI for Science Lab and is affiliated with Together AI. His research focuses on building AI agents for scientific discovery and data science, making AI more reliable and statistically rigorous.
Transcribed and scored by The B2B Podcast Index.
Speaker A: Generative AI is transforming industries at an unprecedented pace. But as AI changes how you or your team work, one thing is clear. Your skills also need to evolve. At, uh, datacamp, we offer everything you
Speaker B: or your team need to adapt and thrive with AI.
Speaker A: Whether it's business users looking to get
Speaker B: the Most out of ChatGPT and Copilot, or developers and data scientists looking to fine tune models.
Speaker A: You can learn the entire AI skills spectrum on Datacamp. Power your AI transformation today. Start learning@datacamp.com we have one project we call the Virtual Biotech. Tens of thousands of AI agents that sort of simulates all the different functions of a pharma company.
Speaker B: We've already seen AI gain superhuman capabilities in analyzing protein structures through projects like AlphaFold. Making AI useful for science in general is still a work in progress. So I'm keen to know where we're up to.
Speaker A: For the past 500 years. The way that humans represent scientific knowledge is in the form of these passive papers which are really like passive artifacts of knowledge. We have this opportunity to basically identify all of these scientific knowledge and within a few days they actually came up with new designs of new proteins that people haven't seen before. And they actually turn out to be better than even previously human expert designed proteins.
Speaker B: Telling us about the frontiers of, um, AI for science is James Zhu. He's a professor at Stanford University and the head of Frontier Agents at Together AI. Uh, James's research spans AI for biomedical applications and, and driving scientific discovery at scale with large swarms of AI agents. His team also built the DS Gym platform for evaluating data science agents. So let's find out where AI is headed for data science and science in general. Hi James, welcome to the show.
Speaker A: Hi Richie. Yeah, thank you for having me. Really excited to participate.
Speaker B: Yeah, great to have you here. Now one of the big trends over the last few years has been having AI agents for, uh, replacing customer service, uh, people and business development people, people and a few other employees. Uh, so how close are ah, we to having AI scientists?
Speaker A: A lot of the existing AI agents are meant more for automating relatively routine and simple workflows. And I think something interesting about science is that good science is never routine. Right. Because the nature is that you want to make new discoveries, so you want to push the frontiers of knowledge, which is what makes science really exciting. And you know, a big part of my work is on building AI scientist agents that can help to really push those frontiers of making discoveries. And that actually I think requires also rethinking how we build and Train the AI agents because a lot of the existing AI agents are more meant to imitate humans. They're taught by essentially try to follow existing workflows. Even the language model is trained by imitating human writings. But to do good science and make new discoveries, you don't want to just imitate. You also want to innovate and try to explore novel ideas. So I think there's still a lot of work we need to do to really teach models how to be more creative, the agents to be like really making novel discoveries. But I think we're making good progress.
Speaker B: Okay, um, yeah, you're right that it's just a very different uh, type of occupation to try and uh, do with AI. So yeah, uh, science necessarily has to be pushing the frontiers. So I'm curious as to where we're up to at the moment. Like what uh, what can AI do towards helping scientists at the moment?
Speaker A: I think AI is making a huge amount of progress in science and I think that's actually going to be one of the most transformative areas of AI, uh, of really automate and to accelerate the way that the speed with which we can make scientific discoveries. I think that's actually going to be one of the most impactful things that AI can do for humanity. And the current kinds of scientific problems that AI are particularly good at in science is more computational problems. Right. So this could be for example mathematical problems or problems that involves uh, sort of like uh, AI research itself or problems, let's say in computational drug discovery or computational protein design, computational biology. So essentially things that can leverage the very strong coding abilities of AI models and also the ability of these models to sort of self refine and self evolve.
Speaker B: Okay, yeah, I mean, I suppose, uh, was it 2024 there was uh, the Nobel Prize awarded for AlphaFold 2, uh, with regards to around protein structures. So there obviously are some advances being made with AI, but I guess uh, in the more general sense it's kind of um, ah, we're not at the point where AI is as good as a human scientist. Is that about right?
Speaker A: Yeah, it's fairly uneven. I would say. There's certain tasks where AI is already very good. So I'll give you one example, uh, which is um, we created what we call the virtual lab, which is a team of AI scientist agents that sort of mirrors a standard human research lab. So there's an AI professor agent, a bunch of different AI student agents, and one of the first tasks that we asked them, agents to basically help us to design new proteins that can bind to the recent SARS COVID variants which can then serve as potential therapeutics or vaccine candidates. And within a few days through a series of group meetings between these agents, they actually came up with uh, new designs of new proteins that people haven't seen before. And we made these experimentally. We actually synthesized these proteins and test them in the wet lab and they actually turn out to be better than even previously human expert designed proteins for binding in terms of binding to these different uh, COVID variants. So that's one example of ah, a computational flavor tasks where AI can already operate at the level of human experts and even better than humans. But there are also a lot of other kind of problems in science that goes um, beyond doing computational modeling that involves actually synthesizing knowledge from different domains and also coming up with experimental evidence to support the discoveries. And that's where human experts are still necessary and needed.
Speaker B: Okay, I mean that is a very cool use case. You mentioned the COVID idea and obviously very impactful. It seems a long time ago now but yeah, of course that was like a world changing event, uh, the COVID pandemic. And uh, having AI contribute to that is like helped solve the problem is pretty amazing. So you're saying computation stuff. I mean I guess that's the easy part. We've had powerful uh, computers for a while and having like slightly smarter uh, approaches. I suppose that's a natural uh, area for AI to be good at. But then come up with like novel hypotheses about what's going on. That's kind of more of a, I guess of more human creativity would you say?
Speaker A: I think that's a good way of putting it. Yeah. And we and other people have seen that uh, for example, if you ask AI to come up with hypothesis or ideas for a new scientific problem, it often ends up having this mode collapse behavior by which we mean maybe the model will come up with one or two good ideas, but if you want to ask it to come up with new and different ideas, it ends up returning back to the first one or two ideas or slight variants of that. So uh, so it doesn't really, it's not able to come up with a lot of very diverse ideas which often is needed to make new scientific progress. Now I think there are ways and techniques that we've been developing that can try to increase the creativity and diversity of these uh, AI agents in generating hypotheses. So for example, one thing that we found to be very useful is to explicitly teach these models to look for analogies because Often a lot of the best ideas in data science and science in general come from taking ideas from adjacent domains, maybe from physics and from telecommunication. And see, oh, maybe uh, a problem in telecommunications is actually very similar to a problem in cell communications in biology. Then taking these analogies and then borrowing ideas to generate new hypotheses. And that's something that we found to actually be quite useful as a way to increase the creativity of AI agents is by teaching it to look for these analogies from very diverse domains.
Speaker B: I suppose a very interesting difference between AI and humans, like these, uh, large language models have seen basically read every book. So being able to understand similarities between domains is something that I guess no human can possibly read, uh, across all these different domains. Okay, if you're a scientist and you want to make more use of AI, then like how do you change your workflows to accommodate this?
Speaker A: You know, one of our visions is that I uh, mentioned this idea of the virtual lab, right? Teams of AI agents. And our vision here is that behind every scientist, behind every human scientist, they should have a virtual lab of AI agents that can help them to do a lot of things, from summarizing literature to generating hypothesis, to designing experiments, analyzing data, even providing critiques and feedback to the human ideas. So I think really across the entire workflow of scientific research, all the way from coming up with research questions to designing experiments, to analyzing data, even writing reports and generating reproducible code, I think all of that can be greatly, uh, accelerated and assisted by AI agents.
Speaker B: That's pretty cool. I love this idea of a virtual lab. Uh, suppose you want one of these. How do you even get started setting one up?
Speaker A: So, good question. So we actually published a paper, uh, so a paper published in Nature a few months ago that introduced the platform of the virtual lab. It's open source, um, and people can just look for virtual lab, uh, or under my name, look for everything. And then they will have actually the open source platforms for building these virtual lab of AI agents and for using these agents across very diverse scientific discovery tasks.
Speaker B: Okay, so this is just like a lab in a box. Just go install the software and then you've got some there. Or do you need to customize it, like what you need to do to
Speaker A: uh, make it work? It's all open source, um, and I think if you for example point cloud code, edit or codecs or a favorite, uh, coding agents, then they should be able to just uh, implement it. Um, quite, uh, should be quite straightforward.
Speaker B: Okay, and you mentioned that you've got um, a Whole team of different agents in there. So I'm thinking about a, uh, real life laboratory where you've got um, like maybe like a biologist or a chemist who is kind of designing the experiments. You've got lab technicians to kind of do the hands on work, and then maybe you've got some data analysts in there to analyze the results. Do you have uh, different types of agents within this virtual lab then?
Speaker A: Yes. Yeah. So it's actually very much, um, a team of different specialist agents. So then there's a professor agent that sort of manages the lab. And then working with the professor agent, we have agents with quite diverse expertise. So there's a data science agent, like a machine learning agent. We could have like a biology agent, or in the case of COVID we had like a protein design agent. And uh, the nice thing is that for different projects, right, like we can give the project description to the professor, the manager agent, and then the manager agent will actually then try to think about for a given project, what are the different experts they would want to have on the team. So maybe for one project it says, oh, it's useful to have like say, uh, a clinician on a team. Right. Then we'll actually go out and train and create a sub agent with expertise in pathology or cardiology. Maybe for a different project, the manager professor agent would say it's useful to have a chemist. Right? They'll actually create a chemist agent. So for different projects, it's actually a lot of flexibility. The PI agent, the manager agent can actually create different teams of experts that are best customized and well suited for that project. And once they are created automatically, then the agents can start to have these group meetings like we do. Right. They will meet together and discuss, come up with research plans. They can also have one on one meetings, right, where one of the agents will meet with the manager agent to review some intermediate subtask so they can start to make progress similar to how we would do.
Speaker B: Okay, uh, that's absolutely fascinating that the idea of like agents stabbing meetings. And so where do, uh, where do humans get involved in this program, in this process? Like do you just leave the agents to it and see what they come up with or do you need to uh, intervene at certain points?
Speaker A: Yeah, it's a good question. So we can provide oversight and we can participate in these agent virtual lab meetings anytime we want. Right. So for example, I could also contribute my ideas in one of these agent virtual lab meetings. And we actually did, um, some tracking to see how often do the humans speak in these virtual Lab meetings and also how often different agents speak. So it turns out that the humans, we don't talk very much and Maybe only about 1% of the time do we actually participate and speak in these virtual lab discussions. I think the role that we tend to see with human researchers is more at the higher level. Like for example, we tell the agents some of the general projects we're interested in. Maybe we say, oh, here's a particular kind of data set we're interested in analyzing. And we also tell the agent some of the constraints we have, like how much time we want to spend on this or how much budget we have to do certain experiments. And then so these are more like high level guidance constraints and feedback to the agents. But otherwise we don't want to micromanage the AI scientists too much. So we want to give them some flexibility to come up with their own plans and to implement and execute those.
Speaker B: Okay, and what's the success rate of these, uh, virtual labs? Like do you come up with, um, do they come up with really good ideas then? Or do you find like, is 1 in 20 is like a good idea or uh, how does it work?
Speaker A: Yeah, so just to give a concrete example, like um, we talked about the STARS Covid, uh, application. So in that case, the virtual lab agents came back to us, right, uh, with A list of 92 new candidate proteins that the agents designed. They said, oh, these are good candidates that you can test as uh, binders to the new COVID variants. So we actually tested all of these 92 candidates, and from these 92, I would say about three or four showed quite promising results and two in particular worked better than previous nanobodies designed by human experts. So you might say, oh, maybe, uh, that means that the stress rate is like maybe 5% or 10%. But in science I think that's actually a very good success rate because usually in drug discovery, where people trying to design these proteins, often they have to design thousands, sometimes millions of candidates and try to find one that works. So if you can actually get one that works out of 10, that's actually really good and can save a huge amount of time and effort and cost.
Speaker B: Yeah, I suppose, uh, whenever you're doing research tasks, the failure rate is incredibly high. Like most science doesn't work, it's novel and you have to spend a lot of time thinking until you get something meaningful.
Speaker A: And I think that's also why it's, I think, somewhat quite different from some of these other applications that you mentioned at the beginning of, let's say customer Support and type, things like that. Because in customer support you have to be very high fidelity. Uh, or in self driving cars or other kind of automations, your agents has to work 99.9% of the time. But in science it's actually the opposite. It's okay if the idea scientists do not work most of the time. In fact for human researchers most of our ideas do not work. That's part of scientific progress. It's great. If just one of ten of our ideas actually works and that's already amazing, can lead to some new breakthroughs, then that's already really good uh, success rate. So we're sort of in this opposite regime of science where it's actually okay to make mistakes for AI and for humans. But we really want to encourage innovation and creativity which is the opposite of things like more standard automation where customer uh, service um, ah, and other applications of AI agents.
Speaker B: Absolutely. I mean you mentioned self driving cars. There a 10% success rate for self driving cars driving that's not very good.
Speaker A: If it crashes, nobody would take those. Yes.
Speaker B: Um, okay, so ah, you mentioned uh, creativity is important and before you were saying how uh, it's changing your approach to like creating these models in the first place. So how do the models need to be different then?
Speaker A: Yeah, so the standard way of training uh, let's say language models, which is sort of the brain behind most of these agents, is that we're training these models essentially to imitate. So if you think about it like all the pre training techniques of next token predictions is basically taking existing corporates of text and then saying can the models imitate and actually reproduce what people have done before? And that's basically the optimization signal, the objective for user training all these models and agents. And that's essentially encouraging incentivizing these models to basically imitate human behaviors. But as we mentioned in science you don't want to just imitate, you want to innovate, you want to come up with new ideas that people haven't thought of before. So one thing that we found to be quite useful there is actually explicitly change the optimization objective to encourage a lot more explorations. So for example instead of having there's kind of imitation learning objectives where you really want the models to reproduce the next tokens and do well, uh, on uh, average we say okay, let's just say it's okay for the models to actually make a lot of mistakes. But as long as let's say one out of the 100 tries that it takes actually leads to a new solution that's much better than before. Then we want to encourage the model and reward it for that. Right. So instead of rewarding it for average performance, which are rewarded for like in some sense like the best performance and that's actually quite a different way of training these models but that's lead to more let's say innovative behaviors.
Speaker B: Okay, so I imagine this is kind of like the equivalent of uh, the crazy professor, you know, with the wild hair and I guess the stereotype. So uh, do you find that um, you get a lot of I guess weirder responses then if you go for creativity, do you also get a lot of nonsense as well alongside that?
Speaker A: Yeah, I think it is a trade off. Um, so we came up with this paradigm, we call it uh, learning to discover where we are explicitly changing the training objective to encourage the models to do this more let's say exploration where risk seeking behaviors that can lead to more new ideas. But as a result of that it also means that uh, maybe it also has more failures. Uh, it can come up by nature a lot of the crazy ideas do not work out. Uh, but I think that goes back to our previous discussion that maybe that's not desirable. When you're talking about self driving cars, you don't want these cars to take crazy routes. But in science it's actually a desirable outcome that you do want AI and humans to try crazy ideas, you know, still within the safe confines. Right. Of science, but try, try new ideas. Right. Um, and it's okay if many of them fail, but one of them works
Speaker B: related to uh, science. Uh, there's obviously uh, uh, agents for data science as well. Uh, do you find that um, is there a similar approach then in terms of having agents for data science? Do you want like a virtual I guess data science lab?
Speaker A: Yeah. So actually with um, a bunch of collaborators and colleagues at together AI, uh we created in some sense like a data science lab. Uh, we call that dsgym. So it's basically a virtual environment, a gym for data science agents to improve data science agents, to train those agents and also to evaluate data science agents. Uh, so in this DS gym we create this fully self contained virtual environment where we have quite diverse kinds of data science tasks. Everything from analyzing data, uh, to derive hypothesis, statistical hypothesis testing, all the way to uh, training predictive models more like Kaggle style kind of data science challenges that we curated to be quite high quality. And we also have created um, the evaluation harness so we can really accurately provide feedback to the data science agents on how well they're doing on these different tasks as well as additional resources like synthetic data pipelines that enables the data science GM DSGM to actually generate a lot of interesting traces, synthetic data that can then be used to train the data science agents. Uh, so that actually creates I think a nice virtual environment for the agents to self uh, improve to become much better at doing these kind of common data science tasks.
Speaker B: That's fascinating. You think uh, humans, I mean, yeah, you spend uh, too much time at the gym trying to figure out uh, getting stronger. Hadn't really thought about agents needing an equivalent system. So uh, in this case um, you said ah, they can self improve. So does uh, that mean like there's no human interaction needed to make the agents better or is this uh, like talk me through how does it work?
Speaker A: Yeah, so I think um, you know in the recent few months there's been a lot of interest in developing agents that can self improve or also called recursive agents or some meta agents or all under the same names of essentially semi idea of like can agents actually with relatively minimal human handholding kind of improve their own capabilities? And how that works typically in general is that you need to have some sort of harness where there's some sort of signal, right? Reward signal or feedback signal comes back automatically, goes uh, back to the agent and then the agent can then use that reward or feedback signal to figure out how to improve either improve their own uh, prompts, instructions, metadata or improve their parameters in the case of the data science uh, in the dsgym. So we basically created that environment where the agents can actually automatically receive these feedback signals from all these different tasks that we designed and also from these leaderboards that we have created. And the agents can actually use those signals either to supervise uh, as a way to basically update their model parameters to more supervised learning or reinforcement learning or they can use it to sort of update their own harnesses which includes like their skills and prompts to approaches like tech grad as a way to um, improve these agents.
Speaker B: Okay, I mean that's very cool that you got this feedback loop and then you're getting better agents out of this with I guess, yeah, minimal human uh, sort of uh, uh interaction. Does this work for all kinds of agents then? Are there ways for any agent to improve in this way?
Speaker A: Yeah, so we tried this on quite a large number of agents. Uh and in particular one of our interests is in can we really create open source data science agents that ah, are very efficient to use, much cheaper to use and also it's more transparent for practitioners. So in this Case we actually showed that we can actually take some of these quite small models. It's like 8 billion, 4 billion parameters. And then by sending those smaller models to the data science gym they get better and stronger. See this self improvement mechanism and they ends up actually performing at the level uh, of Cloud four solids or across many of these data science tasks.
Speaker B: That's very cool. Uh, you mentioned the sort of 8 billion parameter thing. This is like the amount of stuff you can run on a single graphics card, right? So it's kind of available to I guess individual labs or individual researchers.
Speaker A: That's right, yeah. So things that could be run for example on your laptop.
Speaker B: I like the idea of uh, particularly you had me like oh these are going to be cheaper agents to run. I think it's a hot topic at the moment is can you do uh, AI agents cheaply. So okay, so if you've got um, all these agents that are kind of open uh source and uh, they're self improving, talk me through what can you do with these things.
Speaker A: Then now it becomes really interesting. Right Because I think that's what really one of the benefits of the open source community is that different practitioners and researchers and companies can start to train a lot of their own models and agents. They can use the same platform like that we built with dsgen. They can use that same platform actually to customize it to their own use cases to train out their own models. And I think that also opens the door now that there's this larger ecosystem of uh, different agents and different models. This also opens the door for these models to start to cooperate. These more massively multi agent collaborations which is like another topic that we think is super interesting.
Speaker B: Nice stuff. I love that um, you can start having your own custom models as well. I m guess this is part of the beauty of open source is that uh, you can then uh, train things on your own, uh, I guess corporate data or your own personal data and uh, have your own uh, uh version. Okay. Uh, so actually we got slightly sidetracked. I was going to talk about uh, data science agents. So uh, yeah, talk me through like uh, where are we up to with like um, data science agents? Like what's possible, what isn't possible, like what does the frontier look like at the moment?
Speaker A: So with DSGM I think we're able to actually train quite good data science agents. I would say the two main kinds of tasks uh, that we try to optimize the agents to do in DSGM is one is more, let's say exploratory Data analysis kinds of tasks. Right. So given a complex data set. Right. So can you in a more open ended way figure out interesting patterns in those data sets that leads to hypothesis and you can validate those and do rigorous statistics and data science from that. The second kinds of tasks are more predictive modeling. So maybe given a data set, can we, let's say use it to predict the housing prices or predict stock prices or predict different infectious diseases? These are more like the Kago style figure prediction tasks. Uh, so I think in both of those cases the models are now quite good, uh, especially if we provide the agents with the relevant tools and resources. So the tools here, for example, could include things like other kinds of relevant data analysis packages. Uh, for example, if you want the agents to analyze a complex biomedical data set, then it's very useful for the agents to be able to access a lot of the more specialized MCPs and tools that people have developed for those specific domains.
Speaker B: Okay, I mean, that's interesting. So, um, you mentioned you've got exposure data analysis agents and you've got machine learning agents. Is that how specialized you, you want your agent to be to. I know there's a sort of trade off between I've got a single agent that does everything data science or versus I've got a very narrow agent that does like one specific task. Like do you need to go into like, oh, I've got a feature engineering agent or something even more niche. Um, like how, how general specialized should they be?
Speaker A: It's a good question. We haven't seen a huge amount of benefit in having super specialized agents, like agents that focus on maybe just one particular type of data set or particular type of features. Partly because I think data science itself often does benefit from more interdisciplinary knowledge. It has some flavor of statistics and understanding good statistical principles, but also understanding machine learning and also understanding some of the domain knowledge. Right. Uh, and I think that's actually one benefit of having agents that are a little bit broader or at least having a team of agents where they have broader expertise so they can actually start to bring in ideas from different domains. Okay.
Speaker B: All right. So, uh, having them uh, slightly broad, I was thinking about when we talk about like, what are the skills data scientists need? It's uh, always like, uh, you need domain, uh, specific knowledge as well to understand like what's the business or science problem you're trying to solve. And you also need the communication skills and things like that. So do you need other agents for those or are those kind of skills built into the existing data science Agents.
Speaker A: Yeah, I think this is where one area where I think having multi agent collaborations is actually quite natural. Right. Because maybe in those kinds of projects maybe you want to have some agents with relevant domain expertise who can query through the literature or who has a lot of expertise and experience about what are the relevant tools and data sets in that domain. And they really want to pair that agent with other agents that are good at more doing more general purpose like coding or data analysis tasks or training predictive models. So that combination of team of agents then with appropriate coordination and I think often works quite well.
Speaker B: Okay, so back to the idea of like uh, I mean you mentioned the virtual laboratory where it was a whole team of scientists before and now it's like a whole team of uh, uh,
Speaker A: a team of data scientists.
Speaker B: Yeah, yeah. Can you scale this up? Like, can you have like an entire like corporation worth of agents working together on different problems?
Speaker A: Yeah, I think that's definitely the next frontier. I mean we have done some projects. For uh, example we have one project we call the Virtual Biotech that actually has tens of thousands of AI agents, AI scientist agents that sort of simulates uh, all the different functions of a pharma company. So agents that are coming up with potential drug targets, evaluating drug targets, designing therapeutic strategies, designing clinical trials across the whole spectrum. And that involves many different functions, many different expertise. That's why we have potentially so many agents. And I think that's a really exciting next frontier. Seeing can we try to have these fully agent native organizations taking some complex organization, maybe just doing some complex R and D and uh, saying can we really true to have agents, maybe thousands or even millions of agents that sort of simulate and emulate the different cross functional roles of that organization and then bring in the human supervisions and human experts at the relevant places to uh, supervise and mentor the agents.
Speaker B: I mean uh, that sounds pretty amazing. It's very science fiction. Like uh, having this whole uh, grand teams of people working on. Or not people, whole teams of uh, agents working on the uh, big problems. But how do you go about scaling stuff? I think a lot of people like I mean you start with, it's like oh, I wanted to automate something with an agent. And then how do you get from one agent to like you mentioned thousands.
Speaker A: Yeah, and I think that's also where I think some of these um, recursive process where it's not us manually building one agent and another agent. Right. But we essentially have like let's say a meta agent or supervisor agent whose job is to actually create Sub agents by itself. Right. So that actually can lead to a much more scalable setup where the agents are able to spawn their own sub agents to create their own colleagues as needed for specific projects. So then that makes scaling much easier technically. And the other dimension that's also is really important as we're scaling the number of agents into these more complex tasks. I think it's even more important now to have good evaluations, good ways, uh, to really make sure that the agents are doing reasonable analysis, to catch mistakes and to provide these kind of scalable oversight to these large organizations of agents. And this is where having combinations of more deterministic hooks on top of agents, in addition to having good verifiable reward and also have a language model judges with good rubrics, combinations of all of that becomes really important as we're scaling up the complexity of these agentic teams.
Speaker B: Oh man. Yeah, uh, certainly having uh, some way of measuring how good these things are, uh, seems completely essential. You mentioned a few different things there. So you mentioned having like you said, deterministic M hooks. This is like as I presume, like a specific test or measure that um, uh, the agent's any good. And then you mentioned also using LLMs to judge the work of all of their LLMs. Do you want to talk through all these different approaches and when you might want to use each one for a
Speaker A: lot of the scientific discovery or data science discovery tasks that we're looking at, I think our starting point is often that we need to have some way to evaluate the quality of the agent's solutions. And that's often in the form of ideally if we have some verifiable reward, like a deterministic reward that does not involve an LLM, I, um, think that's sort of the best option. It's not always possible, but when that's possible, I think that's often um, maybe the most reliable as a way to prevent reward hacking and um, in other cases in more open ended areas. Right. So then we can also design customized rubrics for maybe having LLM judges and critics, uh, to use those rubrics to evaluate the performance of the submitted solutions. That also often requires having some human supervision. So we also have human domain experts to provide the feedback. So for example, uh, in the case of the COVID design there's uh, some computational evaluations we can do, but ultimately the feedback come from we actually physically make these proteins and then we test them in the real world, uh, and then to see how well do they work and that becomes the Feedback that goes back to the agent.
Speaker B: Okay, so I like the idea that you've got uh, some sort of real world measure of is this good or not? And then providing feedback to the agent is going to tell me or cable, did you do the right thing or not? Okay, uh, all right, so um, suppose you like stolen this dream of having like uh, teams or departments of agents solving problems for you. Uh, how do you go about um, adopting this uh, particularly organization like how do you make sure this happens?
Speaker A: Yeah, I mean I think this is also um, where it's quite different in different use cases. So what we found is that actually scientists are quite, many of them are actually quite open minded and they're quite, actually excited to try and use these agents to help to accelerate discovery. For example, in the recent weeks we're seeing a lot of discussions about AI, uh, in math in particular. And that's also largely driven by the ability now of these frontier models and agents to be able to solve quite complex open problems. And then I think a lot of mathematicians are not natively, uh, the proof is in the pudding. Once they see the results then even if they were skeptical before, now they're actually very excited about the ability of using AI to help to make mathematical discoveries. And we've actually seen that um, firsthand ourselves. So we created this uh, platform called Einstein arena, which is kind of like one of the first agent native platforms for AI scientist agents in the wild, right. To come to collaborate to solve open research problems. And we curated a bunch of these open problems which includes many of these uh, math problems that people are interested in. And our main criteria is that for each of those problems we do have a deterministic verifier to evaluate how good is the solution. So we know that uh, these are really correct. Uh, and we open up this platform so that any agents from anywhere in the world, they can participate. Right. And for free. And just within a few weeks, um, so the agents on Einstein arena, they actually came up and discovered the best new solutions to I think 12 well known problems. Right. So which means that these agents by just by collaborating and interacting with each other on the platform, they came up with better solutions than anything that humans or AI have previously discovered before. Right. Uh, and some of these I think were actually pretty impressive, like breakthroughs in certain areas of uh, the optimization and certain areas of math. And that actually led to a lot of attention and also led to I think quite, uh, a lot of adoption of these kinds of AI agents in those domains.
Speaker B: Yeah, Einstein Rin sounds pretty Amazing. So I love the idea of having a competition platform for agents. Um, are the prizes, I presume, for the human teams behind the agents rather than the, uh, agents themselves. But, uh, can you win things from, um, this, or is it just for glory?
Speaker A: Right now it's just for the glory of discovery. Uh, and what's really kind of interesting here is that we set up the platform to be really agent native. Right. So in some, to participate, you have to prove that you are AI and you are not human in order to enter the arena. And, uh, when there's also it's part of the arena, there's like a platform where the agents can talk to each other, they can ask questions or ask for help from their colleagues from other AI agents. There's also a leaderboard so you can see how the solutions generated by other AI agents as both like a competition but also a collaboration platform. And the other part of this is that we don't know, uh, who are the humans behind this. Uh, because it's all purely designed for the agent interface. And I think that's maybe like a new, potentially a future of what a lot of research will look like, where it will be a lot of somewhat anonymous AI agents making discoveries and generating new results. Um, and I think one of the challenges we're able to figure out is how do we, uh, trace it back to potentially the human teams behind those?
Speaker B: Absolutely. The idea of having to prove that you're AI to join a website kind of boggles my mind a bit. Uh, I think about all those capture things where you've got to select pictures of bicycles or whatever. What's the equivalent test to prove your AI?
Speaker A: Yeah, it's sort of like a reverse capture. Like, basically to interact with the platform, each time, the agent would have to solve, like a numerical puzzle, which will be very easy for AI to do, but it'll be very tedious for humans to do.
Speaker B: All right, that makes sense. Uh, you want to get them to show off the calculation skills. Uh, all right, um, so suppose, uh, you've got all these agents then making discoveries. What, uh, happens to the research then? Uh, do you publish it as a paper and then does the agent get cited? Like, how does that work?
Speaker A: Yeah, I think this is also where we have an opportunity now to reimagine what even papers and more broadly, what scientific knowledge itself should look like. Basically, for the past 500 years, the way that humans represent scientific knowledge is in the form of these passive papers, which are really passive artifacts of knowledge. Um, because if somebody could spend years Doing research. And if you are a reader just by reading the static words on a page, it's often not clear what are the true insights behind that research. Um, but now I think we have this opportunity to basically what I call, to identify all of these scientific knowledge. Uh, we're building this platform called Paper to Agent, which is to basically convert all of these passive artifacts of knowledge from PDFs and papers into dynamic interactive, uh, agent authors. So essentially the paper agent becomes sort of like the virtual corresponding author of that paper. So it knows how to reproduce the results from that paper. It can also know how to apply the data or the methods from that paper to new problems. Uh, so it then becomes this sort uh, of almost like um, in some sense like a living embodiment of the paper. Right. That can help to, to disseminate knowledge and also facilitate new collaborations.
Speaker B: That does sound amazing. I mean it depends a bit on your field, but I think a lot of papers, it's like even if you've read the paper and you're an expert in the field, it can still be quite difficult to reproduce what the original authors did. Uh, and of course they go out of date very quickly. And because it's like from doing the experiments to writing up to getting it published, that can often be m, like more than a year. So I love the idea of being able to speed up that process and having more reproducible science. Okay, so, um, I guess, yeah. What does um, a paper for like, just for agents, uh, look like? Like would you want the output of science to be different for agents versus for humans?
Speaker A: I think that's where it's very interesting. Um, because the current conception of a paper is really something that's designed for human consumption. Consumption. You have this uh, HTML or PDF with the figures, um, and captions and tables. But if we imagine in the near future where the readers of many of these research, the uh, entities that are consuming this research, um, more of those are going to be AI agents rather than humans. And I think there are ways to make the research artifact itself much more agent friendly and more agent native. So as an example of this. So instead of having this static PDF, what we do in Paper to agent is actually create essentially a customized uh, MCP or that captures that research project. So the MCP then itself contains the different tools, the insights and know hows from that research project. And the MCP is actually optimized in a way that enables the agents to be able to reproduce results from that particular research project. So then that mcp, I think Becomes much more of like an agent native artifact of research rather than like a static PDF.
Speaker B: That sounds very cool. I love that idea that, uh, yeah, you've got uh, the MCP interface and then the agent can just pull all the information there. Um, in terms of creating these things, I presume you can have AI create the MCP interface where it's not going to be like humans have to try and figure out programming stuff for the.
Speaker A: Yeah, that's exactly right. Yeah. So like the paper to agent, it's a platform. It's also open source. So the platform that released, uh, essentially automatically will convert research projects and papers into this customized paper mcp.
Speaker B: Okay, so I love this idea of like how science is changing then. So suppose this all works. What does the sort of agentic science future look like in 5, 10 years time?
Speaker A: Well, at first I would want to say that I think, uh, and we're still quite in emerging stages of this agentic science. Um, a lot of progress been made in the last six to 12 months, but we're still in the early stages. And I think one of the big open questions, if we go outside of these domains, these computational domains that we have been discussing as sort of the sweet spot for AI, if we go into these other domains that involves more experimentation or longer horizon things, but it's still not clear. I think we need to do more work on how to really get AI agents to make meaningful progress in those domains. And I think that might be the big, I think the big open challenge for the next five years. How do we get AI to really work? Well, uh, in these domains where we don't have a very clean, verifiable rewards and the projects natively much longer horizon that might take multiple years to conduct.
Speaker B: Yeah, some huge challenges that as you think about, um, anything with social sciences where it involves people and, and things like that, as much more of a challenge. Um, okay. I don't know whether you have any ideas on what might happen here.
Speaker A: Well, I think in the social science, um, particularly interesting, uh, one of the challenges you mentioned is that it's very difficult to do experiments because anything that involves actually doing experiments in real society becomes very costly. And that's also why it's very hard to really get causal signals which often the agents want to have in order to design better algorithms and so on. I think one interesting approach there would be using other AI agents as a way to even simulate human societies. Uh, essentially having maybe I want to understand what's the effect of a particular policy. It's hard to test that policy in the real world. But if I can have a high fidelity simulation of the real world using AI agents that simulate the uh, different human populations, then I can use that much more, much faster, more cost, uh, effective and safer virtual um, world as a way to estimate the causal impact of these different policies.
Speaker B: I mean that does seem incredibly important for anyone involved in lawmaking, uh, governments where if you're able to economically you're able to uh, simulate what's going to happen, make a prediction about, well, what's the impact of all these policies I'm going to create, uh, that seems like a pretty good win for I guess most of humanity there.
Speaker A: I think so, yeah. Because um, for example, there are a lot of uncertainties for any of these policies like what's going to be an impact on gas prices or how it's going to affect people's behaviors if you increase tariffs or if you do this, the world that we have, uh, we only have one timeline. It's not like a multiverse. You can try a bunch of things. But if we actually have uh, a good AI simulator at this virtual society, then that does provide sort of like a multiverse where we can try out a bunch of these things first right in these virtual worlds and see what works well, uh, and then maybe use that to inform how we design the optimal social policies.
Speaker B: Absolutely. I do love that idea of uh, lawmakers being able to think about what they're uh, introducing before they introduce them. Uh, so yeah, certainly the gas price things, uh, it's very topical at the moment. Um, okay, uh, so, uh, I'd like to talk a bit more about uh, the differences in research environments. You're uh, a researcher at Stanford University, but you also work uh, in industry at uh, Together AI. Is there a difference between the approach to uh, AI research like between the two?
Speaker A: I think so. I mean, I think first there's also a lot of commonalities because I think um, one of the really nice things about Together AI is sort of the emphasis on research, an emphasis on publications and open source building open source models, which means that the work that the team together does is actually very much integrated and disseminated among the academic communities and vice versa. They hear a lot of ideas on academic communities and they also have a lot of collaborations with professors and students. So I think that's where it's really exciting. I think one big difference is that um, I think the emphasis on scale and also on efficiency, uh, industry compared to academia. So for example, at Together AI, the teams, there it's really optimizing the inference engine. That's one part of the team. Optimizing inference engine so they can actually serve trillions of tokens. And when you talk about things on that scale, uh, then efficiency becomes super important. Optimizing, uh, the infrastructure, optimizing the kernels, optimizing every components of how to do faster inference. That becomes really important and really very uh, useful. And that's very different from the kinds of considerations often people in academia work on. Because there we're working on tends to be smaller scale problems. And um, maybe a little bit, um, uh, because it's not facing directly customers, then things like efficiency, uh, is often less of a direct consideration.
Speaker B: Okay, yes, they do feel very complementary there. I like the idea that um, if you don't want to worry about customers so much, you can explore more uh, the novel side of things in academia. But industry's generally better at sort of making things scale and do things efficiently.
Speaker A: Yeah, I think it's super complementary. Um, and I mean I've actually think it's actually very nice that um, there are a lot of interesting research problems that arise when we do think about this industry level scale, for example, a lot of things that we talk about. How do we start to manage thousands and millions of AI agents and do that effectively? Uh, that's something that we're starting to run into on the industry side, uh, with partners and customers. And I think that becomes really interesting also on the research side. And similarly these questions we talk about, how do we ensure alignment and safety and guardrails when you have all these agents running around? Right. Um, how do we do scalable oversights? That's something that's really important as we think about adoption and deployment. But that also I think requires a lot of fundamental research to be able to really do that.
Speaker B: Well, absolutely. Uh, yes, you need to keep uh, feeding the progress with new ideas. Uh, I like that. Okay, all right. So uh, to wrap up, I always want more people to learn from. So uh, whose work are you most excited about right now?
Speaker A: Cool. Um, uh, a lot of people when we talk about this kinds of simulation setups, we were talking about simulated societies. I think some of my colleagues at Stanford, like uh, Michael Bernstein, Percy Leon, are doing very interesting work using AI teams to assimilate societies in different settings. I think that's really interesting. More like the AI for science side. Using AI to create these AI scientist agents to make discoveries. Right. Um, yeah, I think there's some even just like in the last week, I think there's some really nice publications from the Google DeepMind teams on these AI co scientist agents. And I think there's also a lot of academic research groups that are building these AI scientist agents.
Speaker B: Tons of exciting work on it. It's difficult to uh, narrow it down, isn't it? But uh, yeah, uh, that's good to know about all your colleagues research and uh, some of the interesting stuff coming out of Google. Okay, uh, nice, uh, thank you so much for your time, James.
Speaker A: Well, thanks for having me. Really enjoyed the conversation.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.