BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments
Yusuf H. Roohani, Andrew H. Lee, Qian Huang, Jian Vora, Zachary Steinhart, Kexin Huang, Alexander Marson, Percy Liang, Jure Leskovec
Abstract
Agents based on large language models have shown great potential in accelerating scientific discovery by leveraging their rich background knowledge and reasoning capabilities. In this paper, we introduce BioDiscoveryAgent, an agent that designs new experiments, reasons about their outcomes, and efficiently navigates the hypothesis space to reach desired solutions. We demonstrate our agent on the problem of designing genetic perturbation experiments, where the aim is to find a small subset out of many possible genes that, when perturbed, result in a specific phenotype (e.g., cell growth). Utilizing its biological knowledge, BioDiscoveryAgent can uniquely design new experiments without the need to train a machine learning model or explicitly design an acquisition function as in Bayesian optimization. Moreover, BioDiscoveryAgent using Claude 3.5 Sonnet achieves an average of 21% improvement in predicting relevant genetic perturbations across six datasets, and a 46% improvement in the harder task of non-essential gene perturbation, compared to existing Bayesian optimization baselines specifically trained for this task. Our evaluation includes one dataset that is unpublished, ensuring it is not part of the language model's training data. Additionally, BioDiscoveryAgent predicts gene combinations to perturb more than twice as accurately as a random baseline, a task so far not explored in the context of closed-loop experiment design. The agent also has access to tools for searching the biomedical literature, executing code to analyze biological datasets, and prompting another agent to critically evaluate its predictions. Overall, BioDiscoveryAgent is interpretable at every stage, representing an accessible new paradigm in the computational design of biological experiments with the potential to augment scientists' efficacy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a4d560b-e90d-4ffb-99e0-001e8b2952bdCited by top-tier papers14
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang et al.NeurIPS 2025 · 387 citations
- RAG-Enhanced Collaborative LLM Agents for Drug DiscoveryNamkyeong Lee, Edward De Brouwer, Ehsan Hajiramezanali, Tommaso Biancalani et al.AAAI 2026 · 21 citations
- HeurekaBench: A Benchmarking Framework for AI Co-scientistSiba Smarak Panigrahi, Jovana Videnovic, Maria BrbicICLR 2026 · 10 citations
- scPilot: Large Language Model Reasoning Toward Automated Single-Cell Analysis and DiscoveryYiming Gao, Zhen Wang, Jefferson Chen, Mark Antkowiak et al.NeurIPS 2025 · 8 citations
- ABC-Bench: An Agentic Bio-Capabilities Benchmark for BiosecurityAndrew Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman et al.ICML 2026 · 6 citations
Builds on5
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- GeneDisco: A Benchmark for Experimental Design in Drug DiscoveryArash Mehrjou, Ashkan Soleymani, Andrew Jesson, Pascal Notin et al.ICLR 2022 · 25 citations
- DiscoBAX: Discovery of optimal intervention sets in genomic experiment designClare Lyle, Arash Mehrjou, Pascal Notin, Andrew Jesson et al.ICML 2023 · 16 citations
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
Related papers
- Contextualizing biological perturbation experiments through languageMenghua Wu, Russell Littman, Jacob Levine, Lin Qiu et al.ICLR 2025
- LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific DiscoveryPingchuan Ma, Tsun-Hsuan Wang, Minghao Guo, Zhiqing Sun et al.ICML 2024 · 76 citations
- BioBO: Biology-informed Bayesian Optimization for Perturbation DesignYanke Li, Tianyu Cui, Tommaso Mansi, Mangal Prakash et al.ICLR 2026 · 2 citations
- Protein Design with Agent Rosetta: A Case Study for Specialized Scientific AgentsJacopo Teneggi, SM Turzo, Tanya Marwah, Alberto Bietti et al.ICML 2026
- TusoAI: Agentic Optimization for Scientific MethodsAlistair Turcan, Kexin Huang, Lei Li, Martin J. ZhangICLR 2026 · 3 citations
