TusoAI: Agentic Optimization for Scientific Methods
Alistair Turcan, Kexin Huang, Lei Li, Martin J. Zhang
Abstract
Scientific discovery is often slowed by the manual development of computational tools needed to analyze complex experimental data. Building such tools is costly and time-consuming because scientists must iteratively review literature, test modeling and scientific assumptions against empirical data, and implement these insights into efficient software. Large language models (LLMs) have demonstrated strong capabilities in synthesizing literature, reasoning with empirical data, and generating domain-specific code, offering new opportunities to accelerate computational method development. Existing LLM-based systems either focus on performing scientific analyses using existing computational methods or on developing computational methods or models for general machine learning without effectively integrating the often unstructured knowledge specific to scientific domains. Here, we introduce TusoAI , an agentic AI system that takes a scientific task description with an evaluation function and autonomously develops and optimizes computational methods for the application. TusoAI integrates domain knowledge into a knowledge tree representation and performs iterative, domain-specific optimization and model diagnosis, improving performance over a pool of candidate solutions. We conducted comprehensive benchmark evaluations demonstrating that TusoAI outperforms state-of-the-art expert methods, MLE agents, and scientific AI agents across diverse tasks, such as single-cell RNA-seq data denoising and satellite-based earth monitoring. Applying TusoAI to two key open problems in genetics improved existing computational methods and uncovered novel biology, including 9 new associations between autoimmune diseases and T cell subtypes and 7 previously unreported links between disease variants linked to their target genes. Our code is publicly available at https://github.com/Alistair-Turcan/TusoAI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3097cce2-0c5d-4c06-aa4c-65ad706965f7Builds on6
- DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based ReasoningSiyuan Guo, Cheng Deng, Ying Wen, Hechang Chen et al.ICML 2024 · 107 citations
- HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care UnitsShenda Hong, Yanbo Xu, Alind Khare, Satria Priambada et al.KDD 2020 · 81 citations
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted RefinementJaehyun Nam, Jinsung Yoon, Jiefeng Chen, Jinwoo Shin et al.NeurIPS 2025 · 58 citations
- Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and FeedbackJiakang Yuan, Xiangchao Yan, Bo Zhang, Tao Chen et al.ACL 2025 · 8 citations
- BoxLM: Unifying Structures and Semantics of Medical Concepts for Diagnosis Prediction in HealthcareYanchao Tan, Hang Lv, Yunfei Zhan, Guofang Ma et al.ICML 2025
Related papers
- LLM Agents Making Agent ToolsGeorg Wölflein, Dyke Ferber, Daniel Truhn, Ognjen Arandjelovic et al.ACL 2025 · 41 citations
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryZiru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang et al.ICLR 2025 · 6 citations
- SR-Scientist: Scientific Equation Discovery With Agentic AIShijie Xia, Yuhan Sun, Pengfei LiuICLR 2026 · 28 citations
- AI-Researcher: Autonomous Scientific InnovationJiabin Tang, Lianghao Xia, Zhonghang Li, Chao HuangNeurIPS 2025 · 101 citations
- AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoMLPatara Trirat, Wonyong Jeong, Sung Ju HwangICML 2025
