Protein Design with Agent Rosetta: A Case Study for Specialized Scientific Agents
Jacopo Teneggi, SM Turzo, Tanya Marwah, Alberto Bietti, P. Douglas Renfrew, Vikram Mulligan, Siavash Golkar
摘要
Large language models (LLMs) are capable of emulating reasoning and using tools, creating opportunities for autonomous agents that execute complex scientific tasks. Protein design provides a natural testbed: although machine learning (ML) methods achieve strong results, these are largely restricted to canonical amino acids and narrow objectives, leaving unfilled need for a generalist tool for broad design pipelines. We introduce Agent Rosetta, an LLM agent paired with a structured environment for operating Rosetta, the leading physics-based heteropolymer design software, capable of modeling non-canonical building blocks and geometries. Agent Rosetta iteratively refines designs to achieve user-defined objectives, combining LLM reasoning with Rosetta's generality. We evaluate Agent Rosetta on design with canonical amino acids, matching specialized models and expert baselines, and with non-canonical residues---where ML approaches fail---achieving comparable performance. Critically, prompt engineering alone often fails to generate Rosetta actions, demonstrating that environment design is essential for integrating LLM agents with specialized software. Our results show that properly designed environments enable LLM agents to make scientific software accessible while matching specialized tools and human experts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
- DiffDock: Diffusion Steps, Twists, and Turns for Molecular DockingGabriele Corso, Hannes Stärk, Bowen Jing, Regina Barzilay 等ICLR 2023 · 被引用 331 次
相关 Paper
- PDAgent: An LLM-Driven Autonomous Agent Framework Towards In Silico Protein Design via Directed MutationSong Ouyang, Zhijie Dong, Yong Luo, Kehua Su 等ICML 2026
- Gödel Agent: A Self-Referential Agent Framework for Recursively Self-ImprovementXunjian Yin, Xinyi Wang, Liangming Pan, Li Lin 等ACL 2025 · 被引用 7 次
- Proteo-R1: Reasoning Foundation Models for De Novo Protein DesignFang Wu, Weihao Xuan, Heli Qi, Hanqun CAO 等ICML 2026 · 被引用 5 次
- BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation ExperimentsYusuf H. Roohani, Andrew H. Lee, Qian Huang, Jian Vora 等ICLR 2025
- Retro-R1: LLM-based Agentic RetrosynthesisWei Liu, Jiangtao Feng, Hongli Yu, Yuxuan Song 等NeurIPS 2025 · 被引用 8 次
