Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimization
Xuefei Wang, Kai A. Horstmann, Ethan Lin, Jonathan Chen, Alexander Farhang, Sophia Stiles, Atharva Sehgal, Jonathan Light, David Valen, Yisong Yue, Jennifer J. Sun
Abstract
Adapting production-level computer vision tools to bespoke scientific datasets is a critical "last mile" bottleneck. Current solutions are impractical: fine-tuning requires large annotated datasets scientists often lack, while manual code adaptation costs scientists weeks to months of effort. We consider using AI agents to automate this manual coding, and focus on the open question of optimal agent design for this targeted task. We introduce a systematic evaluation framework for agentic code optimization and use it to study three production-level biomedical imaging pipelines. We demonstrate that a simple agent framework consistently generates adaptation code that outperforms human-expert solutions. Our analysis reveals that common, complex agent architectures are not universally beneficial, leading to a practical roadmap for agent design. We open source our framework and validate our approach by deploying agentgenerated functions into a production pipeline, demonstrating a clear pathway for real-world impact. The code can be found here: https://github.com/xuefei-wang/simple-agent- opt
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- MLAgentBench: Evaluating Language Agents on Machine Learning ExperimentationQian Huang, Jian Vora, Percy Liang, Jure LeskovecICML 2024 · 209 citations
- Learning Performance-Improving Code EditsAlexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon et al.ICLR 2024 · 141 citations
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-benchEdan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra et al.NeurIPS 2025 · 71 citations
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringJun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung et al.ICLR 2025 · 9 citations
Related papers
- ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT PipelinesTengjun Jin, Yuxuan Zhu, Daniel KangVLDB 2026 · 13 citations
- BioAgent Bench: An AI Agent Evaluation Suite for BioinformaticsDionizije Fa, Marko Culjak, Bruno Pandza, Mateo CupicICML 2026 · 7 citations
- AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoMLPatara Trirat, Wonyong Jeong, Sung Ju HwangICML 2025
- CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding ChallengesKechi Zhang, Jia Li, Ge Li, Xianjie Shi et al.ACL 2024
- AdaptAgent: A Multi-agent, Domain-Guided Reasoning Framework for Code AdaptationXiaokai Rong, Hridya Dhulipala, Aashish Yadavally, Tien N. NguyenISSTA 2026
