SciNav: A General Agent Framework for Scientific Coding Tasks
Tianshu Zhang, Huan Sun
Abstract
Autonomous science agents, built on large language models (LLMs), are increasingly being investigated to generate hypotheses, design experiments, and produce reports. Prior science agents primarily focus on open-ended scientific problems, where such outputs-hypotheses, experiments, or analyses are inherently subjective and thus difficult to evaluate rigorously. In contrast, existing scientific coding benchmarks provide tasks with clearly defined, executable outputs that enable objective assessment. However, current agent-based approaches to these benchmarks remain engineering-driven pipelines, lacking structured framework design. This mismatch exposes a gap: the absence of end-to-end, structured science agent frameworks for scientific coding tasks. We address this gap by focusing on scientific coding tasks, where evaluation can be made rigorously, and introducing an agent framework SciNav (Scientific Navigator) that enables more effective solution exploration. Our framework is designed to operate under constrained search budgets, moving beyond reliance on pre-defined success metrics and prolonged search cycles. Inspired by findings that comparative judgments often reveal finer-grained quality differences and therefore provide greater discriminative power than absolute scoring, our framework leverages pairwise relative judgments within a tree search process to select top-K promising solution branches, prune low-potential ones, and progressively narrow down the solution candidates on the selected branches guided by relative comparisons. We demonstrate our agent's effectiveness across different types of tasks on two benchmarks. Experiments show that SciNav significantly outperforms direct prompting and prior agents like OpenHands and Self-Debug across different base models, task types, and difficulty levels, and exceeds different frontier comparators such as random selection and LLM absolute scoring. These results confirm the strength of our agent design and highlight the effectiveness of relative judgment-guided top-K search for high-quality scientific coding, marking a step toward more practical science agents. 1 INTRODUCTION Large language models (LLMs) have recently shown strong potential to advance scientific discovery, giving rise to science agents that aim to automate the research process end-to-end. Science agents such as Agent Laboratory (Schmidgall et al., 2025 ), ResearchAgent (Baek et al., 2024), and AlphaEvolve (Cui et al., 2021), aspire to generate research ideas, design experiments, and draft reports. While this vision of broad-scope automation is compelling, it raises a central challenge: how to evaluate the outputs of such agents. Unlike standard benchmarks with unambiguous correctness criteria, the artifacts produced, including novel hypotheses, experimental protocols, and written analyses, are inherently open-ended and subjective, often demanding expert review or costly human studies to judge their scientific validity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- AlphaEvolve: A Learning Framework to Discover Novel Alphas in Quantitative InvestmentCan Cui, Wei Wang, Meihui Zhang, Gang Chen et al.SIGMOD 2021 · 25 citations
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang et al.EMNLP 2024 · 7 citations
- OpenHands: An Open Platform for AI Software Developers as Generalist AgentsXingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu et al.ICLR 2025 · 7 citations
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryZiru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang et al.ICLR 2025 · 6 citations
Related papers
- Agent-as-a-Judge: Evaluate Agents with AgentsMingchen Zhuge, Changsheng Zhao, Dylan R. Ashley, Wenyi Wang et al.ICML 2025
- LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language ModelsParshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani et al.ICML 2025
- An Agent-based Evaluation Framework for Complex Code GenerationXinchen Wang, Ruida Hu, Pengfei Gao, Chao Peng et al.ASE 2025 · 7 citations
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific WorkflowsQiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding et al.ICLR 2026 · 45 citations
- FeatureBench: Benchmarking Agentic Coding for Complex Feature DevelopmentQixing Zhou, Jiacheng Zhang, Haiyang Wang, Rui Hao et al.ICLR 2026 · 30 citations
