SciNav: A General Agent Framework for Scientific Coding Tasks
Tianshu Zhang, Huan Sun
摘要
Autonomous science agents, built on large language models (LLMs), are increasingly being investigated to generate hypotheses, design experiments, and produce reports. Prior science agents primarily focus on open-ended scientific problems, where such outputs-hypotheses, experiments, or analyses are inherently subjective and thus difficult to evaluate rigorously. In contrast, existing scientific coding benchmarks provide tasks with clearly defined, executable outputs that enable objective assessment. However, current agent-based approaches to these benchmarks remain engineering-driven pipelines, lacking structured framework design. This mismatch exposes a gap: the absence of end-to-end, structured science agent frameworks for scientific coding tasks. We address this gap by focusing on scientific coding tasks, where evaluation can be made rigorously, and introducing an agent framework SciNav (Scientific Navigator) that enables more effective solution exploration. Our framework is designed to operate under constrained search budgets, moving beyond reliance on pre-defined success metrics and prolonged search cycles. Inspired by findings that comparative judgments often reveal finer-grained quality differences and therefore provide greater discriminative power than absolute scoring, our framework leverages pairwise relative judgments within a tree search process to select top-K promising solution branches, prune low-potential ones, and progressively narrow down the solution candidates on the selected branches guided by relative comparisons. We demonstrate our agent's effectiveness across different types of tasks on two benchmarks. Experiments show that SciNav significantly outperforms direct prompting and prior agents like OpenHands and Self-Debug across different base models, task types, and difficulty levels, and exceeds different frontier comparators such as random selection and LLM absolute scoring. These results confirm the strength of our agent design and highlight the effectiveness of relative judgment-guided top-K search for high-quality scientific coding, marking a step toward more practical science agents. 1 INTRODUCTION Large language models (LLMs) have recently shown strong potential to advance scientific discovery, giving rise to science agents that aim to automate the research process end-to-end. Science agents such as Agent Laboratory (Schmidgall et al., 2025 ), ResearchAgent (Baek et al., 2024), and AlphaEvolve (Cui et al., 2021), aspire to generate research ideas, design experiments, and draft reports. While this vision of broad-scope automation is compelling, it raises a central challenge: how to evaluate the outputs of such agents. Unlike standard benchmarks with unambiguous correctness criteria, the artifacts produced, including novel hypotheses, experimental protocols, and written analyses, are inherently open-ended and subjective, often demanding expert review or costly human studies to judge their scientific validity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- AlphaEvolve: A Learning Framework to Discover Novel Alphas in Quantitative InvestmentCan Cui, Wei Wang, Meihui Zhang, Gang Chen 等SIGMOD 2021 · 被引用 25 次
- DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang 等EMNLP 2024 · 被引用 7 次
- OpenHands: An Open Platform for AI Software Developers as Generalist AgentsXingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu 等ICLR 2025 · 被引用 7 次
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific DiscoveryZiru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang 等ICLR 2025 · 被引用 6 次
相关 Paper
- Agent-as-a-Judge: Evaluate Agents with AgentsMingchen Zhuge, Changsheng Zhao, Dylan R. Ashley, Wenyi Wang 等ICML 2025
- LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language ModelsParshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani 等ICML 2025
- An Agent-based Evaluation Framework for Complex Code GenerationXinchen Wang, Ruida Hu, Pengfei Gao, Chao Peng 等ASE 2025 · 被引用 7 次
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific WorkflowsQiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding 等ICLR 2026 · 被引用 45 次
- FeatureBench: Benchmarking Agentic Coding for Complex Feature DevelopmentQixing Zhou, Jiacheng Zhang, Haiyang Wang, Rui Hao 等ICLR 2026 · 被引用 30 次
