Lune

ICLR2026顶会

SciNav: A General Agent Framework for Scientific Coding Tasks

Tianshu Zhang, Huan Sun

2026年份
1被引次数

摘要

Autonomous science agents, built on large language models (LLMs), are increasingly being investigated to generate hypotheses, design experiments, and produce reports. Prior science agents primarily focus on open-ended scientific problems, where such outputs-hypotheses, experiments, or analyses are inherently subjective and thus difficult to evaluate rigorously. In contrast, existing scientific coding benchmarks provide tasks with clearly defined, executable outputs that enable objective assessment. However, current agent-based approaches to these benchmarks remain engineering-driven pipelines, lacking structured framework design. This mismatch exposes a gap: the absence of end-to-end, structured science agent frameworks for scientific coding tasks. We address this gap by focusing on scientific coding tasks, where evaluation can be made rigorously, and introducing an agent framework SciNav (Scientific Navigator) that enables more effective solution exploration. Our framework is designed to operate under constrained search budgets, moving beyond reliance on pre-defined success metrics and prolonged search cycles. Inspired by findings that comparative judgments often reveal finer-grained quality differences and therefore provide greater discriminative power than absolute scoring, our framework leverages pairwise relative judgments within a tree search process to select top-K promising solution branches, prune low-potential ones, and progressively narrow down the solution candidates on the selected branches guided by relative comparisons. We demonstrate our agent's effectiveness across different types of tasks on two benchmarks. Experiments show that SciNav significantly outperforms direct prompting and prior agents like OpenHands and Self-Debug across different base models, task types, and difficulty levels, and exceeds different frontier comparators such as random selection and LLM absolute scoring. These results confirm the strength of our agent design and highlight the effectiveness of relative judgment-guided top-K search for high-quality scientific coding, marking a step toward more practical science agents. 1 INTRODUCTION Large language models (LLMs) have recently shown strong potential to advance scientific discovery, giving rise to science agents that aim to automate the research process end-to-end. Science agents such as Agent Laboratory (Schmidgall et al., 2025 ), ResearchAgent (Baek et al., 2024), and AlphaEvolve (Cui et al., 2021), aspire to generate research ideas, design experiments, and draft reports. While this vision of broad-scope automation is compelling, it raises a central challenge: how to evaluate the outputs of such agents. Unlike standard benchmarks with unambiguous correctness criteria, the artifacts produced, including novel hypotheses, experimental protocols, and written analyses, are inherently open-ended and subjective, often demanding expert review or costly human studies to judge their scientific validity.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖