TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents
Yanyu Chen, Jiyue Jiang, Jiahong Liu, Yifei Zhang, Xiao Guo, Irwin King
摘要
The evaluation of Deep Research Agents is a critical challenge, as conventional outcome-based metrics fail to capture the nuances of their complex reasoning. Current evaluation faces two primary challenges: 1) a reliance on singular metrics like Pass@1, creating a ''high-score illusion'' that ignores the quality, efficiency, and soundness of the reasoning process; and 2) the failure of static benchmarks to quantify crucial attributes like robustness and latent capability. To address these gaps, we introduce TRACE (Trajectory-Aware Comprehensive Evaluation), a framework that holistically assesses the entire problem-solving trajectory. To counter the ''high-score illusion'', we propose a Hierarchical Trajectory Utility Function that quantifies process efficiency and cognitive quality, including evidence grounding, alongside accuracy. To measure deeper attributes, TRACE introduces a Scaffolded Capability Assessment protocol, quantifying an agent's latent ability by determining the minimum guidance needed for success. Our contributions include the TRACE framework, its novel metrics, and the accompanying DeepResearch-Bench with controllable complexity. Experiments show TRACE delivers a granular ranking that uncovers critical trade-offs between agent accuracy, efficiency, and robustness entirely missed by singular metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang 等ICLR 2026 · 被引用 406 次
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang 等NeurIPS 2025 · 被引用 387 次
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement LearningKuan Li, Zhongwang Zhang, Huifeng Yin, Rui Ye 等ICLR 2026 · 被引用 65 次
相关 Paper
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented AgentsWonjoong Kim, Sangwu Park, Yeonjun In, Sein Kim 等ICML 2026 · 被引用 15 次
- DREAM: Deep Research Evaluation with Agentic MetricsElad Ben-Avraham, Changhao Li, Ron Dorfman, Roy Ganz 等ACL 2026 · 被引用 2 次
- Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time ExplorationDadi Guo, Tianyi Zhou, Dongrui Liu, Chen Qian 等ICLR 2026 · 被引用 3 次
- A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-Controlled and Dynamic Test GenerationShide Zhou, Kailong Wang, Ling Shi, Haoyu WangISSTA 2026
- TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool UsePengfei He, Zhenwei Dai, Bing He, Hui Liu 等ICLR 2026 · 被引用 46 次
