Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
Wonjoong Kim, Sangwu Park, Yeonjun In, Sein Kim, Dongha Lee, Chanyoung Park
Abstract
Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory aspects like efficiency, hallucination, and adaptivity. The most straightforward method for evaluation is to compare an agent’s trajectory with the ground-truth, but annotating all valid ground-truth trajectories is prohibitively expensive. In this manner, we introduce TRACE, a reference free framework for the multi-dimensional evaluation of tool-augmented LLMs. By incorporating an evidence bank which accumulates knowledge from preceding steps, TRACE assesses an agent’s reasoning trajectory effectively. To validate our framework, we develop a new meta-evaluation dataset with diverse and flawed trajectories, each labeled with multi-faceted performance scores. Our results confirm that TRACE accurately evaluates complex trajectories even with small open source LLMs. Furthermore, we apply our method to evaluate the trajectories that agents produce while solving tool-augmented tasks, presenting previously unreported observations and their corresponding insights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f59d0fb-2e7a-4ec6-a67f-5106112fa098Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem SolvingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 289 citations
- Deductive Verification of Chain-of-Thought ReasoningZhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang et al.NeurIPS 2023 · 234 citations
Related papers
- SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QASher Badshah, Ali Emami, Hassan SajjadACL 2026 · 1 citation
- TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool UsePengfei He, Zhenwei Dai, Bing He, Hui Liu et al.ICLR 2026 · 46 citations
- PROBE: Dense Process Rewards with Observation Evidence for Tool-Augmented Visual ReasoningZongsheng Cao, Anran Liu, Jun Xie, Feng Chen et al.KDD 2026
- TaskCraft: Automated Generation of Agentic TasksDingfeng Shi, Jingyi Cao, Qianben Chen, Weichen Sun et al.ICLR 2026 · 49 citations
- Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time ExplorationDadi Guo, Tianyi Zhou, Dongrui Liu, Chen Qian et al.ICLR 2026 · 3 citations
