CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
Jiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung
Abstract
Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability. This neglects a crucial capability: agents'ability to devise and adjust cost-optimal plans in response to changing environments. To bridge this gap, we introduce CostBench, a scalable, cost-centric benchmark designed to evaluate agents'economic reasoning and replanning abilities. Situated in the travel-planning domain, CostBench comprises tasks solvable via multiple sequences of atomic and composite tools with diverse, customizable costs. It also supports four types of dynamic blocking events, such as tool failures and cost changes, to simulate real-world unpredictability and necessitate agents to adapt in real time. Evaluating leading open-sourced and proprietary models on CostBench reveals a substantial gap in cost-aware planning: agents frequently fail to identify cost-optimal solutions in static settings, with even GPT-5 achieving less than 75% exact match rate on the hardest tasks, and performance further dropping by around 40% under dynamic conditions. By diagnosing these weaknesses, CostBench lays the groundwork for developing future agents that are both economically rational and robust.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning ModelsDadi Guo, Jiayu Liu, Zhiyuan Fan, Zhitao He et al.ACL 2026 · 17 citations
- Diversity-Enhanced Reasoning for Subjective QuestionsYumeng Wang, Zhiyuan Fan, Jiayu Liu, Jen-Tse Huang et al.ICLR 2026 · 13 citations
- Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated ReasoningQisheng Su, Shiting Huang, Zhen Fang, Ziyan Chen et al.ACL 2026 · 3 citations
- Towards Self-Evolving Agent Benchmarks : Validatable Agent Trajectory via Test-Time ExplorationDadi Guo, Tianyi Zhou, Dongrui Liu, Chen Qian et al.ICLR 2026 · 3 citations
Builds on16
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- TravelPlanner: A Benchmark for Real-World Planning with Language AgentsJian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu et al.ICML 2024 · 376 citations
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen et al.ICLR 2024 · 308 citations
Related papers
- Beyond Itinerary Planning - A Real-World Benchmark for Multi-Turn and Tool-Using Travel TasksXiang Cheng, Yulan Hu, Xiangwen Zhang, Lu Xu et al.ACL 2026 · 4 citations
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 35 citations
- DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable ConstraintsYinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu et al.ACL 2026 · 21 citations
- Market-Bench: Benchmarking Large Language Models on Economic and Trade CompetitionYushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu et al.ACL 2026 · 1 citation
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
