WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
Zexuan Wang, Chenghao Yang, Yingqi Que, Zhoufutu Wen, Zaiyuan Wang, Jiashuo Liu, Zhixin Yao, Zhenzhu Yang, Huaqing Yuan, Yiwen Wang, Zhengxuan Jiang, Shengjie Fang
摘要
Real-world autonomous planning requires coordinating tightly coupled constraints where a single decision dictates the feasibility of all subsequent actions. However, existing benchmarks predominantly feature loosely coupled constraints solvable through local greedy decisions and rely on idealized data, failing to capture constraint acquisition from realistic web interfaces. We introduce , a benchmark comprising 150 real-world travel scenarios across 5 cities, requiring agents to satisfy an average of 15+ interdependent temporal and logical constraints. To evaluate realistic deployment settings, we further develop , a multi-modal environment with over 2,000 rendered webpages that preserve layout-dependent and information-dense travel interfaces, requiring agents to recover executable constraints from rendered web interfaces. Evaluating 10 frontier models reveals a severe performance collapse: GPT-5.2 achieves only 28.0% feasibility in text-only settings, dropping to 3.4% in multi-modal environments. We observe substantial degradation in planning feasibility when agents must recover executable constraints from rendered webpages, alongside a Planning Horizon threshold at approximately 10 constraints where reasoning reliability collapses. These findings suggest that realistic constraint acquisition and long-horizon planning remain complementary bottlenecks for current agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
相关 Paper
- Beyond Itinerary Planning - A Real-World Benchmark for Multi-Turn and Tool-Using Travel TasksXiang Cheng, Yulan Hu, Xiangwen Zhang, Lu Xu 等ACL 2026 · 被引用 4 次
- TravelPlanner: A Benchmark for Real-World Planning with Language AgentsJian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu 等ICML 2024 · 被引用 376 次
- DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable ConstraintsYinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu 等ACL 2026 · 被引用 21 次
- LongCoT: Benchmarking Long-Horizon Chain-of-Thought ReasoningSumeet Motwani, Daniel Nichols, Charles London, Peggy Li 等ICML 2026 · 被引用 2 次
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use AgentsJiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong 等ACL 2026 · 被引用 19 次
