ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
Jiangyuan Wang, Kejun Xiao, Qi Sun, Huaipeng Zhao, Tao Luo, Jiandong Zhang, Xiaoyi Zeng
摘要
Existing benchmarks in e-commerce primarily focus on basic user intents, such as finding or purchasing products. However, real-world users often pursue more complex goals, such as applying vouchers, managing budgets, and finding multiproducts seller. To bridge this gap, we propose Shopping-Bench, a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products. Experimental results demonstrate that even state-of-the-art language agents (such as GPT-4.1) achieve absolute success rates under 50% on our benchmark tasks, highlighting the significant challenges posed by our ShoppingBench. In addition, we propose a trajectory distillation strategy and leverage supervised fine-tuning, along with reinforcement learning on synthetic trajectories, to distill the capabilities of a large language agent into a smaller one. As a result, our trained agent achieves competitive performance compared to GPT-4.1 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic EncodersYupeng Hou, Jiacheng Li, Xiangjun Fu, Zhankui He 等ACL 2026 · 被引用 346 次
- AgenticShop: Benchmarking Agentic Product Curation for Personalized Web ShoppingSunghwan Kim, Ryang Heo, Yongsik Seo, Jinyoung Yeo 等WWW 2026 · 被引用 3 次
- JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional TasksLanbo Lin, Jiayao Liu, Tianyuan Yang, Li Cai 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper6
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang 等NeurIPS 2025 · 被引用 387 次
相关 Paper
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsShunyu Yao, Noah Shinn, Pedram Razavi, Karthik R. NarasimhanICLR 2025
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag 等ACL 2025 · 被引用 15 次
- ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping AssistantsPei Wang, Yanan Wu, Xiaoshuai Song, Weixun Wang 等ACL 2026 · 被引用 5 次
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world ApplicationsWei He, Yueqing Sun, Hongyan Hao, Xueyuan Hao 等ICLR 2026 · 被引用 41 次
- Beyond Itinerary Planning - A Real-World Benchmark for Multi-Turn and Tool-Using Travel TasksXiang Cheng, Yulan Hu, Xiangwen Zhang, Lu Xu 等ACL 2026 · 被引用 4 次
