AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping
Sunghwan Kim, Ryang Heo, Yongsik Seo, Jinyoung Yeo, Dongha Lee
摘要
The proliferation of e-commerce has made web shopping platforms key gateways for customers navigating the vast digital marketplace. Yet this rapid expansion has led to a noisy and fragmented information environment, increasing cognitive burden as shoppers explore and purchase products online. With promising potential to alleviate this challenge, agentic systems have garnered growing attention for automating user-side tasks in web shopping. Despite significant advancements, existing benchmarks fail to comprehensively evaluate how well agentic systems can curate products in open-web settings. Specifically, they have limited coverage of shopping scenarios, focusing only on simplified single-platform lookups rather than exploratory search. Moreover, they overlook personalization in evaluation, leaving unclear whether agents can adapt to diverse user preferences in realistic shopping contexts. To address this gap, we present AgenticShop, the first benchmark for evaluating agentic systems on personalized product curation in open-web environment. Crucially, our approach features realistic shopping scenarios, diverse user profiles, and a verifiable, checklistdriven personalization evaluation framework. Through extensive experiments, we demonstrate that current agentic systems remain largely insufficient, emphasizing the need for user-side systems that effectively curate tailored products across the modern web. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine OptimizationSunghwan Kim, Wooseok Jeong, Serin Kim, Sangam Lee 等KDD 2026 · 被引用 8 次
- Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User HistorySerin Kim, Sangam Lee, Dongha LeeICML 2026
它引用的顶会 Paper14
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- Multimodal Web Navigation with Instruction-Finetuned Foundation ModelsHiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo 等ICLR 2024 · 被引用 160 次
相关 Paper
- A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce DomainsXianren Zhang, Shreyas Prasad, Di Wang, Qiuhai Zeng 等ACL 2026 · 被引用 8 次
- AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic WebShanshan Zhong, Kate Shen, Chenyan XiongICML 2026 · 被引用 2 次
- Towards Personalized Deep Research: Benchmarks and EvaluationsYuan Liang, Jiaxian Li, Yuqing Wang, Piaohong Wang 等ICLR 2026 · 被引用 13 次
- GTA: Generating Long-horizon Tasks for Web Agents at ScaleTenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou 等ACL 2026
- Ego2Web: A Web Agent Benchmark Grounded in Egocentric VideosShoubin Yu, Lei Shu, Antoine Yang, Yao Fu 等CVPR 2026 · 被引用 4 次
