EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association
Weiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag, Wenju Xu, Chen Luo, Sheikh Muhammad Sarwar, Yang Li, Hansu Gu, Hui Liu, Changlong Yu, Jiaxin Bai
Abstract
Goal-oriented script planning, or the ability to plan coherent sequences of actions toward specific goals, is commonly used by humans to plan for daily activities. In e-commerce, customers increasingly seek LLM-based assistants to plan for them with a script and recommend products at each step, thereby facilitating convenient and efficient shopping experiences. However, this capability remains underexplored due to several challenges, including the inability of LLMs to simultaneously conduct script planning and product retrieval, difficulties in matching products caused by semantic discrepancies between planned actions and search queries, and a lack of methods and benchmark data for evaluation. In this paper, we step forward by formally defining the task of E-commerce Script Planning (ECOMSCRIPT) as three sequential subtasks. We propose a novel framework that enables the scalable generation of product-enriched scripts by associating products with each step based on the semantic similarity between the actions and their purchase intentions. By applying our framework to real-world e-commerce data, we construct the very first large-scale ECOMSCRIPT dataset, ECOMSCRIPTBENCH, which includes 605,229 scripts sourced from 2.4 million products. Human annotations are then conducted to provide gold labels for a sampled subset, forming an evaluation benchmark. Extensive experiments reveal that current (L)LMs face significant challenges with ECOMSCRIPT tasks, even after fine-tuning, while injecting product purchase intentions improves their performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4c2dda0-3f83-418a-9277-3c640b80cc7fCited by top-tier papers2
- ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based AgentsJiangyuan Wang, Kejun Xiao, Qi Sun, Huaipeng Zhao et al.AAAI 2026 · 14 citations
- Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain RulesChenyu Zhou, Xiaoming Shi, Hui Qiu, Yankai Jiang et al.ICLR 2026 · 4 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Distilling Script Knowledge from Large Language Models for Constrained Language PlanningSiyu Yuan, Jiangjie Chen, Ziquan Fu, Xuyang Ge et al.ACL 2023 · 14 citations
- ABSEval: An Agent-based Framework for Script EvaluationSirui Liang, Baoli Zhang, Jun Zhao, Kang LiuEMNLP 2024 · 2 citations
- eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction DataBo Peng, Xinyi Ling, Ziru Chen, Huan Sun et al.ICML 2024 · 53 citations
- Wizard of Shopping: Target-Oriented E-commerce Dialogue Generation with Decision Tree BranchingXiangci Li, Zhiyu Chen, Jason Ingyu Choi, Nikhita Vedula et al.ACL 2025
- ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping AssistantsPei Wang, Yanan Wu, Xiaoshuai Song, Weixun Wang et al.ACL 2026 · 5 citations
