Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules
Chenyu Zhou, Xiaoming Shi, Hui Qiu, Yankai Jiang, ShaoGuo Liu, Tingting Gao, Haitao Leng, Xiawu Zheng, Rongrong Ji
Abstract
E-commerce agents contribute greatly to helping users complete their e-commerce needs. To promote further research and application of e-commerce agents, benchmarking frameworks are introduced for evaluating LLM agents in the e-commerce domain. Despite the progress, current benchmarks lack evaluating agents' capability to handle mixed-type e-commerce dialogue and complex domain rules. To address the issue, this work first introduces a novel corpus, termed Mix-ECom, which is constructed based on real-world customer-service dialogues with post-processing to remove user privacy and add CoT process. Specifically, Mix-ECom contains 4,799 samples with multiply dialogue types in each e-commerce dialogue, covering four dialogue types (QA, recommendation, task-oriented dialogue, and chit-chat), three e-commerce task types (pre-sales, logistics, after-sales), and 82 e-commerce rules. Furthermore, this work build baselines on Mix-Ecom and propose a dynamic framework to further improve the performance. Results show that current e-commerce agents lack sufficient capabilities to handle e-commerce dialogues, due to the hallucination cased by complex domain rules. The dataset will be publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6bae5d7-f3e4-419d-8cd0-b07827d3af59Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language ModelsLei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu et al.ACL 2023 · 249 citations
- A Multimodal Foundation Agent for Financial Trading: Tool-Augmented, Diversified, and GeneralistWentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun et al.KDD 2024 · 50 citations
- WebCPM: Interactive Web Search for Chinese Long-form Question AnsweringYujia Qin, Zihan Cai, Dian Jin, Lan Yan et al.ACL 2023 · 25 citations
Related papers
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag et al.ACL 2025 · 15 citations
- EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerceYangning Li, Shirong Ma, Xiaobin Wang, Shen Huang et al.AAAI 2024 · 85 citations
- LegalAgentBench: Evaluating LLM Agents in Legal DomainHaitao Li, Junjie Chen, Jingli Yang, Qingyao Ai et al.ACL 2025
- eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction DataBo Peng, Xinyi Ling, Ziru Chen, Huan Sun et al.ICML 2024 · 53 citations
- A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce DomainsXianren Zhang, Shreyas Prasad, Di Wang, Qiuhai Zeng et al.ACL 2026 · 8 citations
