Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior Data
Yuxuan Lu, Jing Huang, Yan Han, Bingsheng Yao, Sisong Bei, Yaochen Xie, Yisi Sang, Qi He, Dakuo Wang
摘要
Recent research shows that LLM Agents can generate ``believable''human behaviors via prompt-only methods, and such agents have been increasingly adopted in downstream applications. However, existing evaluation of these agents only focuses on qualitative believability (whether human raters think they are accurate), leaving open questions of whether LLM agents can accurately generate step-by-step actions mimicking a particular human's behavior in a multi-turn interaction task. In this work, we take shopping as a case study and present the first large-scale quantitative evaluation of state-of-the-art LLMs'ability to accurately simulate human behavior. Using real-world data from 31,865 online shopping sessions containing 230,965 user actions, our evaluation reveals that prompt-based LLMs (DeepSeek-R1, Llama, Claude) achieve only 11.86% accuracy in generating human actions, highlighting a substantial gap in actual behavioral accuracy. Through experiments, we also showcase that strategies as simple as fine-tuning LLMs on real human click-through data augmented with synthesized reasoning traces can greatly enhance models'performance. The fine-tuned Qwen2.5-7B achieves 17.26% action generation accuracy and 33.86% F1 score on final purchase prediction, representing substantial improvements of 5.4% and 13.85% over prompt-only baselines. This work establishes the first rigorous benchmark for human behavior simulation and provides actionable insights for developing more accurate LLM agents for future downstream applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic EvaluationsPreethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde, Sameer Singh 等ACL 2026 · 被引用 19 次
- Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement LearningYimeng Zhang, Tian Wang, Jiri Gesi, Ziyi Wang 等ICLR 2026 · 被引用 18 次
- HumanLM: Simulating Users with State Alignment Beats Response ImitationShirley Wu, Evelyn Choi, Arpandeep Khatua, Zhanghan Wang 等ICML 2026
它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsShunyu Yao, Howard Chen, John Yang, Karthik NarasimhanNeurIPS 2022 · 被引用 1,477 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari 等ICLR 2024 · 被引用 359 次
相关 Paper
- OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior SimulationZiyi Wang, Yuxuan Lu, Wenbo Li, Amirali Amini 等ACL 2026 · 被引用 26 次
- ShopSimulator: Evaluating and Exploring RL-Driven LLM Agent for Shopping AssistantsPei Wang, Yanan Wu, Xiaoshuai Song, Weixun Wang 等ACL 2026 · 被引用 5 次
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag 等ACL 2025 · 被引用 15 次
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas 等ICLR 2026 · 被引用 1 次
- ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue AgentsTianjian Liu, Fanqi Wan, Jiajian Guo, Xiaojun QuanACL 2026
