Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language Models
Xiaolei Wang, Xinyu Tang, Xin Zhao, Jingyuan Wang, Ji-Rong Wen
摘要
The recent success of large language models (LLMs) has shown great potential to develop more powerful conversational recommender systems (CRSs), which rely on natural language conversations to satisfy user needs. In this paper, we embark on an investigation into the utilization of ChatGPT for CRSs, revealing the inadequacy of the existing evaluation protocol. It might overemphasize the matching with ground-truth items annotated by humans while neglecting the interactive nature of CRSs. To overcome the limitation, we further propose an interactive Evaluation approach based on LLMs, named iEvaLM, which harnesses LLM-based user simulators. Our evaluation approach can simulate various system-user interaction scenarios. Through the experiments on two public CRS datasets, we demonstrate notable improvements compared to the prevailing evaluation protocol. Furthermore, we emphasize the evaluation of explainability, and Chat-GPT showcases persuasive explanation generation for its recommendations. Our study contributes to a deeper comprehension of the untapped potential of LLMs for CRSs and provides a more flexible and realistic evaluation approach for future research about LLMbased CRSs. The code is available at https: //github.com/RUCAIBox/iEvaLM-CRS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan 等ICLR 2024 · 被引用 188 次
- Adapting Large Language Models by Integrating Collaborative Semantics for RecommendationBowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen 等ICDE 2024 · 被引用 132 次
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng 等ICLR 2024 · 被引用 86 次
- Generative News RecommendationShen Gao, Jiabao Fang, Quan Tu, Zhitao Yao 等WWW 2024 · 被引用 24 次
- A LLM-based Controllable, Scalable, Human-Involved User Simulator Framework for Conversational Recommender SystemsLixi Zhu, Xiaowen Huang, Jitao SangWWW 2025 · 被引用 18 次
它引用的顶会 Paper7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen 等EMNLP 2023 · 被引用 449 次
- Improving Conversational Recommender Systems via Knowledge Graph based Semantic FusionKun Zhou, Wayne Xin Zhao, Shuqing Bian, Yuanhang Zhou 等KDD 2020 · 被引用 309 次
- Towards Unified Conversational Recommender Systems via Knowledge-Enhanced Prompt LearningXiaolei Wang, Kun Zhou, Ji-Rong Wen, Wayne Xin ZhaoKDD 2022 · 被引用 143 次
- Evaluating Conversational Recommender Systems via User SimulationShuo Zhang, Krisztian BalogKDD 2020 · 被引用 80 次
相关 Paper
- Refining Text Generation for Realistic Conversational Recommendation via Direct Preference OptimizationManato Tajiri, Michimasa InabaEMNLP 2025
- ESC-Eval: Evaluating Emotion Support Conversations in Large Language ModelsHaiquan Zhao, Lingyu Li, Shisong Chen, Shuqi Kong 等EMNLP 2024 · 被引用 5 次
- User Experience with LLM-powered Conversational Recommendation Systems: A Case of Music RecommendationSojeong Yun, Youn-kyung LimCHI 2025 · 被引用 13 次
- Collaborative Retrieval for Large Language Model-based Conversational Recommender SystemsYaochen Zhu, Chao Wan, Harald Steck, Dawen Liang 等WWW 2025 · 被引用 15 次
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim 等CHI 2024 · 被引用 81 次
