Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language Models
Xiaolei Wang, Xinyu Tang, Xin Zhao, Jingyuan Wang, Ji-Rong Wen
Abstract
The recent success of large language models (LLMs) has shown great potential to develop more powerful conversational recommender systems (CRSs), which rely on natural language conversations to satisfy user needs. In this paper, we embark on an investigation into the utilization of ChatGPT for CRSs, revealing the inadequacy of the existing evaluation protocol. It might overemphasize the matching with ground-truth items annotated by humans while neglecting the interactive nature of CRSs. To overcome the limitation, we further propose an interactive Evaluation approach based on LLMs, named iEvaLM, which harnesses LLM-based user simulators. Our evaluation approach can simulate various system-user interaction scenarios. Through the experiments on two public CRS datasets, we demonstrate notable improvements compared to the prevailing evaluation protocol. Furthermore, we emphasize the evaluation of explainability, and Chat-GPT showcases persuasive explanation generation for its recommendations. Our study contributes to a deeper comprehension of the untapped potential of LLMs for CRSs and provides a more flexible and realistic evaluation approach for future research about LLMbased CRSs. The code is available at https: //github.com/RUCAIBox/iEvaLM-CRS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcd6b567-9924-4108-b3fa-ae67025d9647Cited by top-tier papers29
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan et al.ICLR 2024 · 188 citations
- Adapting Large Language Models by Integrating Collaborative Semantics for RecommendationBowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen et al.ICDE 2024 · 132 citations
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng et al.ICLR 2024 · 86 citations
- Generative News RecommendationShen Gao, Jiabao Fang, Quan Tu, Zhitao Yao et al.WWW 2024 · 24 citations
- A LLM-based Controllable, Scalable, Human-Involved User Simulator Framework for Conversational Recommender SystemsLixi Zhu, Xiaowen Huang, Jitao SangWWW 2025 · 18 citations
Builds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
- Improving Conversational Recommender Systems via Knowledge Graph based Semantic FusionKun Zhou, Wayne Xin Zhao, Shuqing Bian, Yuanhang Zhou et al.KDD 2020 · 309 citations
- Towards Unified Conversational Recommender Systems via Knowledge-Enhanced Prompt LearningXiaolei Wang, Kun Zhou, Ji-Rong Wen, Wayne Xin ZhaoKDD 2022 · 143 citations
- Evaluating Conversational Recommender Systems via User SimulationShuo Zhang, Krisztian BalogKDD 2020 · 80 citations
Related papers
- Refining Text Generation for Realistic Conversational Recommendation via Direct Preference OptimizationManato Tajiri, Michimasa InabaEMNLP 2025
- ESC-Eval: Evaluating Emotion Support Conversations in Large Language ModelsHaiquan Zhao, Lingyu Li, Shisong Chen, Shuqi Kong et al.EMNLP 2024 · 5 citations
- User Experience with LLM-powered Conversational Recommendation Systems: A Case of Music RecommendationSojeong Yun, Youn-kyung LimCHI 2025 · 13 citations
- Collaborative Retrieval for Large Language Model-based Conversational Recommender SystemsYaochen Zhu, Chao Wan, Harald Steck, Dawen Liang et al.WWW 2025 · 15 citations
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim et al.CHI 2024 · 81 citations
