Beyond Utility: Evaluating LLM as Recommender
Chumeng Jiang, Jiayin Wang, Weizhi Ma, Charles L. A. Clarke, Shuai Wang, Chuhan Wu, Min Zhang
摘要
With the rapid development of Large Language Models (LLMs), recent studies employed LLMs as recommenders to provide personalized information services for distinct users. Despite efforts to improve the accuracy of LLM-based recommendation models, relatively little attention is paid to beyond-utility dimensions. Moreover, there are unique evaluation aspects of LLM-based recommendation models, which have been largely ignored. To bridge this gap, we explore four new evaluation dimensions and propose a multidimensional evaluation framework. The new evaluation dimensions include: 1) history length sensitivity, 2) candidate position bias, 3) generation-involved performance, and 4) hallucinations. All four dimensions have the potential to impact performance, but are largely unnecessary for consideration in traditional systems. Using this multidimensional evaluation framework, along with traditional aspects, we evaluate the performance of seven LLM-based recommenders, with three prompting strategies, comparing them with six traditional models on both ranking and re-ranking tasks on four datasets. We find that LLMs excel at handling tasks with prior knowledge and shorter input histories in the ranking setting, and perform better in the re-ranking setting, beating traditional models across multiple dimensions. However, LLMs exhibit substantial candidate position bias issues, and some models hallucinate nonexistent items much more often than others. We intend our evaluation framework and observations to benefit future research on the use of LLMs as recommenders. The code and data are available at https://github.com/JiangDeccc/EvaLLMasRecommender .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic EncodersYupeng Hou, Jiacheng Li, Xiangjun Fu, Zhankui He 等ACL 2026 · 被引用 346 次
- Tree of Preferences for Diversified RecommendationHanyang Yuan, Ning Tang, Tongya Zheng, Jiarong Xu 等NeurIPS 2025 · 被引用 3 次
- Token-level Collaborative Alignment for LLM-based Generative RecommendationFake Lin, Binbin Hu, Zhi Zheng, Xi Zhu 等WWW 2026 · 被引用 1 次
- Does LLM Focus on the Right Words? Mitigating Context Bias in LLM-based RecommendersBohao Wang, Jiawei Chen, Feng Liu, Changwang Zhang 等WWW 2026 · 被引用 1 次
- InterQuest: A Mixed-Initiative Framework for Dynamic User Interest Modeling in Conversational SearchYu Mei, Yuanxi Wang, Shiyi Wang, Qingyang Wan 等UIST 2025
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
相关 Paper
- GUIDER: Uncertainty Guided Dynamic Re-ranking for Large Language Models Based Recommender SystemsCai Xu, Xujing Wang, Ziyu Guan, Wei Zhao 等AAAI 2026
- ProMax: Exploring the Potential of LLM-derived Profiles with Distribution Shaping for Recommender SystemsYi Zhang, Yiwen Zhang, Kai Zheng, Tong Chen 等SIGIR 2026
- Uncertainty Quantification and Decomposition for LLM-based RecommendationWonbin Kweon, Sanghwan Jang, SeongKu Kang, Hwanjo YuWWW 2025 · 被引用 13 次
- Logit Space Constrained Fine-Tuning for Mitigating Hallucinations in LLM-Based Recommender SystemsJianfeng Deng, Qingfeng Chen, Debo Cheng, Jiuyong Li 等EMNLP 2025 · 被引用 1 次
- Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?Sein Kim, Hongseok Kang, Kibum Kim, Jiwan Kim 等KDD 2025 · 被引用 3 次
