Disentangling Likes and Dislikes in Personalized Generative Explainable Recommendation
Ryotaro Shimizu, Takashi Wada, Yu Wang, Johannes Kruse, Sean O'Brien, Sai Htaung Kham, Linxin Song, Yuya Yoshikawa, Yuki Saito, Fugee Tsung, Masayuki Goto, Julian J. McAuley
Abstract
Recent research on explainable recommendation generally frames the task as a standard text generation problem, and evaluates models simply based on the textual similarity between the predicted and ground-truth explanations. However, this approach fails to consider one crucial aspect of the systems: whether their outputs accurately reflect the users' (post-purchase) sentiments, i.e., whether and why they would like and/or dislike the recommended items. To shed light on this issue, we introduce new datasets and evaluation methods that focus on the users' sentiments. Specifically, we construct the datasets by explicitly extracting users' positive and negative opinions from their post-purchase reviews using an LLM, and propose to evaluate systems based on whether the generated explanations 1) align well with the users' sentiments, and 2) accurately identify both positive and negative opinions of users on the target items. We benchmark several recent models on our datasets and demonstrate that achieving strong performance on existing metrics does not ensure that the generated explanations align well with the users' sentiments. Lastly, we find that existing models can provide more sentiment-aware explanations when the users' (predicted) ratings for the target items are directly fed into the models as input. The datasets and benchmark implementation are available at: https://github.com/jchanxtarov/sent_xrec . CCS Concepts • Information systems → Personalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7638e973-96cb-4b49-9b32-697635d49e9cBuilds on11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng et al.ICLR 2024 · 368 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- Dual Learning for Explainable Recommendation: Towards Unifying User Preference Prediction and Review GenerationPeijie Sun, Le Wu, Kun Zhang, Yanjie Fu et al.WWW 2020 · 94 citations
Related papers
- Coherency Improved Explainable Recommendation via Large Language ModelShijie Liu, Ruixin Ding, Weihai Lu, Jun Wang et al.AAAI 2025 · 10 citations
- ReXPlug: Explainable Recommendation using Plug-and-Play Language ModelDeepesh V. Hada, Vijaikumar M, Shirish K. ShevadeSIGIR 2021 · 55 citations
- Comparative Explanations of RecommendationsAobo Yang, Nan Wang, Renqin Cai, Hongbo Deng et al.WWW 2022 · 16 citations
- Following the TRAIL: Predicting and Explaining Tomorrow's Hits with a Fine-Tuned LLMYinan Zhang, Zhixi Chen, Jiazheng Jing, Zhiqi ShenWWW 2026
- Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language ModelsXiaolei Wang, Xinyu Tang, Xin Zhao, Jingyuan Wang et al.EMNLP 2023 · 69 citations
