More Than What Was Chosen: LLM-based Explainable Recommendation Beyond Noisy User Preferences
Chung Park, Hyeongjun Yun, Taesan Kim, Junui Hong, Dongjoon Hong, Mira Myong, Jihoon Oh, Mincheol Cho, Kijung Park, Minsung Choi, Jihwan Seok, Jaegul Choo
Abstract
Recommender systems traditionally rely on the principle of Revealed Preference (RP), which assumes that observed user behaviors faithfully reflect underlying interests. While effective at scale, this assumption is fragile in practice, as real-world choices are often noisy and inconsistent. Thus, even LLM-based recommendation models (LLM-Rec) equipped with advanced reasoning capabilities may fail to capture genuine user preferences and often produce rationales of limited persuasiveness. To address this issue, we introduce the concept of Coherent Preference (CP), which complements RP by favoring items that are logically and causally coherent with user interaction history. Building on this perspective, we propose Conflict-Aware Direct Preference Optimization (C-APO), an LLM-Rec framework that jointly optimizes RP and CP while adaptively reconciling their agreement and conflict, delivering robust recommendation performance and logically consistent rationales. We construct a unified ordering approach that combines the RP signal, based on chosen versus unobserved items, with the CP signal, which ranks items by their logical consistency with past interaction history. In this unified preference ordering, we dynamically adjust the influence of each signal depending on whether RP and CP agree or conflict, allowing the model to better capture user intent and generate more plausible recommendations. On the Amazon Review dataset, our approach consistently outperforms approximately 20 state-of-the-art baseline models in both recommendation performance and rationale quality, achieving a 1.65 relative improvement in click-through rate during deployment, thereby demonstrating its practical utility. The code and dataset are available at https://github.com/cpark88/C-APO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ae52cc5-3d9a-4d96-97d5-63bc121cd2a6Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine et al.ICLR 2026 · 218 citations
- On Softmax Direct Preference Optimization for RecommendationYuxin Chen, Junfei Tan, An Zhang, Zhengyi Yang et al.NeurIPS 2024 · 126 citations
Related papers
- Think Wise, Collaborate Effectively: A Rationale-Aware LLM-Based Recommender with Reinforcement Learning from Collaborative SignalsChung Park, Taesan Kim, Hyeongjun Yun, Dongjoon Hong et al.AAAI 2026
- What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal ContextZhongyu Ouyang, Qianlong Wen, Chunhui Zhang, Yanfang Ye et al.ACL 2026
- Catalog-Native LLM: Speaking Item-ID dialect with Less Entanglement for RecommendationReza Shirkavand, Xiaokai Wei, Chen Wang, Zheng Hui et al.ICLR 2026 · 4 citations
- CoT4Rec: Revealing User Preferences Through Chain of Thought for Recommender SystemsWeiqi Yue, Yuyu Yin, Xin Zhang, Binbin Shi et al.AAAI 2025 · 10 citations
- Mining Informative Interests via Latent Cross Reasoning for Search Enhanced RecommendationTeng Shi, Weicong Qin, Weijie Yu, Xiao Zhang et al.SIGIR 2026
