Personalized Reward Learning with Interaction-Grounded Learning (IGL)
Jessica Maghakian, Paul Mineiro, Kishan Panaganti, Mark Rucker, Akanksha Saran, Cheng Tan
Abstract
In an era of countless content offerings, recommender systems alleviate information overload by providing users with personalized content suggestions. Due to the scarcity of explicit user feedback, modern recommender systems typically optimize for the same fixed combination of implicit feedback signals across all users. However, this approach disregards a growing body of work highlighting that (i) implicit signals can be used by users in diverse ways, signaling anything from satisfaction to active dislike, and (ii) different users communicate preferences in different ways. We propose applying the recent Interaction Grounded Learning (IGL) paradigm to address the challenge of learning representations of diverse user communication modalities. Rather than requiring a fixed, human-designed reward function, IGL is able to learn personalized reward functions for different users and then optimize directly for the latent user satisfaction. We demonstrate the success of IGL with experiments using simulations as well as with real-world production traces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- P-GenRM: Personalized Generative Reward Model with Test-time User-based ScalingPinyi Zhang, Ting-En Lin, Yuchuan Wu, Jingyang Chen et al.ICLR 2026 · 5 citations
- Provably Efficient Interactive-Grounded Learning with Personalized RewardMengxiao Zhang, Yuheng Zhang, Haipeng Luo, Paul MineiroNeurIPS 2024 · 3 citations
- An Information Theoretic Approach to Interaction-Grounded LearningXiaoyan Hu, Farzan Farnia, Ho-fung LeungICML 2024 · 3 citations
Builds on5
- Clicks can be Cheating: Counterfactual Recommendation for Mitigating Clickbait IssueWenjie Wang, Fuli Feng, Xiangnan He, Hanwang Zhang et al.SIGIR 2021 · 173 citations
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 130 citations
- Interaction-Grounded Learning with Action-Inclusive FeedbackTengyang Xie, Akanksha Saran, Dylan J. Foster, Lekan P. Molu et al.NeurIPS 2022 · 12 citations
- X2T: Training an X-to-Text Typing Interface with Online Learning from User FeedbackJensen Gao, Siddharth Reddy, Glen Berseth, Nicholas Hardy et al.ICLR 2021 · 10 citations
- Interaction-Grounded LearningTengyang Xie, John Langford, Paul Mineiro, Ida MomennejadICML 2021 · 3 citations
Related papers
- Unbiased Reward Modeling from Implicit Feedback for LLM AlignmentHao Wang, Haocheng Yang, Licheng Pan, Zhichao Chen et al.ICML 2026 · 2 citations
- Embed Progressive Implicit Preference in Unified Space for Deep Collaborative FilteringZhongjin Zhang, Yu Liang, Cong Fu, Yuxuan Zhu et al.KDD 2025 · 1 citation
- Multi-Objective Intrinsic Reward Learning for Conversational Recommender SystemsZhendong Chu, Nan Wang, Hongning WangNeurIPS 2023 · 5 citations
- G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior SimulationBoyu Chen, Siran Chen, Zhengrong Yue, Kainan Yan et al.AAAI 2026 · 7 citations
- MTRec: Learning to Align with User Preferences via Mental Reward ModelsMengchen Zhao, Yifan Gao, Yaqing Hou, Xiangyang Li et al.NeurIPS 2025 · 1 citation
