Personalized Reward Learning with Interaction-Grounded Learning (IGL)
Jessica Maghakian, Paul Mineiro, Kishan Panaganti, Mark Rucker, Akanksha Saran, Cheng Tan
摘要
In an era of countless content offerings, recommender systems alleviate information overload by providing users with personalized content suggestions. Due to the scarcity of explicit user feedback, modern recommender systems typically optimize for the same fixed combination of implicit feedback signals across all users. However, this approach disregards a growing body of work highlighting that (i) implicit signals can be used by users in diverse ways, signaling anything from satisfaction to active dislike, and (ii) different users communicate preferences in different ways. We propose applying the recent Interaction Grounded Learning (IGL) paradigm to address the challenge of learning representations of diverse user communication modalities. Rather than requiring a fixed, human-designed reward function, IGL is able to learn personalized reward functions for different users and then optimize directly for the latent user satisfaction. We demonstrate the success of IGL with experiments using simulations as well as with real-world production traces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- P-GenRM: Personalized Generative Reward Model with Test-time User-based ScalingPinyi Zhang, Ting-En Lin, Yuchuan Wu, Jingyang Chen 等ICLR 2026 · 被引用 5 次
- Provably Efficient Interactive-Grounded Learning with Personalized RewardMengxiao Zhang, Yuheng Zhang, Haipeng Luo, Paul MineiroNeurIPS 2024 · 被引用 3 次
- An Information Theoretic Approach to Interaction-Grounded LearningXiaoyan Hu, Farzan Farnia, Ho-fung LeungICML 2024 · 被引用 3 次
它引用的顶会 Paper5
- Clicks can be Cheating: Counterfactual Recommendation for Mitigating Clickbait IssueWenjie Wang, Fuli Feng, Xiangnan He, Hanwang Zhang 等SIGIR 2021 · 被引用 173 次
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 被引用 130 次
- Interaction-Grounded Learning with Action-Inclusive FeedbackTengyang Xie, Akanksha Saran, Dylan J. Foster, Lekan P. Molu 等NeurIPS 2022 · 被引用 12 次
- X2T: Training an X-to-Text Typing Interface with Online Learning from User FeedbackJensen Gao, Siddharth Reddy, Glen Berseth, Nicholas Hardy 等ICLR 2021 · 被引用 10 次
- Interaction-Grounded LearningTengyang Xie, John Langford, Paul Mineiro, Ida MomennejadICML 2021 · 被引用 3 次
相关 Paper
- Unbiased Reward Modeling from Implicit Feedback for LLM AlignmentHao Wang, Haocheng Yang, Licheng Pan, Zhichao Chen 等ICML 2026 · 被引用 2 次
- Embed Progressive Implicit Preference in Unified Space for Deep Collaborative FilteringZhongjin Zhang, Yu Liang, Cong Fu, Yuxuan Zhu 等KDD 2025 · 被引用 1 次
- Multi-Objective Intrinsic Reward Learning for Conversational Recommender SystemsZhendong Chu, Nan Wang, Hongning WangNeurIPS 2023 · 被引用 5 次
- G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior SimulationBoyu Chen, Siran Chen, Zhengrong Yue, Kainan Yan 等AAAI 2026 · 被引用 7 次
- MTRec: Learning to Align with User Preferences via Mental Reward ModelsMengchen Zhao, Yifan Gao, Yaqing Hou, Xiangyang Li 等NeurIPS 2025 · 被引用 1 次
