Multi-User Reinforcement Learning with Low Rank Rewards
Dheeraj Mysore Nagaraj, Suhas S. Kowshik, Naman Agarwal, Praneeth Netrapalli, Prateek Jain
摘要
We consider collaborative multi-user reinforcement learning, where multiple users have the same state-action space and transition probabilities but different rewards. Under the assumption that the reward matrix of the N users has a low-rank structure -a standard and practically successful assumption in the collaborative filtering setting -we design algorithms with significantly lower sample complexity compared to the ones that learn the MDP individually for each user. Our main contribution is an algorithm which explores rewards collaboratively with N user-specific MDPs and can learn rewards efficiently in two key settings: tabular MDPs and linear MDPs. When N is large and the rank is constant, the sample complexity per MDP depends logarithmically over the size of the state-space, which represents an exponential reduction (in the state-space size) when compared to the standard "non-collaborative" algorithms. Our main technical contribution is a method to construct policies which obtain data such that low rank matrix completion is possible (without a generative model). This goes beyond the regular RL framework and is closely related to mean field limits of multiagent RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The Bandit Whisperer: Communication Learning for Restless BanditsYunfan Zhao, Tonghan Wang, Dheeraj Mysore Nagaraj, Aparna Taneja 等AAAI 2025 · 被引用 6 次
- What Reward Structure Enables Efficient Sparse-Reward RL? A Proof-of-Concept with Policy-Aware Matrix CompletionIbne Farabi Shihab, SANJEDA AKTER, Anuj SharmaICML 2026
它引用的顶会 Paper8
- Multi-Task Reinforcement Learning with Context-based RepresentationsShagun Sodhani, Amy Zhang, Joelle PineauICML 2021 · 被引用 241 次
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
- Sharing Knowledge in Multi-Task Deep Reinforcement LearningCarlo D'Eramo, Davide Tateo, Andrea Bonarini, Marcello Restelli 等ICLR 2020 · 被引用 148 次
- Reward-Free RL is No Harder Than Reward-Aware RL in Linear Markov Decision ProcessesAndrew J. Wagenmaker, Yifang Chen, Max Simchowitz, Simon S. Du 等ICML 2022 · 被引用 61 次
- Near-Optimal Representation Learning for Linear Bandits and Linear RLJiachen Hu, Xiaoyu Chen, Chi Jin, Lihong Li 等ICML 2021 · 被引用 60 次
相关 Paper
- Online Low Rank Matrix CompletionSoumyabrata Pal, Prateek JainICLR 2023 · 被引用 2 次
- Provably efficient multi-task reinforcement learning with model transferChicheng Zhang, Zhi WangNeurIPS 2021 · 被引用 20 次
- Spectral Entry-wise Matrix Estimation for Low-Rank Reinforcement LearningStefan Stojanovic, Yassir Jedra, Alexandre ProutièreNeurIPS 2023 · 被引用 9 次
- The Limits of Transfer Reinforcement Learning with Latent Low-rank StructureTyler Sam, Yudong Chen, Christina Lee YuNeurIPS 2024 · 被引用 1 次
- Estimating α-Rank from A Few Entries with Low Rank Matrix CompletionYali Du, Xue Yan, Xu Chen, Jun Wang 等ICML 2021 · 被引用 9 次
