Multi-User Reinforcement Learning with Low Rank Rewards
Dheeraj Mysore Nagaraj, Suhas S. Kowshik, Naman Agarwal, Praneeth Netrapalli, Prateek Jain
Abstract
We consider collaborative multi-user reinforcement learning, where multiple users have the same state-action space and transition probabilities but different rewards. Under the assumption that the reward matrix of the N users has a low-rank structure -a standard and practically successful assumption in the collaborative filtering setting -we design algorithms with significantly lower sample complexity compared to the ones that learn the MDP individually for each user. Our main contribution is an algorithm which explores rewards collaboratively with N user-specific MDPs and can learn rewards efficiently in two key settings: tabular MDPs and linear MDPs. When N is large and the rank is constant, the sample complexity per MDP depends logarithmically over the size of the state-space, which represents an exponential reduction (in the state-space size) when compared to the standard "non-collaborative" algorithms. Our main technical contribution is a method to construct policies which obtain data such that low rank matrix completion is possible (without a generative model). This goes beyond the regular RL framework and is closely related to mean field limits of multiagent RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35edfff5-bb7a-41eb-ab86-8701db795a6dCited by top-tier papers2
- The Bandit Whisperer: Communication Learning for Restless BanditsYunfan Zhao, Tonghan Wang, Dheeraj Mysore Nagaraj, Aparna Taneja et al.AAAI 2025 · 6 citations
- What Reward Structure Enables Efficient Sparse-Reward RL? A Proof-of-Concept with Policy-Aware Matrix CompletionIbne Farabi Shihab, SANJEDA AKTER, Anuj SharmaICML 2026
Builds on8
- Multi-Task Reinforcement Learning with Context-based RepresentationsShagun Sodhani, Amy Zhang, Joelle PineauICML 2021 · 241 citations
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 226 citations
- Sharing Knowledge in Multi-Task Deep Reinforcement LearningCarlo D'Eramo, Davide Tateo, Andrea Bonarini, Marcello Restelli et al.ICLR 2020 · 148 citations
- Reward-Free RL is No Harder Than Reward-Aware RL in Linear Markov Decision ProcessesAndrew J. Wagenmaker, Yifang Chen, Max Simchowitz, Simon S. Du et al.ICML 2022 · 61 citations
- Near-Optimal Representation Learning for Linear Bandits and Linear RLJiachen Hu, Xiaoyu Chen, Chi Jin, Lihong Li et al.ICML 2021 · 60 citations
Related papers
- Online Low Rank Matrix CompletionSoumyabrata Pal, Prateek JainICLR 2023 · 2 citations
- Provably efficient multi-task reinforcement learning with model transferChicheng Zhang, Zhi WangNeurIPS 2021 · 20 citations
- Spectral Entry-wise Matrix Estimation for Low-Rank Reinforcement LearningStefan Stojanovic, Yassir Jedra, Alexandre ProutièreNeurIPS 2023 · 9 citations
- The Limits of Transfer Reinforcement Learning with Latent Low-rank StructureTyler Sam, Yudong Chen, Christina Lee YuNeurIPS 2024 · 1 citation
- Estimating α-Rank from A Few Entries with Low Rank Matrix CompletionYali Du, Xue Yan, Xu Chen, Jun Wang et al.ICML 2021 · 9 citations
