Reinforcement Learning Can Be More Efficient with Multiple Rewards
Christoph Dann, Yishay Mansour, Mehryar Mohri
Abstract
Reward design is one of the most critical and challenging aspects when formulating a task as a reinforcement learning (RL) problem. In practice, it often takes several attempts of reward specification and learning with it in order to find one that leads to sample-efficient learning of the desired behavior. Instead, in this work, we study whether directly incorporating multiple alternate reward formulations of the same task in a single agent can lead to faster learning. We analyze multi-reward extensions of action-elimination algorithms and prove more favorable instance-dependent regret bounds compared to their single-reward counterparts, both in multi-armed bandits and in tabular Markov decision processes. Our bounds scale for each state-action pair with the inverse of the largest gap among all reward functions. This suggests that learning with multiple rewards can indeed be more sample-efficient, as long as the rewards agree on an optimal policy. We further prove that when rewards do not agree, multireward action elimination in multi-armed bandits still learns a policy that is good across all reward functions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b534fb12-97f2-49b0-8e3d-5a220e5f959fCited by top-tier papers8
- Multi-Reward Best Policy IdentificationAlessio Russo, Filippo VannellaNeurIPS 2024 · 6 citations
- Autoregressive Multi-trait Essay Scoring via Reinforcement Learning with Scoring-aware Multiple RewardsHeejin Do, Sangwon Ryu, Gary Geunbae LeeEMNLP 2024 · 5 citations
- Distributional Reinforcement Learning with Regularized Wasserstein LossKe Sun, Yingnan Zhao, Wulong Liu, Bei Jiang et al.NeurIPS 2024 · 2 citations
- RESTL: Reinforcement Learning Guided by Multi-Aspect Rewards for Signal Temporal Logic TransformationYue Fang, Zhi Jin, Jie An, Hongshen Chen et al.AAAI 2026 · 1 citation
- Discovering Implicit Large Language Model Alignment ObjectivesEdward Chen, Sanmi Koyejo, Carlos GuestrinICML 2026
Builds on10
- Model Selection in Contextual Stochastic Bandit ProblemsAldo Pacchiano, My Phan, Yasin Abbasi-Yadkori, Anup Rao et al.NeurIPS 2020 · 107 citations
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu et al.ICML 2020 · 87 citations
- Guarantees for Epsilon-Greedy Reinforcement Learning with Function ApproximationChristoph Dann, Yishay Mansour, Mehryar Mohri, Ayush Sekhari et al.ICML 2022 · 76 citations
- Near-Optimal Representation Learning for Linear Bandits and Linear RLJiachen Hu, Xiaoyu Chen, Chi Jin, Lihong Li et al.ICML 2021 · 60 citations
- The best of both worlds: stochastic and adversarial episodic MDPs with unknown transitionTiancheng Jin, Longbo Huang, Haipeng LuoNeurIPS 2021 · 51 citations
Related papers
- Task-agnostic Exploration in Reinforcement LearningXuezhou Zhang, Yuzhe Ma, Adish SinglaNeurIPS 2020 · 56 citations
- Understanding the Complexity Gains of Single-Task RL with a CurriculumQiyang Li, Yuexiang Zhai, Yi Ma, Sergey LevineICML 2023 · 21 citations
- Improved Corruption Robust Algorithms for Episodic Reinforcement LearningYifang Chen, Simon S. Du, Kevin JamiesonICML 2021 · 27 citations
- Rewriting History with Inverse RL: Hindsight Inference for Policy ImprovementBen Eysenbach, Xinyang Geng, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2020 · 96 citations
- Orchestrated Value Mapping for Reinforcement LearningMehdi Fatemi, Arash TavakoliICLR 2022 · 8 citations
