To the Max: Reinventing Reward in Reinforcement Learning
Grigorii Veviurko, Wendelin Boehmer, Mathijs de Weerdt
Abstract
In reinforcement learning (RL), different reward functions can define the same optimal policy but result in drastically different learning performance. For some, the agent gets stuck with a suboptimal behavior, and for others, it solves the task efficiently. Choosing a good reward function is hence an extremely important yet challenging problem. In this paper, we explore an alternative approach for using rewards for learning. We introduce max-reward RL, where an agent optimizes the maximum rather than the cumulative reward. Unlike earlier works, our approach works for deterministic and stochastic environments and can be easily combined with state-of-the-art RL algorithms. In the experiments, we study the performance of max-reward RL algorithms in two goal-reaching environments from Gymnasium-Robotics and demonstrate its benefits over standard RL. The code is available at https://github.com/veviurko/To-the-Max.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64b49bd9-ac1e-4cc4-8e2b-8e6a14bbd510Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu et al.ICLR 2021 · 222 citations
- Reachability Constrained Reinforcement LearningDongjie Yu, Haitong Ma, Sheng-bo Li, Jianyu ChenICML 2022 · 90 citations
- Planning with General Objective Functions: Going Beyond Total RewardsRuosong Wang, Peilin Zhong, Simon S. Du, Ruslan Salakhutdinov et al.NeurIPS 2020 · 23 citations
Related papers
- Test-driven Reinforcement Learning in Continuous ControlZhao Yu, Xiuping Wu, Liangjun KeAAAI 2026
- BC-IRL: Learning Generalizable Reward Functions from DemonstrationsAndrew Szot, Amy Zhang, Dhruv Batra, Zsolt Kira et al.ICLR 2023 · 1 citation
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu et al.ICML 2020 · 87 citations
- When Maximum Entropy Misleads Policy OptimizationRuipeng Zhang, Ya-Chien Chang, Sicun GaoICML 2025
- Learning to Incentivize Other Learning AgentsJiachen Yang, Ang Li, Mehrdad Farajtabar, Peter Sunehag et al.NeurIPS 2020 · 105 citations
