Taylor Expansion of Discount Factors
Yunhao Tang, Mark Rowland, Rémi Munos, Michal Valko
Abstract
In practical reinforcement learning (RL), the discount factor used for estimating value functions often differs from that used for defining the evaluation objective. In this work, we study the effect that this discrepancy of discount factors has during learning, and discover a family of objectives that interpolate value functions of two distinct discount factors. Our analysis suggests new ways for estimating value functions and performing policy optimization updates, which demonstrate empirical performance gains. This framework also leads to new insights on commonly-used deep RL heuristic modifications to policy optimization algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82a1139b-e231-4ec2-bed2-21f009672401Cited by top-tier papers2
- Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount FactorJulien Grand-Clément, Marek PetrikNeurIPS 2023 · 25 citations
- Thresholds for sensitive optimality and Blackwell optimality in stochastic gamesStephane Gaubert, Julien Grand-Clément, Ricardo KatzNeurIPS 2025 · 1 citation
Builds on2
Related papers
- A Parametric Class of Approximate Gradient Updates for Policy OptimizationRamki Gummadi, Saurabh Kumar, Junfeng Wen, Dale SchuurmansICML 2022
- On the Role of Discount Factor in Offline Reinforcement LearningHao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie ZhangICML 2022 · 26 citations
- Learning Fair Policies in Multi-Objective (Deep) Reinforcement Learning with Average and Discounted RewardsUmer Siddique, Paul Weng, Matthieu ZimmerICML 2020 · 1 citation
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement LearningAhmed Hendawy, Henrik Metternich, Théo Vincent, Mahdi Kallel et al.ICLR 2026 · 4 citations
- Correcting discount-factor mismatch in on-policy policy gradient methodsFengdi Che, Gautham Vasan, A. Rupam MahmoodICML 2023 · 10 citations
