Reinforcement Learning with Non-Exponential Discounting
Matthias Schultheis, Constantin A. Rothkopf, Heinz Koeppl
Abstract
Commonly in reinforcement learning (RL), rewards are discounted over time using an exponential function to model time preference, thereby bounding the expected long-term reward. In contrast, in economics and psychology, it has been shown that humans often adopt a hyperbolic discounting scheme, which is optimal when a specific task termination time distribution is assumed. In this work, we propose a theory for continuous-time model-based reinforcement learning generalized to arbitrary discount functions. This formulation covers the case in which there is a non-exponential random termination time. We derive a Hamilton-Jacobi-Bellman (HJB) equation characterizing the optimal policy and describe how it can be solved using a collocation method, which uses deep learning for function approximation. Further, we show how the inverse RL problem can be approached, in which one tries to recover properties of the discount function given decision data. We validate the applicability of our proposed approach on two simulated problems. Our approach opens the way for the analysis of human discounting in sequential decision-making tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7110a8d5-93c3-40e5-8d49-945213445985Cited by top-tier papers8
- Goal-Conditioned On-Policy Reinforcement LearningXudong Gong, Dawei Feng, Kele Xu, Bo Ding et al.NeurIPS 2024 · 16 citations
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 13 citations
- Probabilistic inverse optimal control for non-linear partially observable systems disentangles perceptual uncertainty and behavioral costsDominik Straub, Matthias Schultheis, Heinz Koeppl, Constantin A. RothkopfNeurIPS 2023 · 8 citations
- A Continuous-time Tractable Model for Present-biased AgentsYasunori Akagi, Hideaki Kim, Takeshi KurashimaAAAI 2025 · 2 citations
- SVL: Goal-Conditioned Reinforcement Learning as Survival LearningFranki Nguimatsia-Tiofack, Fabian Schramm, Théotime Le Hellard, Justin CarpentierICML 2026 · 1 citation
Builds on4
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Inferring learning rules from animal decision-makingZoe Ashwood, Nicholas A. Roy, Ji Hyun Bak, Jonathan W. PillowNeurIPS 2020 · 33 citations
- Inverse Optimal Control Adapted to the Noise Characteristics of the Human Sensorimotor SystemMatthias Schultheis, Dominik Straub, Constantin A. RothkopfNeurIPS 2021 · 25 citations
- POMDPs in Continuous Time and Discrete SpacesBastian Alt, Matthias Schultheis, Heinz KoepplNeurIPS 2020 · 10 citations
Related papers
- Deep learning for continuous-time stochastic control with jumpsPatrick Cheridito, Jean-Loup Dupret, Donatien HainautNeurIPS 2025 · 8 citations
- Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential DiscountingHojin Ko, Jeonggyu HuhICML 2026
- A Temporal Difference Method for Stochastic Continuous DynamicsHaruki Settai, Naoya Takeishi, Takehisa YairiNeurIPS 2025 · 3 citations
- Delta Matters: An Analytically Tractable Model for beta-delta Discounting AgentsYasunori Akagi, Takeshi KurashimaAAAI 2026
- Enhancing Value Function Estimation through First-Order State-Action Dynamics in Offline Reinforcement LearningYun-Hsuan Lien, Ping-Chun Hsieh, Tzu-Mao Li, Yu-Shuen WangICML 2024 · 4 citations
