Reinforcement Learning with Non-Exponential Discounting
Matthias Schultheis, Constantin A. Rothkopf, Heinz Koeppl
摘要
Commonly in reinforcement learning (RL), rewards are discounted over time using an exponential function to model time preference, thereby bounding the expected long-term reward. In contrast, in economics and psychology, it has been shown that humans often adopt a hyperbolic discounting scheme, which is optimal when a specific task termination time distribution is assumed. In this work, we propose a theory for continuous-time model-based reinforcement learning generalized to arbitrary discount functions. This formulation covers the case in which there is a non-exponential random termination time. We derive a Hamilton-Jacobi-Bellman (HJB) equation characterizing the optimal policy and describe how it can be solved using a collocation method, which uses deep learning for function approximation. Further, we show how the inverse RL problem can be approached, in which one tries to recover properties of the discount function given decision data. We validate the applicability of our proposed approach on two simulated problems. Our approach opens the way for the analysis of human discounting in sequential decision-making tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Goal-Conditioned On-Policy Reinforcement LearningXudong Gong, Dawei Feng, Kele Xu, Bo Ding 等NeurIPS 2024 · 被引用 16 次
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 被引用 13 次
- Probabilistic inverse optimal control for non-linear partially observable systems disentangles perceptual uncertainty and behavioral costsDominik Straub, Matthias Schultheis, Heinz Koeppl, Constantin A. RothkopfNeurIPS 2023 · 被引用 8 次
- A Continuous-time Tractable Model for Present-biased AgentsYasunori Akagi, Hideaki Kim, Takeshi KurashimaAAAI 2025 · 被引用 2 次
- SVL: Goal-Conditioned Reinforcement Learning as Survival LearningFranki Nguimatsia-Tiofack, Fabian Schramm, Théotime Le Hellard, Justin CarpentierICML 2026 · 被引用 1 次
它引用的顶会 Paper4
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Inferring learning rules from animal decision-makingZoe Ashwood, Nicholas A. Roy, Ji Hyun Bak, Jonathan W. PillowNeurIPS 2020 · 被引用 33 次
- Inverse Optimal Control Adapted to the Noise Characteristics of the Human Sensorimotor SystemMatthias Schultheis, Dominik Straub, Constantin A. RothkopfNeurIPS 2021 · 被引用 25 次
- POMDPs in Continuous Time and Discrete SpacesBastian Alt, Matthias Schultheis, Heinz KoepplNeurIPS 2020 · 被引用 10 次
相关 Paper
- Deep learning for continuous-time stochastic control with jumpsPatrick Cheridito, Jean-Loup Dupret, Donatien HainautNeurIPS 2025 · 被引用 8 次
- Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential DiscountingHojin Ko, Jeonggyu HuhICML 2026
- A Temporal Difference Method for Stochastic Continuous DynamicsHaruki Settai, Naoya Takeishi, Takehisa YairiNeurIPS 2025 · 被引用 3 次
- Delta Matters: An Analytically Tractable Model for beta-delta Discounting AgentsYasunori Akagi, Takeshi KurashimaAAAI 2026
- Enhancing Value Function Estimation through First-Order State-Action Dynamics in Offline Reinforcement LearningYun-Hsuan Lien, Ping-Chun Hsieh, Tzu-Mao Li, Yu-Shuen WangICML 2024 · 被引用 4 次
