Learning the Arrow of Time for Problems in Reinforcement Learning
Nasim Rahaman, Steffen Wolf, Anirudh Goyal, Roman Remme, Yoshua Bengio
Abstract
We humans have an innate understanding of the asymmetric progression of time, which we use to efficiently and safely perceive and manipulate our environment. Drawing inspiration from that, we approach the problem of learning an arrow of time in a Markov (Decision) Process. We illustrate how a learned arrow of time can capture salient information about the environment, which in turn can be used to measure reachability, detect side-effects and to obtain an intrinsic reward signal. Finally, we propose a simple yet effective algorithm to parameterize the problem at hand and learn an arrow of time with a function approximator (here, a deep neural network). Our empirical results span a selection of discrete and continuous environments, and demonstrate for a class of stochastic processes that the learned arrow of time agrees reasonably well with a well known notion of an arrow of time due to Jordan, Kinderlehrer and Otto (1998).
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- When to Ask for Help: Proactive Interventions in Autonomous Reinforcement LearningAnnie Xie, Fahim Tajwar, Archit Sharma, Chelsea FinnNeurIPS 2022 · 30 citations
- There Is No Turning Back: A Self-Supervised Approach for Reversibility-Aware Reinforcement LearningNathan Grinsztajn, Johan Ferret, Olivier Pietquin, Philippe Preux et al.NeurIPS 2021 · 23 citations
- Robust Imitation of a Few Demonstrations with a Backwards ModelJung Yeon Park, Lawson L. S. WongNeurIPS 2022 · 21 citations
Related papers
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu et al.ICML 2020 · 87 citations
- Flexible inference for animal learning rules using neural networksYuhan Helena Liu, Victor Geadah, Jonathan W. PillowNeurIPS 2025
- Seeing the Arrow of Time in Large Multimodal ModelsZihui Xue, Romy Luo, Kristen GraumanNeurIPS 2025 · 30 citations
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 79 citations
- Finite-Time Analysis of Adaptive Temporal Difference Learning with Deep Neural NetworksTao Sun, Dongsheng Li, Bao WangNeurIPS 2022 · 11 citations
