Learning the Arrow of Time for Problems in Reinforcement Learning
Nasim Rahaman, Steffen Wolf, Anirudh Goyal, Roman Remme, Yoshua Bengio
摘要
We humans have an innate understanding of the asymmetric progression of time, which we use to efficiently and safely perceive and manipulate our environment. Drawing inspiration from that, we approach the problem of learning an arrow of time in a Markov (Decision) Process. We illustrate how a learned arrow of time can capture salient information about the environment, which in turn can be used to measure reachability, detect side-effects and to obtain an intrinsic reward signal. Finally, we propose a simple yet effective algorithm to parameterize the problem at hand and learn an arrow of time with a function approximator (here, a deep neural network). Our empirical results span a selection of discrete and continuous environments, and demonstrate for a class of stochastic processes that the learned arrow of time agrees reasonably well with a well known notion of an arrow of time due to Jordan, Kinderlehrer and Otto (1998).
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- When to Ask for Help: Proactive Interventions in Autonomous Reinforcement LearningAnnie Xie, Fahim Tajwar, Archit Sharma, Chelsea FinnNeurIPS 2022 · 被引用 30 次
- There Is No Turning Back: A Self-Supervised Approach for Reversibility-Aware Reinforcement LearningNathan Grinsztajn, Johan Ferret, Olivier Pietquin, Philippe Preux 等NeurIPS 2021 · 被引用 23 次
- Robust Imitation of a Few Demonstrations with a Backwards ModelJung Yeon Park, Lawson L. S. WongNeurIPS 2022 · 被引用 21 次
相关 Paper
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu 等ICML 2020 · 被引用 87 次
- Flexible inference for animal learning rules using neural networksYuhan Helena Liu, Victor Geadah, Jonathan W. PillowNeurIPS 2025
- Seeing the Arrow of Time in Large Multimodal ModelsZihui Xue, Romy Luo, Kristen GraumanNeurIPS 2025 · 被引用 30 次
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 被引用 79 次
- Finite-Time Analysis of Adaptive Temporal Difference Learning with Deep Neural NetworksTao Sun, Dongsheng Li, Bao WangNeurIPS 2022 · 被引用 11 次
