Learning Retrospective Knowledge with Reverse Reinforcement Learning
Shangtong Zhang, Vivek Veeriah, Shimon Whiteson
Abstract
We present a Reverse Reinforcement Learning (Reverse RL) approach for representing retrospective knowledge. General Value Functions (GVFs) have enjoyed great success in representing predictive knowledge, i.e., answering questions about possible future outcomes such as "how much fuel will be consumed in expectation if we drive from A to B?". GVFs, however, cannot answer questions like "how much fuel do we expect a car to have given it is at B at time ?". To answer this question, we need to know when that car had a full tank and how that car came to B. Since such questions emphasize the influence of possible past events on the present, we refer to their answers as retrospective knowledge. In this paper, we show how to represent retrospective knowledge with Reverse GVFs, which are trained via Reverse RL. We demonstrate empirically the utility of Reverse GVFs in both representation learning and anomaly detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88e13fb4-3dc8-436c-8aaf-f8d3cbca80abCited by top-tier papers6
- Forethought and Hindsight in Credit AssignmentVeronica Chelu, Doina Precup, Hado van HasseltNeurIPS 2020 · 29 citations
- Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement LearningYun Qu, Yuhang Jiang, Boyuan Wang, Yixiu Mao et al.AAAI 2025 · 29 citations
- Robust Imitation of a Few Demonstrations with a Backwards ModelJung Yeon Park, Lawson L. S. WongNeurIPS 2022 · 21 citations
- Learning Expected Emphatic Traces for Deep RLRay Jiang, Shangtong Zhang, Veronica Chelu, Adam White et al.AAAI 2022 · 14 citations
- From Past to Future: Rethinking Eligibility TracesDhawal Gupta, Scott M. Jordan, Shreyas Chaudhari, Bo Liu et al.AAAI 2024 · 5 citations
Builds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
- GradientDICE: Rethinking Generalized Offline Estimation of Stationary ValuesShangtong Zhang, Bo Liu, Shimon WhitesonICML 2020 · 107 citations
- The Value-Improvement Path: Towards Better Representations for Reinforcement LearningWill Dabney, André Barreto, Mark Rowland, Robert Dadashi et al.AAAI 2021 · 76 citations
Related papers
- A Unifying Framework of Off-Policy General Value Function EvaluationTengyu Xu, Zhuoran Yang, Zhaoran Wang, Yingbin LiangNeurIPS 2022 · 2 citations
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing et al.NeurIPS 2020 · 12 citations
- Gamma-Nets: Generalizing Value Estimation over TimescaleCraig Sherstan, Shibhansh Dohare, James MacGlashan, Johannes Günther et al.AAAI 2020 · 14 citations
- Discovering Object-Centric Generalized Value Functions From PixelsSomjit Nath, Gopeshh Raaj Subbaraj, Khimya Khetarpal, Samira Ebrahimi KahouICML 2023 · 2 citations
- TimeTraveler: Reinforcement Learning for Temporal Knowledge Graph ForecastingHaohai Sun, Jialun Zhong, Yunpu Ma, Zhen Han et al.EMNLP 2021 · 164 citations
