Principled Offline RL in the Presence of Rich Exogenous Information
Riashat Islam, Manan Tomar, Alex Lamb, Yonathan Efroni, Hongyu Zang, Aniket Rajiv Didolkar, Dipendra Misra, Xin Li, Harm van Seijen, Remi Tachet des Combes, John Langford
摘要
Learning to control an agent from offline data collected in a rich pixel-based visual observation space is vital for real-world applications of reinforcement learning (RL). A major challenge in this setting is the presence of input information that is hard to model and irrelevant to controlling the agent. This problem has been approached by the theoretical RL community through the lens of exogenous information, i.e., any control-irrelevant information contained in observations. For example, a robot navigating in busy streets needs to ignore irrelevant information, such as other people walking in the background, textures of objects, or birds in the sky. In this paper, we focus on the setting with visually detailed exogenous information and introduce new offline RL benchmarks that offer the ability to study this problem. We find that contemporary representation learning techniques can fail on datasets where the noise is a complex and time-dependent process, which is prevalent in practical applications. To address these, we propose to use multi-step inverse models to learn Agent-Centric Representations for Offline-RL (ACRO). Despite being simple and rewardfree, we show theoretically and empirically that the representation created by this objective greatly outperforms baselines. Code is provided at https://github.com/manantomar/ agent-centric-representations .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive RepresentationsMingqi Yuan, Tao Yu, Haolin Song, Bo Li 等CVPR 2026 · 被引用 3 次
- Action-Sufficient Goal RepresentationsJinu Hyeon, Woobin Park, Hongjoon Ahn, Taesup MoonICML 2026 · 被引用 1 次
- Model-based Offline Reinforcement Learning with Lower Expectile Q-LearningKwanyoung Park, Youngwoon LeeICLR 2025
- Stable Offline Value Function Learning with Bisimulation-based RepresentationsBrahma S. Pavse, Yudong Chen, Qiaomin Xie, Josiah P. HannaICML 2025
它引用的顶会 Paper26
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
相关 Paper
- DRIBO: Robust Deep Reinforcement Learning via Multi-View Information BottleneckJiameng Fan, Wenchao LiICML 2022 · 被引用 49 次
- Provably Filtering Exogenous Distractors using Multistep Inverse DynamicsYonathan Efroni, Dipendra Misra, Akshay Krishnamurthy, Alekh Agarwal 等ICLR 2022 · 被引用 38 次
- MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement LearningYi Wang, Ningze Zhong, Zhiheng Fu, Longguang Wang 等CVPR 2026
- OGBench: Benchmarking Offline Goal-Conditioned RLSeohong Park, Kevin Frans, Benjamin Eysenbach, Sergey LevineICLR 2025
- An Investigation into Pre-Training Object-Centric Representations for Reinforcement LearningJaesik Yoon, Yi-Fu Wu, Heechul Bae, Sungjin AhnICML 2023 · 被引用 59 次
