Offline Transition Modeling via Contrastive Energy Learning
Ruifeng Chen, Chengxing Jia, Zefang Huang, Tian-Shuo Liu, Xu-Hui Liu, Yang Yu
摘要
Learning a high-quality transition model is of great importance for sequential decision-making tasks, especially in offline settings. Nevertheless, the complex behaviors of transition dynamics in real-world environments pose challenges for the standard forward models because of their inductive bias towards smooth regressors, conflicting with the inherent nature of transitions such as discontinuity or large curvature. In this work, we propose to model the transition probability implicitly through a scalar-value energy function, which enables not only flexible distribution prediction but also capturing complex transition behaviors. The Energy-based Transition Models (ETM) are shown to accurately fit the discontinuous transition functions and better generalize to out-of-distribution transition data. Furthermore, we demonstrate that energy-based transition models improve the evaluation accuracy and significantly outperform other off-policy evaluation methods in DOPE benchmark. Finally, we show that energy-based transition models also benefit reinforcement learning and outperform prior offline RL algorithms in D4RL Gym-Mujoco tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Policy Learning from Tutorial Books via Understanding, Rehearsing and IntrospectingXiong-Hui Chen, Ziyan Wang, Yali Du, Shengyi Jiang 等NeurIPS 2024 · 被引用 4 次
- Multi-Agent Imitation by Learning and Sampling from Factorized Soft Q-FunctionYi-Chen Li, Zhongxiang Ling, Tao Jiang, Fuxiang Zhang 等NeurIPS 2025 · 被引用 3 次
- Policy-conditioned Environment Models are More GeneralizableRuifeng Chen, Xiong-Hui Chen, Yihao Sun, Siyuan Xiao 等ICML 2024 · 被引用 1 次
- ADM-v2: Pursuing Full-Horizon Roll-out in Dynamics Models for Offline Policy Learning and EvaluationHaoxin Lin, Siyuan Xiao, Yi-Chen Li, Zhilong Zhang 等ICLR 2026
- Trajectory World Models for Heterogeneous EnvironmentsShaofeng Yin, Jialong Wu, Siqiao Huang, Xingjian Su 等ICML 2025
它引用的顶会 Paper27
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 1,294 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 被引用 1,115 次
相关 Paper
- Variational Latent Branching Model for Off-Policy EvaluationQitong Gao, Ge Gao, Min Chi, Miroslav PajicICLR 2023 · 被引用 2 次
- Autoregressive Dynamics Models for Offline Policy Evaluation and OptimizationMichael R. Zhang, Thomas Paine, Ofir Nachum, Cosmin Paduraru 等ICLR 2021 · 被引用 52 次
- Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement LearningFan-Ming Luo, Tian Xu, Xingchen Cao, Yang YuICLR 2024 · 被引用 16 次
- EMFuse: Energy-based Model Fusion for Decision MakingKejie He, Yi-Chen Li, Yang YuICLR 2026
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing 等NeurIPS 2020 · 被引用 12 次
