Temporal Distance-aware Transition Augmentation for Offline Model-based Reinforcement Learning
Dongsu Lee, Minhae Kwon
摘要
The goal of offline reinforcement learning (RL) is to extract a high-performance policy from the fixed datasets, minimizing performance degradation due to out-of-distribution (OOD) samples. Offline model-based RL (MBRL) is a promising approach that ameliorates OOD issues by enriching state-action transitions with augmentations synthesized via a learned dynamics model. Unfortunately, seminal offline MBRL methods often struggle in sparse-reward, longhorizon tasks. In this work, we introduce a novel MBRL framework, dubbed Temporal Distance-Aware Transition Augmentation (TempDATA), that generates augmented transitions in a temporally structured latent space rather than in raw state space. To model long-horizon behavior, TempDATA learns a latent abstraction that captures a temporal distance from both trajectory and transition levels of state space. Our experiments confirm that TempDATA outperforms previous offline MBRL methods and achieves matching or surpassing the performance of diffusion-based trajectory augmentation and goal-conditioned RL on the D4RL AntMaze, FrankaKitchen, CALVIN, and pixel-based FrankaKitchen. https://dongsuleetech.github.io/ projects/tempdata/ 4. TempDATA: Temporal Distance-aware Transition Augmentation This section introduces TempDATA, our offline modelbased scheme that augments new transitions, which help
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Policy Compatible Skill Incremental Learning via Lazy Learning InterfaceDaehee Lee, Dongsu Lee, TaeYoon Kwack, Wonje Choi 等NeurIPS 2025 · 被引用 3 次
- Compositional Transduction with Latent Analogies for Offline Goal-Conditioned Reinforcement LearningJunseok Kim, Dohyeong Kim, Mineui Hong, Songhwai OhICML 2026 · 被引用 1 次
- : Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive EnvironmentsSangeun Park, Minhae KwonICML 2026
- Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-LearningSungyoung Lee, Dohyeong Kim, Eshan Balachandar, Zelal Mustafaoglu 等ICML 2026
它引用的顶会 Paper49
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
相关 Paper
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 被引用 47 次
- GTA: Generative Trajectory Augmentation with Guidance for Offline Reinforcement LearningJaewoo Lee, Sujin Yun, Taeyoung Yun, Jinkyoo ParkNeurIPS 2024 · 被引用 35 次
- Offline Trajectory Optimization for Offline Reinforcement LearningZiqi Zhao, Zhaochun Ren, Liu Yang, Yunsen Liang 等KDD 2025
- RTDiff: Reverse Trajectory Synthesis via Diffusion for Offline Reinforcement LearningQianlan Yang, Yu-Xiong WangICLR 2025
- Reasoning with Latent Diffusion in Offline Reinforcement LearningSiddarth Venkatraman, Shivesh Khaitan, Ravi Tej Akella, John M. Dolan 等ICLR 2024 · 被引用 39 次
