Reasoning with Latent Diffusion in Offline Reinforcement Learning
Siddarth Venkatraman, Shivesh Khaitan, Ravi Tej Akella, John M. Dolan, Jeff Schneider, Glen Berseth
Abstract
Offline reinforcement learning (RL) holds promise as a means to learn high-reward policies from a static dataset, without the need for further environment interactions. However, a key challenge in offline RL lies in effectively stitching portions of suboptimal trajectories from the static dataset while avoiding extrapolation errors arising due to a lack of support in the dataset. Existing approaches use conservative methods that are tricky to tune and struggle with multi-modal data (as we show) or rely on noisy Monte Carlo return-to-go samples for reward conditioning. In this work, we propose a novel approach that leverages the expressiveness of latent diffusion to model in-support trajectory sequences as compressed latent skills. This facilitates learning a Q-function while avoiding extrapolation error via batch-constraining. The latent space is also expressive and gracefully copes with multi-modal data. We show that the learned temporally-abstract latent space encodes richer task-specific information for offline RL tasks as compared to raw state-actions. This improves credit assignment and facilitates faster reward propagation during Q-learning. Our method demonstrates state-of-the-art performance on the D4RL benchmarks, particularly excelling in long-horizon, sparse-reward tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d693ad6-7296-4a5a-950d-5ff86dfc178dCited by top-tier papers20
- Consistency Models as a Rich and Efficient Policy Class for Reinforcement LearningZihan Ding, Chi JinICLR 2024 · 73 citations
- Learning Multimodal Behaviors from Scratch with Diffusion Policy GradientSteven Li, Rickmer Krohn, Tao Chen, Anurag Ajay et al.NeurIPS 2024 · 61 citations
- Generative Trajectory Stitching through Diffusion CompositionYunhao Luo, Utkarsh A. Mishra, Yilun Du, Danfei XuNeurIPS 2025 · 48 citations
- DiffLight: A Partial Rewards Conditioned Diffusion Model for Traffic Signal Control with Missing DataHanyang Chen, Yang Jiang, Shengnan Guo, Xiaowei Mao et al.NeurIPS 2024 · 18 citations
- Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement LearningFranki Nguimatsia Tiofack, Théotime Le Hellard, Fabian Schramm, Nicolas Perrin-Gilbert et al.ICLR 2026 · 8 citations
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
Related papers
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement LearningZhendong Wang, Jonathan J. Hunt, Mingyuan ZhouICLR 2023 · 33 citations
- Constrained Latent Action Policies for Model-Based Offline Reinforcement LearningMarvin Alles, Philip Becker-Ehmck, Patrick van der Smagt, Maximilian KarlNeurIPS 2024 · 5 citations
- Efficient Planning with Latent DiffusionWenhao LiICLR 2024 · 15 citations
- State-Covering Trajectory Stitching for Diffusion PlannersKyowoon Lee, Jaesik ChoiNeurIPS 2025 · 17 citations
- Stitching Sub-trajectories with Conditional Diffusion Model for Goal-Conditioned Offline RLSungyoon Kim, Yunseon Choi, Daiki E. Matsunaga, Kee-Eung KimAAAI 2024 · 19 citations
