BiTrajDiff: Bidirectional Trajectory Generation with Diffusion Models for Offline Reinforcement Learning
Yunpeng Qing, Yixiao Chi, Shuo Chen, Shunyu Liu, Kexuan Zhou, Sixu Lin, Litao Liu, Changqing Zou
Abstract
Offline Reinforcement Learning (RL) relies on static datasets and often enforces conservative constraints to mitigate out-of-distribution errors, but this inevitably gives rise to learning dataset biases and limited behavioral generalization. Recent Data Augmentation (DA) methods leverage generative models to enrich offline data, yet they mainly operate within a single rollout paradigm and tend to preserve the original trajectory-level connectivity of the dataset. As a result, such methods often introduce local variations and fail to recover connections between distinct behavior patterns. In this paper, we propose Bidirectional Trajectory Diffusion (BiTrajDiff), a novel DA framework that explicitly addresses this limitation. BiTrajDiff decomposes trajectory synthesis into two independent diffusion processes that generate forward-future and backward-history segments conditioned on shared intermediate anchor states. By stitching the generated segments at these anchors, BiTrajDiff can synthesize trajectories that bridge disconnected behavior patterns and recover global trajectory-level connectivity absent from the original data. Extensive experiments demonstrate that BiTrajDiff consistently outperforms advanced DA methods across a range of offline RL backbones. Our code is available at https://github.com/Plankson/BiTrajDiff .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bbfbb18-02e6-4f9d-b299-f843555fdc53Cited by top-tier papers1
Ask how each one uses itBuilds on31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
Related papers
- RTDiff: Reverse Trajectory Synthesis via Diffusion for Offline Reinforcement LearningQianlan Yang, Yu-Xiong WangICLR 2025
- GTA: Generative Trajectory Augmentation with Guidance for Offline Reinforcement LearningJaewoo Lee, Sujin Yun, Taeyoung Yun, Jinkyoo ParkNeurIPS 2024 · 35 citations
- ATraDiff: Accelerating Online Reinforcement Learning with Imaginary TrajectoriesQianlan Yang, Yu-Xiong WangICML 2024 · 2 citations
- DiffStitch: Boosting Offline Reinforcement Learning with Diffusion-based Trajectory StitchingGuanghe Li, Yixiang Shan, Zhengbang Zhu, Ting Long et al.ICML 2024 · 41 citations
- Trajectory Generation with Conservative Value Guidance for Offline Reinforcement LearningTieru Wang, Kunbao Wu, Guoshun NanICLR 2026
