Trajectory-Level Data Augmentation for Offline Reinforcement Learning
Tobias Schmähling, Matthias Burkhardt, Tobias Windisch
Abstract
We propose a data augmentation method for offline reinforcement learning, motivated by active positioning problems. Particularly, our approach enables the training of off-policy models from a limited number of suboptimal trajectories. We introduce a trajectory-based augmentation technique that exploits task structure and the geometric relationship between rewards, value functions, and mathematical properties of logging policies. During data collection, our augmentation supports suboptimal logging policies, leading to higher data quality and improved offline reinforcement learning performance. We provide theoretical justification for these strategies and validate them empirically across positioning tasks of varying dimensionality and under partial observability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
- Synthetic Experience ReplayCong Lu, Philip J. Ball, Yee Whye Teh, Jack Parker-HolderNeurIPS 2023 · 148 citations
Related papers
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 47 citations
- On Trajectory Augmentations for Off-Policy EvaluationGe Gao, Qitong Gao, Xi Yang, Song Ju et al.ICLR 2024 · 5 citations
- Offline Trajectory Optimization for Offline Reinforcement LearningZiqi Zhao, Zhaochun Ren, Liu Yang, Yunsen Liang et al.KDD 2025
- GTA: Generative Trajectory Augmentation with Guidance for Offline Reinforcement LearningJaewoo Lee, Sujin Yun, Taeyoung Yun, Jinkyoo ParkNeurIPS 2024 · 35 citations
- Trajectory Generation with Conservative Value Guidance for Offline Reinforcement LearningTieru Wang, Kunbao Wu, Guoshun NanICLR 2026
