On Trajectory Augmentations for Off-Policy Evaluation
Ge Gao, Qitong Gao, Xi Yang, Song Ju, Miroslav Pajic, Min Chi
Abstract
In the realm of reinforcement learning (RL), off-policy evaluation (OPE) holds a pivotal position, especially in high-stake human-centric scenarios such as elearning and healthcare. Applying OPE to these domains is often challenging with scarce and underrepresentative offline training trajectories. Data augmentation has been a successful technique to enrich training data. However, directly employing existing data augmentation methods to OPE may not be feasible, due to the Markovian nature within the offline trajectories and the desire for generalizability across diverse target policies. In this work, we propose an offline trajectory augmentation approach, named OAT, to specifically facilitate OPE in human-involved scenarios. We propose sub-trajectory mining to extract potentially valuable sub-trajectories from offline data, and diversify the behaviors within those sub-trajectories by varying coverage of the state-action space. Our work was empirically evaluated in a wide array of environments, encompassing both simulated scenarios and realworld domains like robotic control, healthcare, and e-learning, where the training trajectories include varying levels of coverage of the state-action space. By enhancing the performance of a variety of OPE methods, our work offers a promising path forward for tackling OPE challenges in situations where human-centric data may be limited or underrepresentative.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d43e83e1-5aa5-479a-8744-ce0b35238882Cited by top-tier papers1
Ask how each one uses itBuilds on25
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
Related papers
- Offline Trajectory Optimization for Offline Reinforcement LearningZiqi Zhao, Zhaochun Ren, Liu Yang, Yunsen Liang et al.KDD 2025
- Trajectory-Level Data Augmentation for Offline Reinforcement LearningTobias Schmähling, Matthias Burkhardt, Tobias WindischICML 2026
- GTA: Generative Trajectory Augmentation with Guidance for Offline Reinforcement LearningJaewoo Lee, Sujin Yun, Taeyoung Yun, Jinkyoo ParkNeurIPS 2024 · 35 citations
- Counterfactual-Augmented Importance Sampling for Semi-Offline Policy EvaluationShengpu Tang, Jenna WiensNeurIPS 2023 · 8 citations
- Off-Policy Evaluation for Human FeedbackQitong Gao, Ge Gao, Juncheng Dong, Vahid Tarokh et al.NeurIPS 2023 · 13 citations
