Boosting Offline Reinforcement Learning with Action Preference Query
Qisen Yang, Shenzhi Wang, Matthieu Gaetan Lin, Shiji Song, Gao Huang
Abstract
Training practical agents usually involve offline and online reinforcement learning (RL) to balance the policy's performance and interaction costs. In particular, online fine-tuning has become a commonly used method to correct the erroneous estimates of out-of-distribution data learned in the offline training phase. However, even limited online interactions can be inaccessible or catastrophic for high-stake scenarios like healthcare and autonomous driving. In this work, we introduce an interaction-free training scheme dubbed Offline-with-Action-Preferences (OAP). The main insight is that, compared to online fine-tuning, querying the preferences between pre-collected and learned actions can be equally or even more helpful to the erroneous estimate problem. By adaptively encouraging or suppressing policy constraint according to action preferences, OAP could distinguish overestimation from beneficial policy improvement and thus attains a more accurate evaluation of unseen data. Theoretically, we prove a lower bound of the behavior policy's performance improvement brought by OAP. Moreover, comprehensive experiments on the D4RL benchmark and state-of-the-art algorithms demonstrate that OAP yields higher (29% on average) scores, especially on challenging AntMaze tasks (98% higher).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc0cf21a-a458-4881-a918-737dcc9f4f84Cited by top-tier papers8
- Rank-DETR for High Quality Object DetectionYifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan et al.NeurIPS 2023 · 138 citations
- Train Once, Get a Family: State-Adaptive Balances for Offline-to-Online Reinforcement LearningShenzhi Wang, Qisen Yang, Jiawei Gao, Matthieu Gaetan Lin et al.NeurIPS 2023 · 41 citations
- Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RLYang Yue, Rui Lu, Bingyi Kang, Shiji Song et al.NeurIPS 2023 · 28 citations
- Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement LearningTenglong Liu, Yang Li, Yixing Lan, Hao Gao et al.ICML 2024 · 15 citations
- Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise DifferentialsYifan Pu, Jixuan Ying, Qixiu Li, Tianzhu Ye et al.NeurIPS 2025 · 9 citations
Builds on9
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
Related papers
- Budgeting Counterfactual for Offline RLYao Liu, Pratik Chaudhari, Rasool FakoorNeurIPS 2023 · 6 citations
- Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RLQin-Wen Luo, Ming-Kun Xie, Ye-Wen Wang, Sheng-Jun HuangNeurIPS 2024 · 15 citations
- Online Pre-Training for Offline-to-Online Reinforcement LearningYongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong et al.ICML 2025
- State Proficiency-Based Adaptive Fine-Tuning for Offline-to-Online Reinforcement LearningSonglin Li, Wei Xiao, Hao Wu, Xiaodan Zhang et al.AAAI 2026
- Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion GenerationXiao Huang, Xu Liu, Enze Zhang, Tong Yu et al.ICML 2025
