User-Interactive Offline Reinforcement Learning
Phillip Swazinna, Steffen Udluft, Thomas A. Runkler
Abstract
Offline reinforcement learning algorithms still lack trust in practice due to the risk that the learned policy performs worse than the original policy that generated the dataset or behaves in an unexpected way that is unfamiliar to the user. At the same time, offline RL algorithms are not able to tune their most important hyperparameter -the proximity of the learned policy to the original policy. We propose an algorithm that allows the user to tune this hyperparameter at runtime, thereby addressing both of the above mentioned issues simultaneously. This allows users to start with the original behavior and grant successively greater deviation, as well as stopping at any time when the policy deteriorates or the behavior is too far from the familiar one. Introduction Recently, offline reinforcement learning (RL) methods have shown that it is possible to learn effective policies from a static pre-collected dataset instead of directly interacting with the environment (Laroche et al., 2019; Fujimoto et al., 2019; Yu et al., 2020; Swazinna et al., 2021b). Since direct interaction is in practice usually very costly, these techniques have alleviated a large obstacle on the path of applying reinforcement learning techniques in real world problems. A major issue that these algorithms still face is tuning their most important hyperparameter: The proximity to the original policy. Virtually all algorithms tackling the offline setting have such a hyperparameter, and it is obviously hard to tune, since no interaction with the real environment is permitted until final deployment. Practitioners thus risk being overly conservative (resulting in no improvement) or overly progressive (risking worse performing policies) in their choice. Additionally, one of the arguably largest obstacles on the path to deployment of RL trained policies in most industrial control problems is that (offline) RL algorithms ignore the presence of domain experts, who can be seen as users of the final product -the policy. Instead, most algorithms today can be seen as trying to make human practitioners obsolete. We argue that it is important to provide these users with a utility -something that makes them want to use RL solutions. Other research fields, such as machine learning for medical diagnoses, have already established the idea that domain experts are crucially important to solve the task and complement human users in various ways Babbar et al. (2022); Cai et al. (2019); De-Arteaga et al. (2021); Fard & Pineau (2011); Tang et al. ( 2020 ). We see our work in line with these and other researchers (Shneiderman, 2020; Schmidt et al., 2021) , who suggest that the next generation of AI systems needs to adopt a user-centered approach and develop systems that behave more like an intelligent tool, combining both high levels of human control and
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd9c0a83-e91b-4c39-9121-608ccacc85cfCited by top-tier papers5
- Train Once, Get a Family: State-Adaptive Balances for Offline-to-Online Reinforcement LearningShenzhi Wang, Qisen Yang, Jiawei Gao, Matthieu Gaetan Lin et al.NeurIPS 2023 · 41 citations
- Hypervolume Maximization: A Geometric View of Pareto Set LearningXiaoyuan Zhang, Xi Lin, Bo Xue, Yifan Chen et al.NeurIPS 2023 · 40 citations
- Model-based Offline RL via Robust Value-Aware Model Learning with Implicitly Differentiable Adaptive WeightingZhongjian Qiao, Jiafei Lyu, Boxiang Lyu, Yao Shu et al.ICLR 2026 · 5 citations
- Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary DynamicsXinyu Zhang, Wenjie Qiu, Yi-Chen Li, Lei Yuan et al.ICML 2024 · 3 citations
- Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit ConservatismTianwei Ni, Esther Derman, Vineet Jain, Vincent Taboga et al.ICML 2026 · 1 citation
Builds on23
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
Related papers
- Behavior Proximal Policy OptimizationZifeng Zhuang, Kun Lei, Jinxin Liu, Donglin Wang et al.ICLR 2023 · 8 citations
- VIPeR: Provably Efficient Algorithm for Offline RL with Neural Function ApproximationThanh Nguyen-Tang, Raman AroraICLR 2023 · 2 citations
- Weighted Policy Constraints for Offline Reinforcement LearningZhiyong Peng, Changlin Han, Yadong Liu, Zongtan ZhouAAAI 2023 · 18 citations
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson et al.EMNLP 2020 · 9 citations
- Adaptive Scaling of Policy Constraints for Offline Reinforcement LearningJing Tan, Xiaorui Li, Chao Yao, Xiaojuan Ban et al.ICLR 2026 · 1 citation
