User-Interactive Offline Reinforcement Learning
Phillip Swazinna, Steffen Udluft, Thomas A. Runkler
摘要
Offline reinforcement learning algorithms still lack trust in practice due to the risk that the learned policy performs worse than the original policy that generated the dataset or behaves in an unexpected way that is unfamiliar to the user. At the same time, offline RL algorithms are not able to tune their most important hyperparameter -the proximity of the learned policy to the original policy. We propose an algorithm that allows the user to tune this hyperparameter at runtime, thereby addressing both of the above mentioned issues simultaneously. This allows users to start with the original behavior and grant successively greater deviation, as well as stopping at any time when the policy deteriorates or the behavior is too far from the familiar one. Introduction Recently, offline reinforcement learning (RL) methods have shown that it is possible to learn effective policies from a static pre-collected dataset instead of directly interacting with the environment (Laroche et al., 2019; Fujimoto et al., 2019; Yu et al., 2020; Swazinna et al., 2021b). Since direct interaction is in practice usually very costly, these techniques have alleviated a large obstacle on the path of applying reinforcement learning techniques in real world problems. A major issue that these algorithms still face is tuning their most important hyperparameter: The proximity to the original policy. Virtually all algorithms tackling the offline setting have such a hyperparameter, and it is obviously hard to tune, since no interaction with the real environment is permitted until final deployment. Practitioners thus risk being overly conservative (resulting in no improvement) or overly progressive (risking worse performing policies) in their choice. Additionally, one of the arguably largest obstacles on the path to deployment of RL trained policies in most industrial control problems is that (offline) RL algorithms ignore the presence of domain experts, who can be seen as users of the final product -the policy. Instead, most algorithms today can be seen as trying to make human practitioners obsolete. We argue that it is important to provide these users with a utility -something that makes them want to use RL solutions. Other research fields, such as machine learning for medical diagnoses, have already established the idea that domain experts are crucially important to solve the task and complement human users in various ways Babbar et al. (2022); Cai et al. (2019); De-Arteaga et al. (2021); Fard & Pineau (2011); Tang et al. ( 2020 ). We see our work in line with these and other researchers (Shneiderman, 2020; Schmidt et al., 2021) , who suggest that the next generation of AI systems needs to adopt a user-centered approach and develop systems that behave more like an intelligent tool, combining both high levels of human control and
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Train Once, Get a Family: State-Adaptive Balances for Offline-to-Online Reinforcement LearningShenzhi Wang, Qisen Yang, Jiawei Gao, Matthieu Gaetan Lin 等NeurIPS 2023 · 被引用 41 次
- Hypervolume Maximization: A Geometric View of Pareto Set LearningXiaoyuan Zhang, Xi Lin, Bo Xue, Yifan Chen 等NeurIPS 2023 · 被引用 40 次
- Model-based Offline RL via Robust Value-Aware Model Learning with Implicitly Differentiable Adaptive WeightingZhongjian Qiao, Jiafei Lyu, Boxiang Lyu, Yao Shu 等ICLR 2026 · 被引用 5 次
- Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary DynamicsXinyu Zhang, Wenjie Qiu, Yi-Chen Li, Lei Yuan 等ICML 2024 · 被引用 3 次
- Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit ConservatismTianwei Ni, Esther Derman, Vineet Jain, Vincent Taboga 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper23
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
相关 Paper
- Behavior Proximal Policy OptimizationZifeng Zhuang, Kun Lei, Jinxin Liu, Donglin Wang 等ICLR 2023 · 被引用 8 次
- VIPeR: Provably Efficient Algorithm for Offline RL with Neural Function ApproximationThanh Nguyen-Tang, Raman AroraICLR 2023 · 被引用 2 次
- Weighted Policy Constraints for Offline Reinforcement LearningZhiyong Peng, Changlin Han, Yadong Liu, Zongtan ZhouAAAI 2023 · 被引用 18 次
- Human-centric dialog training via offline reinforcement learningNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson 等EMNLP 2020 · 被引用 9 次
- Adaptive Scaling of Policy Constraints for Offline Reinforcement LearningJing Tan, Xiaorui Li, Chao Yao, Xiaojuan Ban 等ICLR 2026 · 被引用 1 次
