Model-Free Offline Reinforcement Learning with Enhanced Robustness
Chi Zhang, Zain Ulabedeen Farhat, George K. Atia, Yue Wang
摘要
Offline reinforcement learning (RL) has gained considerable attention for its ability to learn policies from pre-collected data without real-time interaction, which makes it particularly useful for high-risk applications. However, due to its reliance on offline datasets, existing works inevitably introduce assumptions to ensure effective learning, which, however, often lead to a trade-off between robustness to model mismatch and scalability to large environments. In this paper, we enhance both aspects with a novel double-pessimism principle, which conservatively estimates performance and accounts for both limited data and potential model mismatches, two major reasons for the previous trade-off. We then propose a universal, modelfree algorithm to learn a policy that is robust to potential environment mismatches, which enhances robustness in a scalable manner. Furthermore, we provide a sample complexity analysis of our algorithm when the mismatch is modeled by the l α -norm, which also theoretically demonstrates the efficiency of our method. Extensive experiments further demonstrate that our approach significantly improves robustness in a more scalable manner than existing methods.
Published as a conference paper at ICLR 2025 environments, commonly known as the sim-to-real gap (Zhao et al., 2020), can cause significant performance degradation during deployment. Therefore, it is crucial to enhance the robustness of offline RL to ensure that the learned policies can perform reliably in the presence of such uncertainties. A promising solution is to adapt robust RL frameworks (Iyengar, 2005;Nilim & El Ghaoui, 2004) to the offline setting, as explored recently in (Shi & Chi, 2022;Blanchet et al., 2023). However, these methods often come at the cost of scalability. Due to their inherent structure, robust RL methods typically rely on dynamic planning, which requires knowledge of the full transition dynamics, and are predominantly model-based. This necessitates learning and storing a complete transition model, which is resource-intensive (Zhang et al., 2021a) and limits scalability for large-scale problems.
Recognizing the limitations of current methods and the challenges posed by large-scale problems and model uncertainty, a trade-off between robustness and scalability becomes apparent. Enhancing one typically comes at the expense of the other. This naturally leads to the following question:
Can we develop a unified framework that enhances both scalability and robustness in offline RL?
In this paper, we address this question by presenting a model-free algorithm to learn a policy that is both robust to model uncertainty and scalable to large-scale problems. Our method introduces a principle of double pessimism to simultaneously address two key sources of uncertainty: (1) the uncertainty arising from inaccurate estimations due to the underexplored datasets, and (2) model mismatch between the data collection and deployment environments. We then propose a streamlined conceptual framework, design a model-free algorithm, and provide the first theoretical guarantee of convergence and robustness of our approach. Our contributions can be summarized as follows.
• A Double-Pessimism Principle for Offline RL with Model Mismatch. We begin by framing the challenge of enhancing robustness in offline RL within an offline robust RL framework, where an uncertainty set captures potential environmental mismatches. To solve offline robust RL in a scalable manner, we propose the double-pessimism principle that does not require transition kernel estimations. This principle maintains a conservative estimate of robust performance, obtained directly from data collection without requiring model estimation. We then introduce the first model-free pessimistic robust Q-learning algorithm.
Our algorithm optimizes performance under model mismatch using an offline dataset, while offering greater memory efficiency and more scalability than previous methods.
• First and Near-Optimal Model-Free Algorithm for Offline Robust RL. We provide a rigorous sample complexity analysis for our model-free double-pessimistic robust Q-learning algorithm under the widely used l α -norm uncertainty set. Our analysis shows that, given a dataset satisfying the partial coverage condition (to be introduced later), our algorithm can identify an optimal robust policy with near-optimal sample complexity, comparable to that of model-based offline robust RL and model-free offline non-robust RL. This represents the first sample complexity analysis for model-free robust offline RL, demonstrating its applicability to large-scale problems that require high data efficiency.
• Numerical Experimental Verification of Enhanced Robustness. We conduct extensive numerical experiments to demonstrate the improvements in robustness achieved by our algorithms in both simulated environments (Archibald et al., 1995) and real physics-based Classic Control problems (Brockman et al., 2016)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Sample-Efficient Distributionally Robust Multi-Agent Reinforcement Learning via Online InteractionZain Ulabedeen Farhat, Debamita Ghosh, George K. Atia, Yue WangICLR 2026 · 被引用 5 次
- ORVIT: Near-Optimal Online Distributionally Robust Reinforcement LearningDebamita Ghosh, George K. Atia, Yue WangAAAI 2026 · 被引用 2 次
- Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity AnalysisZachary Roch, George Atia, Yue WangICML 2026 · 被引用 1 次
它引用的顶会 Paper40
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
相关 Paper
- A Unified Principle of Pessimism for Offline Reinforcement Learning under Model MismatchYue Wang, Zhongchang Sun, Shaofeng ZouNeurIPS 2024 · 被引用 11 次
- Double Pessimism is Provably Efficient for Distributionally Robust Offline Reinforcement Learning: Generic Algorithm and Robust Partial CoverageJose H. Blanchet, Miao Lu, Tong Zhang, Han ZhongNeurIPS 2023 · 被引用 58 次
- Model-Free Robust ϕ-Divergence Reinforcement Learning Using Both Offline and Online DataKishan Panaganti, Adam Wierman, Eric MazumdarICML 2024 · 被引用 12 次
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement LearningChenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhi-Hong Deng 等ICLR 2022 · 被引用 173 次
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 被引用 130 次
