Model-Free Offline Reinforcement Learning with Enhanced Robustness
Chi Zhang, Zain Ulabedeen Farhat, George K. Atia, Yue Wang
Abstract
Offline reinforcement learning (RL) has gained considerable attention for its ability to learn policies from pre-collected data without real-time interaction, which makes it particularly useful for high-risk applications. However, due to its reliance on offline datasets, existing works inevitably introduce assumptions to ensure effective learning, which, however, often lead to a trade-off between robustness to model mismatch and scalability to large environments. In this paper, we enhance both aspects with a novel double-pessimism principle, which conservatively estimates performance and accounts for both limited data and potential model mismatches, two major reasons for the previous trade-off. We then propose a universal, modelfree algorithm to learn a policy that is robust to potential environment mismatches, which enhances robustness in a scalable manner. Furthermore, we provide a sample complexity analysis of our algorithm when the mismatch is modeled by the l α -norm, which also theoretically demonstrates the efficiency of our method. Extensive experiments further demonstrate that our approach significantly improves robustness in a more scalable manner than existing methods.
Published as a conference paper at ICLR 2025 environments, commonly known as the sim-to-real gap (Zhao et al., 2020), can cause significant performance degradation during deployment. Therefore, it is crucial to enhance the robustness of offline RL to ensure that the learned policies can perform reliably in the presence of such uncertainties. A promising solution is to adapt robust RL frameworks (Iyengar, 2005;Nilim & El Ghaoui, 2004) to the offline setting, as explored recently in (Shi & Chi, 2022;Blanchet et al., 2023). However, these methods often come at the cost of scalability. Due to their inherent structure, robust RL methods typically rely on dynamic planning, which requires knowledge of the full transition dynamics, and are predominantly model-based. This necessitates learning and storing a complete transition model, which is resource-intensive (Zhang et al., 2021a) and limits scalability for large-scale problems.
Recognizing the limitations of current methods and the challenges posed by large-scale problems and model uncertainty, a trade-off between robustness and scalability becomes apparent. Enhancing one typically comes at the expense of the other. This naturally leads to the following question:
Can we develop a unified framework that enhances both scalability and robustness in offline RL?
In this paper, we address this question by presenting a model-free algorithm to learn a policy that is both robust to model uncertainty and scalable to large-scale problems. Our method introduces a principle of double pessimism to simultaneously address two key sources of uncertainty: (1) the uncertainty arising from inaccurate estimations due to the underexplored datasets, and (2) model mismatch between the data collection and deployment environments. We then propose a streamlined conceptual framework, design a model-free algorithm, and provide the first theoretical guarantee of convergence and robustness of our approach. Our contributions can be summarized as follows.
• A Double-Pessimism Principle for Offline RL with Model Mismatch. We begin by framing the challenge of enhancing robustness in offline RL within an offline robust RL framework, where an uncertainty set captures potential environmental mismatches. To solve offline robust RL in a scalable manner, we propose the double-pessimism principle that does not require transition kernel estimations. This principle maintains a conservative estimate of robust performance, obtained directly from data collection without requiring model estimation. We then introduce the first model-free pessimistic robust Q-learning algorithm.
Our algorithm optimizes performance under model mismatch using an offline dataset, while offering greater memory efficiency and more scalability than previous methods.
• First and Near-Optimal Model-Free Algorithm for Offline Robust RL. We provide a rigorous sample complexity analysis for our model-free double-pessimistic robust Q-learning algorithm under the widely used l α -norm uncertainty set. Our analysis shows that, given a dataset satisfying the partial coverage condition (to be introduced later), our algorithm can identify an optimal robust policy with near-optimal sample complexity, comparable to that of model-based offline robust RL and model-free offline non-robust RL. This represents the first sample complexity analysis for model-free robust offline RL, demonstrating its applicability to large-scale problems that require high data efficiency.
• Numerical Experimental Verification of Enhanced Robustness. We conduct extensive numerical experiments to demonstrate the improvements in robustness achieved by our algorithms in both simulated environments (Archibald et al., 1995) and real physics-based Classic Control problems (Brockman et al., 2016)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e34149a0-e6e0-4361-ac79-3f7fad994c50Cited by top-tier papers3
- Sample-Efficient Distributionally Robust Multi-Agent Reinforcement Learning via Online InteractionZain Ulabedeen Farhat, Debamita Ghosh, George K. Atia, Yue WangICLR 2026 · 5 citations
- ORVIT: Near-Optimal Online Distributionally Robust Reinforcement LearningDebamita Ghosh, George K. Atia, Yue WangAAAI 2026 · 2 citations
- Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity AnalysisZachary Roch, George Atia, Yue WangICML 2026 · 1 citation
Builds on40
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao et al.NeurIPS 2021 · 373 citations
Related papers
- A Unified Principle of Pessimism for Offline Reinforcement Learning under Model MismatchYue Wang, Zhongchang Sun, Shaofeng ZouNeurIPS 2024 · 11 citations
- Double Pessimism is Provably Efficient for Distributionally Robust Offline Reinforcement Learning: Generic Algorithm and Robust Partial CoverageJose H. Blanchet, Miao Lu, Tong Zhang, Han ZhongNeurIPS 2023 · 58 citations
- Model-Free Robust ϕ-Divergence Reinforcement Learning Using Both Offline and Online DataKishan Panaganti, Adam Wierman, Eric MazumdarICML 2024 · 12 citations
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement LearningChenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhi-Hong Deng et al.ICLR 2022 · 173 citations
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 130 citations
