Doubly Robust Distributionally Robust Off-Policy Evaluation and Learning
Nathan Kallus, Xiaojie Mao, Kaiwen Wang, Zhengyuan Zhou
Abstract
Off-policy evaluation and learning (OPE/L) use offline observational data to make better decisions, which is crucial in applications where online experimentation is limited. However, depending entirely on logged data, OPE/L is sensitive to environment distribution shifts -- discrepancies between the data-generating environment and that where policies are deployed. proposed distributionally robust OPE/L (DROPE/L) to address this, but the proposal relies on inverse-propensity weighting, whose estimation error and regret will deteriorate if propensities are nonparametrically estimated and whose variance is suboptimal even if not. For standard, non-robust, OPE/L, this is solved by doubly robust (DR) methods, but they do not naturally extend to the more complex DROPE/L, which involves a worst-case expectation. In this paper, we propose the first DR algorithms for DROPE/L with KL-divergence uncertainty sets. For evaluation, we propose Localized Doubly Robust DROPE (LDROPE) and show that it achieves semiparametric efficiency under weak product rates conditions. Thanks to a localization technique, LDROPE only requires fitting a small number of regressions, just like DR methods for standard OPE. For learning, we propose Continuum Doubly Robust DROPL (CDROPL) and show that, under a product rate condition involving a continuum of regressions, it enjoys a fast regret rate of even when unknown propensities are nonparametrically estimated. We empirically validate our algorithms in simulations and further extend our results to general -divergence uncertainty sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 36 citations
- The Benefits of Being Distributional: Small-Loss Bounds for Reinforcement LearningKaiwen Wang, Kevin Zhou, Runzhe Wu, Nathan Kallus et al.NeurIPS 2023 · 31 citations
- Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented ImitationYihong Guo, Yixuan Wang, Yuanyuan Shi, Pan Xu et al.NeurIPS 2024 · 21 citations
- Factored DRO: Factored Distributionally Robust Policies for Contextual BanditsTong Mu, Yash Chandak, Tatsunori B. Hashimoto, Emma BrunskillNeurIPS 2022 · 8 citations
- Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision ProcessesAndrew Bennett, Nathan Kallus, Miruna Oprescu, Wen Sun et al.NeurIPS 2024 · 7 citations
Builds on2
Related papers
- Doubly Robust Distributionally Robust Offline Contextual PricingMin Xu, Xinyi Yin, Yunfan Zhang, Yuxuan Han et al.ICML 2026 · 15 citations
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 60 citations
- Double Reinforcement Learning for Efficient and Robust Off-Policy EvaluationNathan Kallus, Masatoshi UeharaICML 2020 · 6 citations
- Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational DataCheuk Hang Leung, Yiyan Huang, Yijun Li, Qi WuAAAI 2025 · 1 citation
- ORVIT: Near-Optimal Online Distributionally Robust Reinforcement LearningDebamita Ghosh, George K. Atia, Yue WangAAAI 2026 · 2 citations
