Off-Policy Evaluation and Learning for External Validity under a Covariate Shift
Masatoshi Uehara, Masahiro Kato, Shota Yasui
Abstract
We consider evaluating and training a new policy for the evaluation data by using the historical data obtained from a different policy. The goal of off-policy evaluation (OPE) is to estimate the expected reward of a new policy over the evaluation data, and that of off-policy learning (OPL) is to find a new policy that maximizes the expected reward over the evaluation data. Although the standard OPE and OPL assume the same distribution of covariate between the historical and evaluation data, a covariate shift often exists, i.e., the distribution of the covariate of the historical data is different from that of the evaluation data. In this paper, we derive the efficiency bound of OPE under a covariate shift. Then, we propose doubly robust and efficient estimators for OPE and OPL under a covariate shift by using a nonparametric estimator of the density ratio between the historical and evaluation data distributions. We also discuss other possible estimators and compare their theoretical properties. Finally, we confirm the effectiveness of the proposed estimators through experiments. * Equal contributions. Off-Policy Evaluation and Learning for External Validity under a Covariate Shift A PREPRINT importance weighting using the density ratio between the distributions of the covariates of the historical and evaluation data (Shimodaira, 2000; Reddi et al., 2015) .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75bae32d-b7c2-49e9-b427-f86e908771fdCited by top-tier papers12
- Non-Negative Bregman Divergence Minimization for Deep Direct Density Ratio EstimationMasahiro Kato, Takeshi TeshimaICML 2021 · 53 citations
- Doubly Robust Alignment for Large Language ModelsErhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu et al.NeurIPS 2025 · 14 citations
- Active Adaptive Experimental Design for Treatment Effect Estimation with Covariate ChoiceMasahiro Kato, Akihiro Oga, Wataru Komatsubara, Ryo InokuchiICML 2024 · 12 citations
- Offline Minimax Soft-Q-learning Under Realizability and Partial CoverageMasatoshi Uehara, Nathan Kallus, Jason D. Lee, Wen SunNeurIPS 2023 · 10 citations
- Factored DRO: Factored Distributionally Robust Policies for Contextual BanditsTong Mu, Yash Chandak, Tatsunori B. Hashimoto, Emma BrunskillNeurIPS 2022 · 8 citations
Related papers
- Weighted model estimation for offline model-based reinforcement learningToru Hishinuma, Kei SendaNeurIPS 2021 · 15 citations
- Doubly Robust Distributionally Robust Off-Policy Evaluation and LearningNathan Kallus, Xiaojie Mao, Kaiwen Wang, Zhengyuan ZhouICML 2022 · 39 citations
- Double Reinforcement Learning for Efficient and Robust Off-Policy EvaluationNathan Kallus, Masatoshi UeharaICML 2020 · 6 citations
- Marginal Density Ratio for Off-Policy Evaluation in Contextual BanditsMuhammad Faaiz Taufiq, Arnaud Doucet, Rob Cornish, Jean-Francois TonNeurIPS 2023 · 14 citations
- Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational DataCheuk Hang Leung, Yiyan Huang, Yijun Li, Qi WuAAAI 2025 · 1 citation
