Scaling Marginalized Importance Sampling to High-Dimensional State-Spaces via State Abstraction
Brahma S. Pavse, Josiah P. Hanna
Abstract
We consider the problem of off-policy evaluation (OPE) in reinforcement learning (RL), where the goal is to estimate the performance of an evaluation policy, πe, using a fixed dataset, D, collected by one or more policies that may be different from πe. Current OPE algorithms may produce poor OPE estimates under policy distribution shift i.e., when the probability of a particular stateaction pair occurring under πe is very different from the probability of that same pair occurring in D (Voloshin et al. 2021; Fu et al. 2021) . In this work, we propose to improve the accuracy of OPE estimators by projecting the high-dimensional state-space into a low-dimensional state-space using concepts from the state abstraction literature. Specifically, we consider marginalized importance sampling (MIS) OPE algorithms which compute state-action distribution correction ratios to produce their OPE estimate. In the original ground statespace, these ratios may have high variance which may lead to high variance OPE. However, we prove that in the lower-dimensional abstract state-space the ratios can have lower variance resulting in lower variance OPE. We then highlight the challenges that arise when estimating the abstract ratios from data, identify sufficient conditions to overcome these issues, and present a minimax optimization problem whose solution yields these abstract ratios. Finally, our empirical evaluation on difficult, high-dimensional state-space OPE tasks shows that the abstract ratios can make MIS OPE estimators achieve lower mean-squared error and more robust to hyperparameter tuning than the ground ratios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- State-Action Similarity-Based Representations for Off-Policy EvaluationBrahma S. Pavse, Josiah HannaNeurIPS 2023 · 5 citations
- Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy EvaluationShreyas Chaudhari, Ameet Deshpande, Bruno C. da Silva, Philip S. ThomasNeurIPS 2024 · 4 citations
- Stable Offline Value Function Learning with Bisimulation-based RepresentationsBrahma S. Pavse, Yudong Chen, Qiaomin Xie, Josiah P. HannaICML 2025
Builds on11
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 171 citations
- Off-Policy Evaluation via the Regularized LagrangianMengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li et al.NeurIPS 2020 · 125 citations
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker et al.ICLR 2021 · 112 citations
Related papers
- A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor RepresentationScott Fujimoto, David Meger, Doina PrecupICML 2021 · 17 citations
- Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL PoliciesHaanvid Lee, Tri Wahyu Guntara, Jongmin Lee, Yung-Kyun Noh et al.ICLR 2024 · 3 citations
- Variance-Aware Off-Policy Evaluation with Linear Function ApproximationYifei Min, Tianhao Wang, Dongruo Zhou, Quanquan GuNeurIPS 2021 · 43 citations
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang et al.ICML 2025
- Counterfactual-Augmented Importance Sampling for Semi-Offline Policy EvaluationShengpu Tang, Jenna WiensNeurIPS 2023 · 8 citations
