Off-Policy Evaluation and Learning for the Future under Non-Stationarity
Tatsuhiro Shimizu, Kazuki Kawamura, Takanori Muroi, Yusuke Narita, Kei Tateno, Takuma Udagawa, Yuta Saito
Abstract
We study the novel problem of future off-policy evaluation (F-OPE) and learning (F-OPL) for estimating and optimizing the future value of policies in non-stationary environments, where distributions vary over time. In e-commerce recommendations, for instance, our goal is often to estimate and optimize the policy value for the upcoming month using data collected by an old policy in the previous month. A critical challenge is that data related to the future environment is not observed in the historical data. Existing methods assume stationarity or depend on restrictive reward-modeling assumptions, leading to significant bias. To address these limitations, we propose a novel estimator named Off-Policy Estimator for the Future Value (OPFV), designed for accurately estimating policy values at any future time point. The key feature of OPFV is its ability to leverage the useful structure within time-series data. While future data might not be present in the historical log, we can leverage, for example, seasonal, weekly, or holiday effects that are consistent in both the historical and future data. Our estimator is the first to exploit these time-related structures via a new type of importance weighting, enabling effective F-OPE. Theoretical analysis identifies the conditions under which OPFV becomes low-bias. In addition, we extend our estimator to develop a new policy-gradient method to proactively learn a good future policy using only historical data. Empirical results show that our methods substantially outperform existing methods in estimating and optimizing the future policy value under non-stationarity for various experimental setups.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White et al.ICML 2020 · 72 citations
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 62 citations
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 60 citations
Related papers
- Off-Policy Evaluation for Action-Dependent Non-stationary EnvironmentsYash Chandak, Shiv Shankar, Nathaniel D. Bastian, Bruno C. da Silva et al.NeurIPS 2022 · 7 citations
- Future-Dependent Value-Based Off-Policy Evaluation in POMDPsMasatoshi Uehara, Haruka Kiyohara, Andrew Bennett, Victor Chernozhukov et al.NeurIPS 2023 · 31 citations
- Forecasting in Offline Reinforcement Learning for Non-stationary EnvironmentsSuzan Ece Ada, Georg Martius, Emre Ugur, Erhan OztopNeurIPS 2025
- Distributional Offline Policy Evaluation with Predictive Error GuaranteesRunzhe Wu, Masatoshi Uehara, Wen SunICML 2023 · 19 citations
- Off-Policy Learning with Limited SupplyKoichi Tanaka, Ren Kishimoto, Bushun Kawagishi, Yusuke Narita et al.WWW 2026
