Batch Stationary Distribution Estimation
Junfeng Wen, Bo Dai, Lihong Li, Dale Schuurmans
Abstract
We consider the problem of approximating the stationary distribution of an ergodic Markov chain given a set of sampled transitions. Classical simulation-based approaches assume access to the underlying process so that trajectories of sufficient length can be gathered to approximate stationary sampling. Instead, we consider an alternative setting where a fixed set of transitions has been collected beforehand, by a separate, possibly unknown procedure. The goal is still to estimate properties of the stationary distribution, but without additional access to the underlying system. We propose a consistent estimator that is based on recovering a correction ratio function over the given data. In particular, we develop a variational power method (VPM) that provides provably consistent estimates under general conditions. In addition to unifying a number of existing approaches from different subfields, we also find that VPM yields significantly better estimates across a range of problems, including queueing, stochastic differential equations, post-processing MCMC, and off-policy evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6762aa1e-b3f3-4cd8-9db7-37c9e4b3638fCited by top-tier papers14
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker et al.ICLR 2021 · 112 citations
- Provably Good Batch Off-Policy Reinforcement Learning Without Great ExplorationYao Liu, Adith Swaminathan, Alekh Agarwal, Emma BrunskillNeurIPS 2020 · 97 citations
- Autoregressive Dynamics Models for Offline Policy Evaluation and OptimizationMichael R. Zhang, Thomas Paine, Ofir Nachum, Cosmin Paduraru et al.ICLR 2021 · 52 citations
- Continual Learning In Environments With Polynomial Mixing TimesMatthew Riemer, Sharath Chandra Raparthy, Ignacio Cases, Gopeshh Subbaraj et al.NeurIPS 2022 · 18 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
Related papers
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
- Infinite-horizon Off-Policy Policy Evaluation with Multiple Behavior PoliciesXinyun Chen, Lu Wang, Yizhe Hang, Heng Ge et al.ICLR 2020 · 5 citations
- Stein Π-Importance SamplingCongye Wang, Wilson Ye Chen, Heishiro Kanagawa, Chris J. OatesNeurIPS 2023
- Zero-Shot Off-Policy LearningArip Asadulaev, Maksim Bobrin, Salem Lahlou, Dmitry V. Dylov et al.ICML 2026 · 1 citation
- Rank-One Modified Value IterationArman Sharifi Kolarijani, Tolga Ok, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi KolarijaniICML 2025
