Variational Latent Branching Model for Off-Policy Evaluation
Qitong Gao, Ge Gao, Min Chi, Miroslav Pajic
Abstract
Model-based methods have recently shown great potential for off-policy evaluation (OPE); offline trajectories induced by behavioral policies are fitted to transitions of Markov decision processes (MDPs), which are used to rollout simulated trajectories and estimate the performance of policies. Model-based OPE methods face two key challenges. First, as offline trajectories are usually fixed, they tend to cover limited state and action space. Second, the performance of model-based methods can be sensitive to the initialization of their parameters. In this work, we propose the variational latent branching model (VLBM) to learn the transition function of MDPs by formulating the environmental dynamics as a compact latent space, from which the next states and rewards are then sampled. Specifically, VLBM leverages and extends the variational inference framework with the recurrent state alignment (RSA), which is designed to capture as much information underlying the limited training data, by smoothing out the information flow between the variational (encoding) and generative (decoding) part of VLBM. Moreover, we also introduce the branching architecture to improve the model's robustness against randomly initialized model weights. The effectiveness of the VLBM is evaluated on the deep OPE (DOPE) benchmark, from which the training trajectories are designed to result in varied coverage of the state-action space. We show that the VLBM outperforms existing state-of-the-art OPE methods in general.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 223d8d4a-e8c9-4ea7-a72e-02ec94c1e4e9Cited by top-tier papers5
- Off-Policy Evaluation for Human FeedbackQitong Gao, Ge Gao, Juncheng Dong, Vahid Tarokh et al.NeurIPS 2023 · 13 citations
- OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple EstimatorsAllen Nie, Yash Chandak, Christina J. Yuan, Anirudhan Badrinath et al.NeurIPS 2024 · 7 citations
- On Trajectory Augmentations for Off-Policy EvaluationGe Gao, Qitong Gao, Xi Yang, Song Ju et al.ICLR 2024 · 5 citations
- Off-Policy Selection for Initiating Human-Centric Experimental DesignGe Gao, Xi Yang, Qitong Gao, Song Ju et al.NeurIPS 2024 · 1 citation
- Get a Head Start: On-Demand Pedagogical Policy Selection in Intelligent TutoringGe Gao, Xi Yang, Min ChiAAAI 2024 · 1 citation
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran et al.NeurIPS 2021 · 549 citations
Related papers
- Offline Transition Modeling via Contrastive Energy LearningRuifeng Chen, Chengxing Jia, Zefang Huang, Tian-Shuo Liu et al.ICML 2024 · 4 citations
- ADM-v2: Pursuing Full-Horizon Roll-out in Dynamics Models for Offline Policy Learning and EvaluationHaoxin Lin, Siyuan Xiao, Yi-Chen Li, Zhilong Zhang et al.ICLR 2026
- Autoregressive Dynamics Models for Offline Policy Evaluation and OptimizationMichael R. Zhang, Thomas Paine, Ofir Nachum, Cosmin Paduraru et al.ICLR 2021 · 52 citations
- STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy EvaluationHossein Goli, Michael Gimelfarb, Nathan de Lara, Haruki Nishimura et al.NeurIPS 2025 · 3 citations
- Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy EvaluationShreyas Chaudhari, Ameet Deshpande, Bruno C. da Silva, Philip S. ThomasNeurIPS 2024 · 4 citations
