Variational Latent Branching Model for Off-Policy Evaluation
Qitong Gao, Ge Gao, Min Chi, Miroslav Pajic
摘要
Model-based methods have recently shown great potential for off-policy evaluation (OPE); offline trajectories induced by behavioral policies are fitted to transitions of Markov decision processes (MDPs), which are used to rollout simulated trajectories and estimate the performance of policies. Model-based OPE methods face two key challenges. First, as offline trajectories are usually fixed, they tend to cover limited state and action space. Second, the performance of model-based methods can be sensitive to the initialization of their parameters. In this work, we propose the variational latent branching model (VLBM) to learn the transition function of MDPs by formulating the environmental dynamics as a compact latent space, from which the next states and rewards are then sampled. Specifically, VLBM leverages and extends the variational inference framework with the recurrent state alignment (RSA), which is designed to capture as much information underlying the limited training data, by smoothing out the information flow between the variational (encoding) and generative (decoding) part of VLBM. Moreover, we also introduce the branching architecture to improve the model's robustness against randomly initialized model weights. The effectiveness of the VLBM is evaluated on the deep OPE (DOPE) benchmark, from which the training trajectories are designed to result in varied coverage of the state-action space. We show that the VLBM outperforms existing state-of-the-art OPE methods in general.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Off-Policy Evaluation for Human FeedbackQitong Gao, Ge Gao, Juncheng Dong, Vahid Tarokh 等NeurIPS 2023 · 被引用 13 次
- OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple EstimatorsAllen Nie, Yash Chandak, Christina J. Yuan, Anirudhan Badrinath 等NeurIPS 2024 · 被引用 7 次
- On Trajectory Augmentations for Off-Policy EvaluationGe Gao, Qitong Gao, Xi Yang, Song Ju 等ICLR 2024 · 被引用 5 次
- Off-Policy Selection for Initiating Human-Centric Experimental DesignGe Gao, Xi Yang, Qitong Gao, Song Ju 等NeurIPS 2024 · 被引用 1 次
- Get a Head Start: On-Demand Pedagogical Policy Selection in Intelligent TutoringGe Gao, Xi Yang, Min ChiAAAI 2024 · 被引用 1 次
它引用的顶会 Paper19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran 等NeurIPS 2021 · 被引用 549 次
相关 Paper
- Offline Transition Modeling via Contrastive Energy LearningRuifeng Chen, Chengxing Jia, Zefang Huang, Tian-Shuo Liu 等ICML 2024 · 被引用 4 次
- ADM-v2: Pursuing Full-Horizon Roll-out in Dynamics Models for Offline Policy Learning and EvaluationHaoxin Lin, Siyuan Xiao, Yi-Chen Li, Zhilong Zhang 等ICLR 2026
- Autoregressive Dynamics Models for Offline Policy Evaluation and OptimizationMichael R. Zhang, Thomas Paine, Ofir Nachum, Cosmin Paduraru 等ICLR 2021 · 被引用 52 次
- STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy EvaluationHossein Goli, Michael Gimelfarb, Nathan de Lara, Haruki Nishimura 等NeurIPS 2025 · 被引用 3 次
- Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy EvaluationShreyas Chaudhari, Ameet Deshpande, Bruno C. da Silva, Philip S. ThomasNeurIPS 2024 · 被引用 4 次
