Zero-Shot Off-Policy Learning
Arip Asadulaev, Maksim Bobrin, Salem Lahlou, Dmitry V. Dylov, Fakhri Karray, Martin Takac
Abstract
Off-policy learning methods seek to derive an optimal policy directly from a fixed dataset of prior interactions. This objective presents significant challenges, primarily due to the inherent distributional shift and value function overestimation bias. These issues become even more noticeable in zero-shot reinforcement learning, where an agent trained on reward-free data must adapt to new tasks at test time without additional training. In this work, we address the off-policy problem in a zero-shot setting by discovering a theoretical connection of successor measures to stationary density ratios. Using this insight, our algorithm can infer optimal importance sampling ratios, effectively performing a stationary distribution correction with an optimal policy for any task on the fly. We benchmark our method in motion tracking tasks on SMPL Humanoid, continuous control on ExoRL, and for the long-horizon OGBench tasks. Our technique seamlessly integrates into forward-backward representation frameworks and enables fast-adaptation to new tasks in a training-free regime. More broadly, this work bridges off-policy learning and zero-shot adaptation, offering benefits to both research areas.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6adafe1-c73d-4164-9d74-3a14aaa774b1Builds on7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 72 citations
- BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement LearningYitang Li, Zhengyi Luo, Tonghe Zhang, Cunxi Dai et al.ICLR 2026 · 63 citations
- Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement LearningLiyuan Mao, Haoran Xu, Xianyuan Zhan, Weinan Zhang et al.NeurIPS 2024 · 49 citations
- Zero-Shot Reinforcement Learning from Low Quality DataScott R. Jeen, Tom Bewley, Jonathan M. CullenNeurIPS 2024 · 24 citations
Related papers
- Improving Zero-Shot Offline RL via Behavioral Task SamplingNazim Bendib, Nicolas Perrin-Gilbert, Olivier SigaudICML 2026
- Proto Successor Measure: Representing the Behavior Space of an RL AgentSiddhant Agarwal, Harshit Sikchi, Peter Stone, Amy ZhangICML 2025
- A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor RepresentationScott Fujimoto, David Meger, Doina PrecupICML 2021 · 17 citations
- Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement LearningJingbo Sun, Qichao Zhang, Songjun Tu, Xing Fang et al.CVPR 2026 · 1 citation
- Learning from Sparse Offline Datasets via Conservative Density EstimationZhepeng Cen, Zuxin Liu, Zitong Wang, Yihang Yao et al.ICLR 2024 · 12 citations
