Planning, Fast and Slow: Online Reinforcement Learning with Action-Free Offline Data via Multiscale Planners
Chengjie Wu, Hao Hu, Yiqin Yang, Ning Zhang, Chongjie Zhang
摘要
The surge in volumes of video data offers unprecedented opportunities for advancing reinforcement learning (RL). This growth has motivated the development of passive RL, seeking to convert passive observations into actionable insights. This paper explores the prerequisites and mechanisms through which passive data can be utilized to improve online RL. We show that, in identifiable dynamics, where action impact can be distinguished from stochasticity, learning on passive data is statistically beneficial. Building upon the theoretical insights, we propose a novel algorithm named Multiscale State-Centric Planners (MSCP) that leverages two planners at distinct scales to offer guidance across varying levels of abstraction. The algorithm's fast planner targets immediate objectives, while the slow planner focuses on achieving longer-term goals. Notably, the fast planner incorporates pessimistic regularization to address the distributional shift between offline and online data. MSCP effectively handles the practical challenges involving imperfect pretraining and limited dataset coverage. Our empirical evaluations across multiple benchmarks demonstrate that MSCP significantly outperforms existing approaches, underscoring its proficiency in addressing complex, long-horizon tasks through the strategic use of passive data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement LearningHongjoon Ahn, Heewoong Choi, Jisu Han, Taesup MoonNeurIPS 2025 · 被引用 22 次
- RD-HRL: Generating Reliable Sub-Goals for Long-Horizon Sparse-Reward TasksYixiang Shan, Haipeng Liu, Ting Long, Yi ChangICLR 2026
它引用的顶会 Paper35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 被引用 950 次
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
相关 Paper
- Reinforcement Learning from Passive Data via Latent IntentionsDibya Ghosh, Chethan Anand Bhateja, Sergey LevineICML 2023 · 被引用 69 次
- Model-Based Offline Reinforcement Learning with Pessimism-Modulated Dynamics BeliefKaiyang Guo, Yunfeng Shao, Yanhui GengNeurIPS 2022 · 被引用 39 次
- From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-trainingJinwen Wang, Youfang Lin, Xiaobo Hu, Siyu Yang 等ACM MM 2025 · 被引用 1 次
- Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement LearningShentao Yang, Yihao Feng, Shujian Zhang, Mingyuan ZhouICML 2022 · 被引用 14 次
- Progressor: A Perceptually Guided Reward Estimator with Self-Supervised Online RefinementTewodros W. Ayalew, Xiao Zhang, Kevin Yuanbo Wu, Tianchong Jiang 等ICCV 2025 · 被引用 13 次
