PWM: Policy Learning with Multi-Task World Models
Ignat Georgiev, Varun Giridhar, Nicklas Hansen, Animesh Garg
Abstract
Reinforcement Learning (RL) has made significant strides in complex tasks but struggles in multi-task settings with different embodiments. World model methods offer scalability by learning a simulation of the environment but often rely on inefficient gradient-free optimization methods for policy extraction. In contrast, gradient-based methods exhibit lower variance but fail to handle discontinuities. Our work reveals that well-regularized world models can generate smoother optimization landscapes than the actual dynamics, facilitating more effective first-order optimization. We introduce Policy learning with multi-task World Models (PWM), a novel model-based RL algorithm for continuous control. Initially, the world model is pre-trained on offline data, and then policies are extracted from it using first-order optimization in less than 10 minutes per task. PWM effectively solves tasks with up to 152 action dimensions and outperforms methods that use groundtruth dynamics. Additionally, PWM scales to an 80-task setting, achieving up to 27% higher rewards than existing baselines without relying on costly online planning. Visualizations and code are available at imgeorgiev.com/pwm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a2c7637-833b-4519-b1fd-2d7e950684e6Cited by top-tier papers2
- DMWM: Dual-Mind World Model with Long-Term ImaginationLingyi Wang, Rashed Shelim, Walid Saad, Naren RamakrishnanNeurIPS 2025 · 15 citations
- Dream-MPC: Gradient-Based Model Predictive Control with Latent ImaginationJonathan Spieler, Sven BehnkeICML 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 388 citations
Related papers
- TaskLoom: Weaving Knowledge Across Tasks in World ModelsQingzhang Zeng, Peixi Peng, hang li, Luntong Li et al.ICML 2026
- Learning Massively Multitask World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2026 · 14 citations
- World Models as Reference Trajectories for Rapid Motor AdaptationCarlos Stein Brito, Daniel McNameeNeurIPS 2025 · 2 citations
- On the Feasibility of Cross-Task Transfer with Model-Based Reinforcement LearningYifan Xu, Nicklas Hansen, Zirui Wang, Yung-Chieh Chan et al.ICLR 2023 · 3 citations
- Making Better Decision by Directly Planning in Continuous ControlJinhua Zhu, Yue Wang, Lijun Wu, Tao Qin et al.ICLR 2023
