TD-MPC2: Scalable, Robust World Models for Continuous Control
Nicklas Hansen, Hao Su, Xiaolong Wang
Abstract
are designed for single-task learning and rely on per-task hyperparameters, with no principled method for selecting those hyperparameters (Zhang et al., 2021 ). An algorithm that can consume large multitask datasets will invariably need to be robust to variation between different tasks (e.g., action space dimensionality, difficulty of exploration, and reward distribution). In this work, we present TD-MPC2: a significant step towards achieving this goal. TD-MPC2 is a model-based RL algorithm designed for learning generalist world models on large uncurated datasets composed of multiple task domains, embodiments, and action spaces, with data sourced from behavior policies that cover a wide range of skill levels, and without the need for hyperparameter-tuning. Our algorithm, which builds upon TD-MPC (Hansen et al., 2022) , performs local trajectory optimization in the latent space of a learned implicit (decoder-free) world model. While the TD-MPC family of algorithms has demonstrated strong empirical performance in prior work (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3949859-482a-4b93-a9ef-56d203fd75feCited by top-tier papers161
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningMoo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin et al.ICLR 2026 · 325 citations
- iVideoGPT: Interactive VideoGPTs are Scalable World ModelsJialong Wu, Shaofeng Yin, Ningya Feng, Xu He et al.NeurIPS 2024 · 177 citations
- Bigger, Regularized, Optimistic: scaling for compute and sample efficient continuous controlMichal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Milos et al.NeurIPS 2024 · 119 citations
- Stop Regressing: Training Value Functions via Classification for Scalable Deep RLJesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taïga et al.ICML 2024 · 118 citations
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics ModelsUladzislau Sobal, Wancong Zhang, Kyunghyun Cho, Randall Balestriero et al.NeurIPS 2025 · 109 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
Related papers
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 388 citations
- Self-Supervised Reinforcement Learning that Transfers using Random FeaturesBoyuan Chen, Chuning Zhu, Pulkit Agrawal, Kaiqing Zhang et al.NeurIPS 2023 · 16 citations
- GenRL: Multimodal-foundation world models for generalization in embodied agentsPietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Aaron C. Courville et al.NeurIPS 2024 · 37 citations
- Dream-MPC: Gradient-Based Model Predictive Control with Latent ImaginationJonathan Spieler, Sven BehnkeICML 2026
- PWM: Policy Learning with Multi-Task World ModelsIgnat Georgiev, Varun Giridhar, Nicklas Hansen, Animesh GargICLR 2025
