Discrete Codebook World Models for Continuous Control
Aidan Scannell, Mohammadreza Nakhaeinezhadfard, Kalle Kujanpää, Yi Zhao, Kevin Sebastian Luck, Arno Solin, Joni Pajarinen
摘要
In reinforcement learning (RL), world models serve as internal simulators, enabling agents to predict environment dynamics and future outcomes in order to make informed decisions. While previous approaches leveraging discrete latent spaces, such as DreamerV3, have demonstrated strong performance in discrete action settings and visual control tasks, their comparative performance in state-based continuous control remains underexplored. In contrast, methods with continuous latent spaces, such as TD-MPC2, have shown notable success in state-based continuous control benchmarks. In this paper, we demonstrate that modeling discrete latent states has benefits over continuous latent states and that discrete codebook encodings are more effective representations for continuous control, compared to alternative encodings, such as one-hot and label-based encodings. Based on these insights, we introduce DCWM: Discrete Codebook World Model, a self-supervised world model with a discrete and stochastic latent space, where latent states are codes from a codebook. We combine DCWM with decision-time planning to get our model-based RL algorithm, named DC-MPC: Discrete Codebook Model Predictive Control, which performs competitively against recent state-of-the-art algorithms, including TD-MPC2 and DreamerV3, on continuous control benchmarks. See our project website www.aidanscannell.com/dcmpc.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical ImagingIbrahim Ethem Hamamci, Sezgin Er, Suprosanna Shit, Hadrien Reynaud 等NeurIPS 2025 · 被引用 23 次
- DMWM: Dual-Mind World Model with Long-Term ImaginationLingyi Wang, Rashed Shelim, Walid Saad, Naren RamakrishnanNeurIPS 2025 · 被引用 15 次
- Simplicial Embeddings Improve Sample Efficiency in Actor–Critic AgentsJohan Obando-Ceron, Walter Mayor, Samuel Lavoie, Scott Fujimoto 等ICLR 2026 · 被引用 12 次
- Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World ModelDongwon Kim, Gawon Seo, Jinsung Lee, Minsu Cho 等CVPR 2026 · 被引用 6 次
- Efficient Reinforcement Learning by Guiding World Models with Non-Curated DataYi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
相关 Paper
- Temporal Predictive Coding For Model-Based Planning In Latent SpaceTung D. Nguyen, Rui Shu, Tuan Pham, Hung Bui 等ICML 2021 · 被引用 65 次
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 被引用 388 次
- Dream-MPC: Gradient-Based Model Predictive Control with Latent ImaginationJonathan Spieler, Sven BehnkeICML 2026
- SOLD: Slot Object-Centric Latent Dynamics Models for Relational Manipulation Learning from PixelsMalte Mosbach, Jan Niklas Ewertz, Angel Villar-Corrales, Sven BehnkeICML 2025
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot GeneralizationFrank Röder, Jan Benad, Manfred Eppe, Pradeep Kr. BanerjeeNeurIPS 2025 · 被引用 9 次
