Predicting What Matters: Robust Generalist Robot Policy Learning via Future Semantic Mask
Yunfan Lou, Xiaowei Chi, Xiaojie Zhang, Zezhong Qian, Chengxuan Li, Rongyu Zhang, yaoxu lyu, Guoyu Song, Chuyao Fu, Haoxuan Xu, Pengwei Wang, Shanghang Zhang
Abstract
World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often focus on high-fidelity RGB video prediction, this can result in overfitting to irrelevant factors, such as dynamic backgrounds and illumination changes. These distractions reduce the model's ability to generalize, ultimately leading to unreliable and fragile control policies. To address this, we introduce the Mask World Model (MWM), which leverages video diffusion architectures to predict the evolution of semantic masks instead of pixels. This shift imposes a geometric information bottleneck, forcing the model to capture essential physical dynamics and contact relations while filtering out visual noise. We seamlessly integrate this mask dynamics backbone with a diffusion-based policy head to enable robust end-to-end control. Extensive evaluations demonstrate the superiority of MWM on the LIBERO and RLBench simulation benchmarks, significantly outperforming the state-of-the-art RGB-based world models. Furthermore, real-world experiments and robustness evaluation (via random token pruning) reveal that MWM exhibits superior generalization capabilities and robust resilience to texture information loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8524e4ab-d459-4f29-9985-2e1e0678bc75Builds on12
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationYue Liao, Pengfei Zhou, Siyuan Huang, Donglin Yang et al.ICLR 2026 · 136 citations
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel et al.ICLR 2023 · 87 citations
- Denoised MDPs: Learning World Models Better Than the World ItselfTongzhou Wang, Simon S. Du, Antonio Torralba, Phillip Isola et al.ICML 2022 · 63 citations
Related papers
- DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic ManipulationKaizhao Zhang, Tian Niu, Tianyu Liu, Chenen Guo et al.CVPR 2026
- Scaling Real-World Robot Policy Evaluation via Discrete Diffusion World ModelYaxuan Li, Junjie Wen, Zhongyi Zhou, Yefei Chen et al.ICML 2026 · 5 citations
- MaskGWM: A Generalizable Driving World Model with Video Mask ReconstructionJingcheng Ni, Yuxin Guo, Yichen Liu, Rui Chen et al.CVPR 2025
- InternVideo-Next: Towards World-Understanding Video ModelsChenting Wang, Yuhan Zhu, Yicheng Xu, Jiange Yang et al.CVPR 2026
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsYucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen et al.ICML 2025
