Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion
Lunjun Zhang, Yuwen Xiong, Ze Yang, Sergio Casas, Rui Hu, Raquel Urtasun
摘要
Learning world models can teach an agent how the world works in an unsupervised manner. Even though it can be viewed as a special case of sequence modeling, progress for scaling world models on robotic applications such as autonomous driving has been somewhat less rapid than scaling language models with Generative Pre-trained Transformers (GPT). We identify two reasons as major bottlenecks: dealing with complex and unstructured observation space, and having a scalable generative model. Consequently, we propose Copilot4D, a novel world modeling approach that first tokenizes sensor observations with VQVAE, then predicts the future via discrete diffusion. To efficiently decode and denoise tokens in parallel, we recast Masked Generative Image Transformer as discrete diffusion and enhance it with a few simple changes, resulting in notable improvement. When applied to learning world models on point cloud observations, Copilot4D reduces prior SOTA Chamfer distance by more than 65% for 1s prediction, and more than 50% for 3s prediction, across NuScenes, KITTI Odometry, and Argov-erse2 datasets. Our results demonstrate that discrete diffusion on tokenized agent experience can unlock the power of GPT-like unsupervised learning for robotics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Vid2World: Crafting Video Diffusion Models to Interactive World ModelsSiqiao Huang, Jialong Wu, Qixing Zhou, Shangchen Miao 等ICLR 2026 · 被引用 68 次
- Visual Point Cloud Forecasting Enables Scalable Autonomous DrivingZetong Yang, Li Chen, Yanan Sun, Hongyang LiCVPR 2024 · 被引用 40 次
- DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed ImagesXiaoxue Chen, Ziyi Xiong, Yuantao Chen, Gen Li 等CVPR 2026 · 被引用 24 次
- Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal ConsistencyXiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu 等NeurIPS 2025 · 被引用 24 次
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilityYu Yang, Alan Liang, Jianbiao Mei, Yukai Ma 等NeurIPS 2025 · 被引用 22 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Trajeglish: Traffic Modeling as Next-Token PredictionJonah Philion, Xue Bin Peng, Sanja FidlerICLR 2024 · 被引用 61 次
- DriveGPT: Scaling Autoregressive Behavior Models for DrivingXin Huang, Eric M. Wolff, Paul Vernaza, Tung Phan-Minh 等ICML 2025
- GWM: Towards Scalable Gaussian World Models for Robotic ManipulationGuanxing Lu, Baoxiong Jia, Puhao Li, Yixin Chen 等ICCV 2025 · 被引用 1 次
- -World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene ForecastingZhimin Liao, Ping Wei, Ruijie Zhang, Shuaijia Chen 等ICCV 2025 · 被引用 1 次
- Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World ModelXiaodong Wang, Zhirong Wu, Peixi PengAAAI 2026
