4D Visual Pre-Training for Robot Learning
Chengkai Hou, Yanjie Ze, Yankai Fu, Zeyu Gao, Songbo Hu, Yue Yu, Shanghang Zhang, Huazhe Xu
摘要
General visual representations learned from web-scale datasets for robotics have achieved great success in recent years, enabling data-efficient robot learning on manipulation tasks; yet these pre-trained representations are mostly on 2D images, neglecting the inherent 3D nature of the world. However, due to the scarcity of large-scale 3D data, it is still hard to extract a universal 3D representation from web datasets. Instead, we are seeking a general visual pre-training framework that could improve all 3D representations as an alternative. Our framework, called FVP, is a novel 4D Visual Pre-training framework for realworld robot learning. FVP frames the visual pre-training objective as a next-point-cloud-prediction problem, models the prediction model as a diffusion model, and pre-trains the model on the larger public datasets directly. Across Abstracttwelve real-world manipulation tasks, FVP boosts the average success rate of 3D Diffusion Policy (DP3) for these tasks by 28%. The FVP pre-trained DP3 achieves state-of-the-art performance across imitation learning methods. Moreover, the efficacy of FVP adapts across various point cloud encoders and datasets. Finally, we apply FVP to the RDT-1B, a larger Vision-Language-Action robotic model, enhancing its performance on various robot tasks. Our project page is available at: https://4d-visualpretraining.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot LearningQiwei Liang, Boyang Cai, Minghao Lai, Sitong Zhuang 等CVPR 2026 · 被引用 6 次
- Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic ManipulationYunlong Zhao, Xiaoheng Deng, Yichao Cao, Yi Chen 等CVPR 2026
它引用的顶会 Paper23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- 3D Shape Generation and Completion through Point-Voxel DiffusionLinqi Zhou, Yilun Du, Jiajun WuICCV 2021 · 被引用 681 次
相关 Paper
- Spatial-Temporal Aware Visuomotor Diffusion Policy LearningZhenyang Liu, Yikai Wang, Kuanning Wang, Longfei Liang 等ICCV 2025 · 被引用 11 次
- DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic ManipulationKaizhao Zhang, Tian Niu, Tianyu Liu, Chenen Guo 等CVPR 2026
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationsYucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen 等ICML 2025
- Pre-training Auto-regressive Robotic Models with 4D RepresentationsDantong Niu, Yuvan Sharma, Haoru Xue, Giscard Biamby 等ICML 2025
- 3D-MVP: 3D Multiview Pretraining for ManipulationShengyi Qian, Kaichun Mo, Valts Blukis, David F. Fouhey 等CVPR 2025
