DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic Manipulation
Kaizhao Zhang, Tian Niu, Tianyu Liu, Chenen Guo, Zijun Xu, Qingda Hu, Wenchao Ding
Abstract
Robotic manipulation from visual observations remains challenging due to the lack of 3D consistent representations that can generalize across diverse viewpoints and sensor configurations. Existing methods, primarily based on masked autoencoders or neural scene representations, struggle to capture robust view correspondences due to a lack of global 3D consistency or the requirement for scene specific optimization. Crucially, while multi-view diffusion models have recently shown tremendous success in 3D aware generative synthesis, their powerful representations offer a promising direction for achieving viewpoint robust visuomotor control. In this paper, we introduce DiffuView, a novel framework that learns unified 3D aware representations through multi-view diffusion pretraining and deploys them for imitation learning. Specifically, DiffuView models the conditional generation of target views given source observations within a diffusion framework, enabling the network to implicitly recover scene geometry and enforce view consistency. The pretrained diffusion network is then utilized as a powerful visual backbone for an action policy, allowing robust control under varying viewpoints and visual conditions. We evaluate DiffuView on two challenging benchmarks, MetaWorld and Libero. Extensive experiments in both simulation and realworld scenarios demonstrate that DiffuView achieves superior generalization, improving success rates under viewpoint shifts by nearly 20% compared with existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- CAT3D: Create Anything in 3D with Multi-View Diffusion ModelsRuiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee et al.NeurIPS 2024 · 490 citations
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma et al.NeurIPS 2023 · 336 citations
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen et al.ICLR 2024 · 309 citations
Related papers
- Predicting What Matters: Robust Generalist Robot Policy Learning via Future Semantic MaskYunfan Lou, Xiaowei Chi, Xiaojie Zhang, Zezhong Qian et al.ICML 2026
- 4D Visual Pre-Training for Robot LearningChengkai Hou, Yanjie Ze, Yankai Fu, Zeyu Gao et al.ICCV 2025 · 1 citation
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot LearningQiwei Liang, Boyang Cai, Minghao Lai, Sitong Zhuang et al.CVPR 2026 · 6 citations
- Masked Generative Policy for Robotic ControlLipeng Zhuang, Shiyu Fan, Florent P. Audonnet, Yingdong Ru et al.ICLR 2026 · 1 citation
