PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic Manipulation
Jingyi Tian, Le Wang, Sanping Zhou, Sen Wang, Jiayi Li, Haowen Sun, Wei Tang
摘要
Robotic manipulation based on visual observations and natural language instructions is a long-standing challenge in robotics. Yet prevailing approaches model action distribution by adopting explicit or implicit representations, which often struggle to achieve a trade-off between accuracy and efficiency. In response, we propose PDFactor, a novel framework that models action distribution with a hybrid triplane representation. In particular, PDFactor decomposes 3D point cloud into three orthogonal feature planes and leverages a tri-perspective view transformer to produce dense cubic features as a latent diffusion field aligned with observation space representing 6-DoF action probability distribution at an arbitrary location. We employ a small denoising network conceptually as both a parameterized loss function measuring the quality of the learned latent features and an action gradient decoder to sample actions from the latent diffusion field during inference. This design enables our PDFactor to benefit from spatial awareness of explicit representation and arbitrary resolution of implicit representation, rendering it with manipulation accuracy, inference efficiency, and model scalability. Experiments demonstrate that PDFactor outperforms state-of-theart approaches across a diverse range of manipulation tasks in RLBench simulation. Moreover, PDFactor can effectively learn multi-task policies from a limited number of human demonstrations, achieving promising accuracy in a variety of real-world manipulation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal OverheadChaojun Ni, Chen Cheng, Xiaofeng Wang, Zheng Zhu 等CVPR 2026 · 被引用 23 次
- UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human VideosGu Zhang, Qicheng Xu, Haozhe Zhang, Jianhan Ma 等CVPR 2026 · 被引用 23 次
- SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World ModelsSen Wang, Jingyi Tian, Le Wang, Zhimin Liao 等NeurIPS 2025 · 被引用 3 次
- PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud LearningXinxing Yu, Ajian Liu, Sunyuan Qiang, Hui Ma 等CVPR 2026 · 被引用 1 次
- DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic ManipulationKaizhao Zhang, Tian Niu, Tianyu Liu, Chenen Guo 等CVPR 2026
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua 等NeurIPS 2020 · 被引用 1,535 次
相关 Paper
- Action-Geometry Prediction with 3D Geometric Prior for Bimanual ManipulationChongyang Xu, Haipeng Li, Shen Cheng, Haoqiang Fan 等CVPR 2026 · 被引用 10 次
- ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D TrajectoryYing Li, Xiaobao Wei, Xiaowei Chi, Yuming Li 等AAAI 2026
- 4D Visual Pre-Training for Robot LearningChengkai Hou, Yanjie Ze, Yankai Fu, Zeyu Gao 等ICCV 2025 · 被引用 1 次
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot LearningQiwei Liang, Boyang Cai, Minghao Lai, Sitong Zhuang 等CVPR 2026 · 被引用 6 次
- DiffFacto: Controllable Part-Based 3D Point Cloud Generation with Cross DiffusionGeorge Kiyohiro Nakayama, Mikaela Angelina Uy, Jiahui Huang, Shi-Min Hu 等ICCV 2023 · 被引用 46 次
