Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation
Chongyang Xu, Haipeng Li, Shen Cheng, Haoqiang Fan, Ziliang Feng, Shuaicheng Liu
摘要
Bimanual manipulation requires policies that can reason about 3D geometry, anticipate how it evolves under action, and generate smooth, coordinated motions. However, existing methods typically rely on 2D features with limited spatial awareness, or require explicit point clouds that are difficult to obtain reliably in real-world settings. At the same time, recent 3D geometric foundation models show that accurate and diverse 3D structure can be reconstructed directly from RGB images in a fast and robust manner. We leverage this opportunity and propose a framework that builds bimanual manipulation directly on a pre-trained 3D geometric foundation model. Our policy fuses geometry-aware latents, 2D semantic features, and proprioception into a unified state representation, and uses diffusion model to jointly predict a future action chunk and a future 3D latent that decodes into a dense pointmap. By explicitly predicting how the 3D scene will evolve together with the action sequence, the policy gains strong spatial understanding and predictive capability using only RGB observations. We evaluate our method both in simulation on the RoboTwin benchmark and in real-world robot executions. Our approach consistently outperforms 2D-based and point-cloud-based baselines, achieving state-of-the-art performance in manipulation success, inter-arm coordination, and 3D spatial prediction accuracy. Code is available at https://github.com/Chongyang-99/GAP.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- EPS3D: End-to-End Feed-Forward 3D Panoptic SegmentationRunsong Zhu, Jiaxin GUO, Xiaoyang Guo, Zhengzhe Liu 等ICML 2026 · 被引用 3 次
- DMAligner: Enhancing Image Alignment via Diffusion Model Based View SynthesisXinglong Luo, Ao Luo, Zhengning Wang, Yueqi Yang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu 等ICLR 2026 · 被引用 335 次
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang 等ICLR 2026 · 被引用 318 次
相关 Paper
- Diffusion-Based Imaginative Coordination for Bimanual ManipulationHuilin Xu, Jian Ding, Jiakun Xu, Ruixiang Wang 等ICCV 2025
- CUBic: Coordinated Unified Bimanual Perception and Control FrameworkXingyu Wang, Pengxiang Ding, Jingkai Xu, Donglin Wang 等CVPR 2026 · 被引用 1 次
- GeoMoLa: Geometry-Aware Motion Latents for Learning Robust Manipulation PoliciesYunchao Zhang, Yijia Weng, Ruizhe Liu, Ming Hu 等ICML 2026
- RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationSongming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan 等ICLR 2025
- G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object ManipulationTianxing Chen, Yao Mu, Zhixuan Liang, Zanxin Chen 等CVPR 2025
