Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation
Chongyang Xu, Haipeng Li, Shen Cheng, Haoqiang Fan, Ziliang Feng, Shuaicheng Liu
Abstract
Bimanual manipulation requires policies that can reason about 3D geometry, anticipate how it evolves under action, and generate smooth, coordinated motions. However, existing methods typically rely on 2D features with limited spatial awareness, or require explicit point clouds that are difficult to obtain reliably in real-world settings. At the same time, recent 3D geometric foundation models show that accurate and diverse 3D structure can be reconstructed directly from RGB images in a fast and robust manner. We leverage this opportunity and propose a framework that builds bimanual manipulation directly on a pre-trained 3D geometric foundation model. Our policy fuses geometry-aware latents, 2D semantic features, and proprioception into a unified state representation, and uses diffusion model to jointly predict a future action chunk and a future 3D latent that decodes into a dense pointmap. By explicitly predicting how the 3D scene will evolve together with the action sequence, the policy gains strong spatial understanding and predictive capability using only RGB observations. We evaluate our method both in simulation on the RoboTwin benchmark and in real-world robot executions. Our approach consistently outperforms 2D-based and point-cloud-based baselines, achieving state-of-the-art performance in manipulation success, inter-arm coordination, and 3D spatial prediction accuracy. Code is available at https://github.com/Chongyang-99/GAP.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e78b4b44-b57c-4b0f-a61c-4d19ad379ca6Cited by top-tier papers2
- EPS3D: End-to-End Feed-Forward 3D Panoptic SegmentationRunsong Zhu, Jiaxin GUO, Xiaoyang Guo, Zhengzhe Liu et al.ICML 2026 · 3 citations
- DMAligner: Enhancing Image Alignment via Diffusion Model Based View SynthesisXinglong Luo, Ao Luo, Zhengning Wang, Yueqi Yang et al.CVPR 2026 · 1 citation
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai et al.NeurIPS 2023 · 742 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu et al.ICLR 2026 · 335 citations
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang et al.ICLR 2026 · 318 citations
Related papers
- Diffusion-Based Imaginative Coordination for Bimanual ManipulationHuilin Xu, Jian Ding, Jiakun Xu, Ruixiang Wang et al.ICCV 2025
- CUBic: Coordinated Unified Bimanual Perception and Control FrameworkXingyu Wang, Pengxiang Ding, Jingkai Xu, Donglin Wang et al.CVPR 2026 · 1 citation
- GeoMoLa: Geometry-Aware Motion Latents for Learning Robust Manipulation PoliciesYunchao Zhang, Yijia Weng, Ruizhe Liu, Ming Hu et al.ICML 2026
- RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationSongming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan et al.ICLR 2025
- G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object ManipulationTianxing Chen, Yao Mu, Zhixuan Liang, Zanxin Chen et al.CVPR 2025
