A Cross-view Fusion Framework for Robust 6-DoF Grasp Pose Estimation
Kangjian Zhu, Haobo Jiang, Jianjun Qian, Jin Xie
Abstract
In this paper, we propose a cross-view fusion framework that enhances the robustness of 6-DoF grasp pose estimation in corner views. Our framework alleviates occlusion by incorporating an auxiliary view and avoids the time-consuming, task-agnostic multi-view reconstruction through a post-fusion strategy. To enhance cross-view fusion, we propose a self-supervised contrastive learning strategy that leverages cross-view associations to regularize point cloud features. In brief, a cross-view point pair is considered a match if the two points correspond to the same 3D location, and a non-match if they represent distinct grasp directions. The learning strategy significantly enhances the spatial consistency and direction distinctiveness of point features, thereby facilitating cross-view fusion and improving estimation robustness. Furthermore, we propose a cross-view-aligned cylinder integration module to fuse grasp-relevant geometry into a comprehensive representation. Specifically, the module first aligns the cross-view points and features according to their similarity to enhance the robustness against noise. Subsequently, these points are registered into the cylindrical coordinate frame, emphasizing the rotation-symmetric geometry which is important for grasping. Finally, local self-attention and seed cross-attention layers are alternately employed, respectively enabling interactions within single views and across views, which supports fine-grained representation of grasp-relevant geometry. Our framework achieves strong performance on the GraspNet-1Billion benchmark and in real-world applications. Code is available at https://github.com/KJZhuAutomatic/Cross-view-Grasp.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbfd14dc-3283-4424-b8a7-abdf03f7a140Builds on17
- 6-DOF GraspNet: Variational Grasp Generation for Object ManipulationArsalan Mousavian, Clemens Eppner, Dieter FoxICCV 2019 · 673 citations
- Graspness Discovery in Clutters for Fast and Accurate Grasp DetectionChenxi Wang, Haoshu Fang, Minghao Gou, Hongjie Fang et al.ICCV 2021 · 177 citations
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong et al.NeurIPS 2025 · 65 citations
- Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud RegistrationHaobo Jiang, Yaqi Shen, Jin Xie, Jun Li et al.ICCV 2021 · 52 citations
- SE(3) Diffusion Model-based Point Cloud Registration for Robust 6D Object Pose EstimationHaobo Jiang, Mathieu Salzmann, Zheng Dang, Jin Xie et al.NeurIPS 2023 · 51 citations
Related papers
- Completing 3D Partial Assemblies with View-Consistent 2D-3D CorrespondenceWeihao Wang, Yu Lan, Mingyu You, Bin HeICCV 2025 · 1 citation
- Mining Multi-View Information: A Strong Self-Supervised Framework for Depth-based 3D Hand Pose and Mesh EstimationPengfei Ren, Haifeng Sun, Jiachang Hao, Jingyu Wang et al.CVPR 2022 · 23 citations
- Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point CloudsXiaolong Li, Yijia Weng, Li Yi, Leonidas J. Guibas et al.NeurIPS 2021 · 61 citations
- Contrastive Predictive Autoencoders for Dynamic Point Cloud Self-Supervised LearningXiaoxiao Sheng, Zhiqiang Shen, Gang XiaoAAAI 2023 · 14 citations
- GraphGrasp: Lightweight and Efficient Graph-Guided 6-DoF Robotic Grasp Pose Estimation NetworkSheng Yu, Di-Hua Zhai, Yuanqing XiaAAAI 2026
