Deep Fusion Transformer Network with Weighted Vector-Wise Keypoints Voting for Robust 6D Object Pose Estimation
Jun Zhou, Kai Chen, Linlin Xu, Qi Dou, Jing Qin
Abstract
One critical challenge in 6D object pose estimation from a single RGBD image is efficient integration of two different modalities, i.e., color and depth. In this work, we tackle this problem by a novel Deep Fusion Transformer (DFTr) block that can aggregate cross-modality features for improving pose estimation. Unlike existing fusion methods, the proposed DFTr can better model cross-modality semantic correlation by leveraging their semantic similarity, such that globally enhanced features from different modalities can be better integrated for improved information extraction. Moreover, to further improve robustness and efficiency, we introduce a novel weighted vector-wise voting algorithm that employs a non-iterative global optimization strategy for precise 3D keypoint localization while achieving near real-time inference. Extensive experiments show the effectiveness and strong generalization capability of our proposed 3D keypoint voting algorithm. Results on four widely used benchmarks also demonstrate that our method outperforms the state-of-the-art methods by large margins. Code is available at https://github.com/junzastar/DFTr Voting .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a6d362f-1fb1-4cf3-8955-090bf60d83d8Cited by top-tier papers7
- GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose RefinementLinfang Zheng, Tze Ho Elden Tse, Chen Wang, Yinghan Sun et al.CVPR 2024 · 6 citations
- Vision Foundation Model Enables Generalizable Object Pose EstimationKai Chen, Yiyao Ma, Xingyu Lin, Stephen James et al.NeurIPS 2024 · 5 citations
- Unified Category-Level Object Detection and Pose Estimation from RGB Images Using 3D PrototypesTom Fischer, Xiaojie Zhang, Eddy IlgICCV 2025 · 2 citations
- HiPose: Hierarchical Binary Surface Encoding and Correspondence Pruning for RGB-D 6DoF Object Pose EstimationYongliang Lin, Yongzhi Su, Praveen Nathan, Sandeep Inuganti et al.CVPR 2024
- BiGMINT: Biologically-guided Hierarchical Multimodal Integration for Modeling Multiple Compound Activities in Drug DiscoveryPushpak Pati, Bo Li, Abbas Rayabat Khan, Tomé Albuquerque et al.CVPR 2026
Builds on21
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose EstimationKiru Park, Timothy Patten, Markus VinczeICCV 2019 · 527 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 482 citations
Related papers
- PVN3D: A Deep Point-Wise 3D Keypoints Voting Network for 6DoF Pose EstimationYisheng He, Wei Sun, Haibin Huang, Jianran Liu et al.CVPR 2020
- FFB6D: A Full Flow Bidirectional Fusion Network for 6D Pose EstimationYisheng He, Haibin Huang, Haoqiang Fan, Qifeng Chen et al.CVPR 2021
- Keypoint Fusion for RGB-D Based 3D Hand Pose EstimationXingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang et al.AAAI 2024 · 11 citations
- DCNet: Dense Correspondence Neural Network for 6DoF Object Pose Estimation in Occluded ScenesZhi Chen, Wei Yang, Zhenbo Xu, Xike Xie et al.ACM MM 2020 · 3 citations
- PR-GCN: A Deep Graph Convolutional Network with Point Refinement for 6D Pose EstimationGuangyuan Zhou, Huiqun Wang, Jiaxin Chen, Di HuangICCV 2021 · 45 citations
