FFB6D: A Full Flow Bidirectional Fusion Network for 6D Pose Estimation
Yisheng He, Haibin Huang, Haoqiang Fan, Qifeng Chen, Jian Sun
Abstract
In this work, we present FFB6D, a Full Flow Bidirectional fusion network designed for 6D pose estimation from a single RGBD image. Our key insight is that appearance information in the RGB image and geometry information from the depth image are two complementary data sources, and it still remains unknown how to fully leverage them. Towards this end, we propose FFB6D, which learns to combine appearance and geometry information for representation learning as well as output representation selection. Specifically, at the representation learning stage, we build bidirectional fusion modules in the full flow of the two networks, where fusion is applied to each encoding and decoding layer. In this way, the two networks can leverage local and global complementary information from the other one to obtain better representations. Moreover, at the output representation stage, we designed a simple but effective 3D keypoints selection algorithm considering the texture and geometry information of objects, which simplifies keypoint localization for precise pose estimation. Experimental results show that our method outperforms the state-of-the-art by large margins on several benchmarks. Code and video are available at https://github.com/ethnhe/FFB6D.git . Pose Estimation Dense Fusion CNN Encoder CNN Decoder Point Cloud Decoder Point Cloud Encoder (a) The DenseFusion [65] Network. The two networks extract features from different modalities of data separately without any communication, util the final layers of the encoding-decoding architecture. Fusion module Pose Estimation Concatenate CNN Encoder CNN Decoder Point Cloud Decoder Point Cloud Encoder
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86a18484-b1bd-43bf-9d98-1adcbf53dc64Cited by top-tier papers45
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 215 citations
- ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose EstimationYongzhi Su, Mahdi Saleh, Torben Fetzer, Jason R. Rambach et al.CVPR 2022 · 170 citations
- GPV-Pose: Category-level Object Pose Estimation via Geometry-guided Point-wise VotingYan Di, Ruida Zhang, Zhiqiang Lou, Fabian Manhardt et al.CVPR 2022 · 141 citations
- Category-Level 6D Object Pose Estimation in the Wild: A Semi-Supervised Learning Approach and A New DatasetYanjie Ze, Xiaolong WangNeurIPS 2022 · 104 citations
- SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size EstimationHaitao Lin, Zichang Liu, Chilam Cheang, Yanwei Fu et al.CVPR 2022 · 86 citations
Builds on16
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose EstimationKiru Park, Timothy Patten, Markus VinczeICCV 2019 · 527 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 482 citations
- MoreFusion: Multi-object Reasoning for 6D Pose Estimation from Volumetric FusionKentaro Wada, Edgar Sucar, Stephen James, Daniel Lenton et al.CVPR 2020
Related papers
- Deep Fusion Transformer Network with Weighted Vector-Wise Keypoints Voting for Robust 6D Object Pose EstimationJun Zhou, Kai Chen, Linlin Xu, Qi Dou et al.ICCV 2023 · 42 citations
- PVN3D: A Deep Point-Wise 3D Keypoints Voting Network for 6DoF Pose EstimationYisheng He, Wei Sun, Haibin Huang, Jianran Liu et al.CVPR 2020
- Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose EstimationXiaoke Jiang, Donghai Li, Hao Chen, Ye Zheng et al.CVPR 2022 · 54 citations
- PR-GCN: A Deep Graph Convolutional Network with Point Refinement for 6D Pose EstimationGuangyuan Zhou, Huiqun Wang, Jiaxin Chen, Di HuangICCV 2021 · 45 citations
- Keypoint Fusion for RGB-D Based 3D Hand Pose EstimationXingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang et al.AAAI 2024 · 11 citations
