Leveraging Global Stereo Consistency for Category-Level Shape and 6D Pose Estimation from Stereo Images
Junning Qiu, Minglei Lu, Fei Wang, Yu Guo, Yonggen Ling
Abstract
Stereo-based category-level shape and 6D pose estimation methods have the potential to generalize to a wider range of materials than RGBD methods, which often suffer from depth measurement errors. However, without explicit depth from two views, parameters to be estimated can become inherently entangled, negatively impacting performance. To address this, we propose a method that leverages global stereo consistency to constrain optimization directions and mitigate parameter entanglement. We first estimate an intra-category occupancy field to represent a unified shape across views, ensuring consistency and preventing shape ambiguity. Through a divide-and-conquer approach within global shape fitting, we fit this shape to stereo images to obtain the pose, iteratively rendering normalized depth maps and exchanging information across views. This approach improves convergence toward the correct pose and scale. We validated our method on both depth-friendly and depth-challenging materials using our S-RGBD dataset and the TOD benchmark. Our method surpasses RGBD methods on challenging objects and performs comparably on depth-friendly ones. Ablation studies confirm the effectiveness of each component.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08d0763d-d82e-4680-8d23-811fd08901afBuilds on21
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera et al.NeurIPS 2023 · 371 citations
- SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose EstimationKai Chen, Qi DouICCV 2021 · 183 citations
- CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowPhilippe Weinzaepfel, Thomas Lucas, Vincent Leroy, Yohann Cabon et al.ICCV 2023 · 181 citations
Related papers
- Universal Features Guided Zero-Shot Category-Level Object Pose EstimationWentian Qu, Chenyu Meng, Heng Li, Jian Cheng et al.AAAI 2025
- SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose EstimationYan Di, Fabian Manhardt, Gu Wang, Xiangyang Ji et al.ICCV 2021 · 163 citations
- Learning Canonical Shape Space for Category-Level 6D Object Pose and Size EstimationDengsheng Chen, Jun Li, Zheng Wang, Kai XuCVPR 2020
- UniPR: Unified Object-level Real-to-Sim Perception and Reconstruction from a Single Stereo PairChuanrui Zhang, Yingshuang Zou, ZhengXian Wu, Yonggen Ling et al.CVPR 2026 · 1 citation
- MonSter: Marry Monodepth to Stereo Unleashes PowerJunda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang et al.CVPR 2025
