COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation
Yuchen Che, JINGTU WU, Hao ZHENG, Asako Kanezaki
Abstract
making the problem particularly ill-posed. A key to solving this task is to establish reliable cross-view correspondences, since the pose can be recovered by aligning geometric structures between query and reference views. Yet, most existing approaches [39,41] construct correspondences via discrete one-to-one assignments (e.g. argmax), which tend to collapse onto a few dominant keypoints, leaving many of the points unused. Moreover, such non-differentiable assignments breaks differentiability and prevents the model from being trained in an unsupervised manner.
We introduce Confidence-aware Optimal Geometric Correspondence (COG), an unsupervised framework for novel object pose estimation from a single reference image. COG addresses the above issues by formulating soft correspondence as an optimal transport (OT) problem, where point-wise confidences are predicted beforehand and explicitly incorporated as target marginals of the transport plan. Compared to OT-based methods [13,55,72] with uniform marginals and only apply confidence post hoc, This formulation yields globally balanced correspondences that naturally suppresses outliers and non-overlapping regions. Given these correspondences, corresponding points are generated via convex combinations, and a weighted SVD solver [63] is used to recover the pose transformation. The entire process forms an end-to-end correspondence finding and pose estimation pipeline, enabling unsupervised optimization of both correspondence and confidence. To mitigate ambiguity in purely geometric matching, we integrate semantic priors denoised from vision foundation models such as DINO [5,48], which softly encourage correspondences between semantically consistent parts. Furthermore, for unsupervised confidence learning, Gaussian RBF-style kernels of geometric and semantic consistency are used to generate pseudo confidence labels, guiding the network to down weight uncertain points without discarding them entirely and to emphasize reliable regions with high confidence. With these designs, COG naturally extends to the unsupervised setting, where neither groundtruth confidence nor pose supervision is available. Experimental results demonstrate that COG achieves performance comparable to leading supervised approaches, while the supervised variant of COG further outperforms them.
Our contributions are summarized as follows: 1. We formulate correspondence finding as an OT problem with confidence as marginals. Compared to OT with uniform marginals, our formulation yields balanced correspondences by suppressing non-overlapping points. 2. We propose an end-to-end pipeline that jointly learns object pose and point validity confidence without supervision from CAD models, poses, or overlap scores. 3. Unsupervised COG achieves performance competitive with state-of-the-art supervised methods, and its supervised variant further outperforms them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc3bbd52-1144-4442-82bb-cea587cedb0eBuilds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Fully Convolutional Geometric FeaturesChristopher B. Choy, Jaesik Park, Vladlen KoltunICCV 2019 · 807 citations
- Geometric Transformer for Fast and Robust Point Cloud RegistrationZheng Qin, Hao Yu, Changjian Wang, Yulan Guo et al.CVPR 2022 · 436 citations
Related papers
- MAPConNet: Self-supervised 3D Pose Transfer with Mesh and Point Contrastive LearningJiaze Sun, Zhixiang Chen, Tae-Kyun KimICCV 2023 · 2 citations
- OrienPose: Orientation-Guided Novel View Synthesis for Single-Image Unseen Object Pose EstimationYating Liu, Zhaoshuai Qi, Yang Zou, Yongnan Yang et al.CVPR 2026
- MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training SupervisionRuicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang et al.CVPR 2025
- Jamais Vu: Exposing the Generalization Gap in Supervised Semantic CorrespondenceOctave Mariotti, Zhipeng Du, Yash Bhalgat, Oisin Mac Aodha et al.NeurIPS 2025 · 8 citations
- SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft SignalsSoyeon Yoon, Chang Wook Seo, Hyunjung ShimCVPR 2026 · 1 citation
