Lune

CVPR2026Top-tier venue

COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation

Yuchen Che, JINGTU WU, Hao ZHENG, Asako Kanezaki

2026Year
1Citations

Abstract

making the problem particularly ill-posed. A key to solving this task is to establish reliable cross-view correspondences, since the pose can be recovered by aligning geometric structures between query and reference views. Yet, most existing approaches [39,41] construct correspondences via discrete one-to-one assignments (e.g. argmax), which tend to collapse onto a few dominant keypoints, leaving many of the points unused. Moreover, such non-differentiable assignments breaks differentiability and prevents the model from being trained in an unsupervised manner.

We introduce Confidence-aware Optimal Geometric Correspondence (COG), an unsupervised framework for novel object pose estimation from a single reference image. COG addresses the above issues by formulating soft correspondence as an optimal transport (OT) problem, where point-wise confidences are predicted beforehand and explicitly incorporated as target marginals of the transport plan. Compared to OT-based methods [13,55,72] with uniform marginals and only apply confidence post hoc, This formulation yields globally balanced correspondences that naturally suppresses outliers and non-overlapping regions. Given these correspondences, corresponding points are generated via convex combinations, and a weighted SVD solver [63] is used to recover the pose transformation. The entire process forms an end-to-end correspondence finding and pose estimation pipeline, enabling unsupervised optimization of both correspondence and confidence. To mitigate ambiguity in purely geometric matching, we integrate semantic priors denoised from vision foundation models such as DINO [5,48], which softly encourage correspondences between semantically consistent parts. Furthermore, for unsupervised confidence learning, Gaussian RBF-style kernels of geometric and semantic consistency are used to generate pseudo confidence labels, guiding the network to down weight uncertain points without discarding them entirely and to emphasize reliable regions with high confidence. With these designs, COG naturally extends to the unsupervised setting, where neither groundtruth confidence nor pose supervision is available. Experimental results demonstrate that COG achieves performance comparable to leading supervised approaches, while the supervised variant of COG further outperforms them.

Our contributions are summarized as follows: 1. We formulate correspondence finding as an OT problem with confidence as marginals. Compared to OT with uniform marginals, our formulation yields balanced correspondences by suppressing non-overlapping points. 2. We propose an end-to-end pipeline that jointly learns object pose and point validity confidence without supervision from CAD models, poses, or overlap scores. 3. Unsupervised COG achieves performance competitive with state-of-the-art supervised methods, and its supervised variant further outperforms them.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bc3bbd52-1144-4442-82bb-cea587cedb0e

Builds on30

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines