Dual Pose-invariant Embeddings: Learning Category and Object-specific Discriminative Representations for Recognition and Retrieval
Rohan Sarkar, Avinash C. Kak
Abstract
In the context of pose-invariant object recognition and retrieval, we demonstrate that it is possible to achieve significant improvements in performance if both the category-based and the object-identity-based embed-dings are learned simultaneously during training. In hind-sight, that sounds intuitive because learning about the cat-egories is more fundamental than learning about the indi-vidual objects that correspond to those categories. How-ever, to the best of what we know, no prior work in pose-invariant learning has demonstrated this effect. This paper presents an attention-based dual-encoder architecture with specially designed loss functions that optimize the inter-and intra-class distances simultaneously in two different embedding spaces, one for the category embeddings and the other for the object level embeddings. The loss functions we have proposed are pose-invariant ranking losses that are designed to minimize the intra-class distances and maximize the inter-class distances in the dual representation spaces. We demonstrate the power of our approach with three challenging multi-view datasets, Model Net-40, ObjectPI, and FG3D. With our dual approach, for single-view object recognition, we outperform the previous best by 20.0% on ModelNet40, 2.0% on ObjectPI, and 46.5% on FG3D. On the other hand, for single-view object retrieval, we outperform the previous best by 33.7% on ModelNet40, 18.8% on ObjectPI, and 56.9% on FG3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning Canonical View Representation for 3D Shape Recognition with Arbitrary ViewsXin Wei, Yifei Gong, Fudong Wang, Xing Sun et al.ICCV 2021 · 19 citations
- Cross-Batch Memory for Embedding LearningXun Wang, Haozhi Zhang, Weilin Huang, Matthew R. ScottCVPR 2020
Related papers
- Learning Relationships for Multi-View 3D Object RecognitionZe Yang, Liwei WangICCV 2019 · 166 citations
- DualPoseNet: Category-level 6D Object Pose and Size Estimation Using Dual Pose Network with Refined Learning of Pose ConsistencyJiehong Lin, Zewei Wei, Zhihao Li, Songcen Xu et al.ICCV 2021 · 169 citations
- Discovering Relationships Between Object Categories via Universal Canonical MapsNatalia Neverova, Artsiom Sanakoyeu, Patrick Labatut, David Novotný et al.CVPR 2021
- One2Any: One-Reference 6D Pose Estimation for Any ObjectMengya Liu, Siyuan Li, Ajad Chhatkuli, Prune Truong et al.CVPR 2025
- Multi-Path Learning for Object Pose Estimation Across DomainsMartin Sundermeyer, Maximilian Durner, En Yen Puang, Zoltan-Csaba Marton et al.CVPR 2020
