Single Image 3D Shape Retrieval via Cross-Modal Instance and Category Contrastive Learning
Ming-Xian Lin, Jie Yang, He Wang, Yu-Kun Lai, Rongfei Jia, Binqiang Zhao, Lin Gao
摘要
In this work, we tackle the problem of single imagebased 3D shape retrieval (IBSR), where we seek to find the most matched shape of a given single 2D image from a shape repository. Most of the existing works learn to embed 2D images and 3D shapes into a common feature space and perform metric learning using a triplet loss. Inspired by the great success in recent contrastive learning works on self-supervised representation learning, we propose a novel IBSR pipeline leveraging contrastive learning. We note that adopting such cross-modal contrastive learning between 2D images and 3D shapes into IBSR tasks is non-trivial and challenging: contrastive learning requires very strong data augmentation in constructed positive pairs to learn the feature invariance, whereas traditional metric learning works do not have this requirement. Moreover, object shape and appearance are entangled in 2D query images, thus making the learning task more difficult than contrasting single-modal data. To mitigate the challenges, we propose to use multi-view grayscale rendered images from the 3D shapes as a shape representation. We then introduce a strong data augmentation technique based on color transfer, which can significantly but naturally change the appearance of the query image, effectively satisfying the need for contrastive learning. Finally, we propose to incorporate a novel category-level contrastive loss that helps distinguish similar objects from different categories, in addition to classic instance-level contrastive loss. Our experiments demonstrate that our approach achieves the best performance on * Corresponding Author is Lin Gao all the three popular IBSR benchmarks, including Pix3D, Stanford Cars, and Comp Cars, outperforming the previous state-of-the-art from 4% -15% on retrieval accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Cross-Modal Contrastive Learning for Domain Adaptation in 3D Semantic SegmentationBowei Xing, Xianghua Ying, Ruibin Wang, Jinfa Yang 等AAAI 2023 · 被引用 23 次
- Fine-grained Prototypical Voting with Heterogeneous Mixup for Semi-supervised 2D-3D Cross-modal RetrievalFan Zhang, Xian-Sheng Hua, Chong Chen, Xiao LuoCVPR 2024 · 被引用 5 次
- DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal RetrievalFan Zhang, Changhu Wang, Zebang Cheng, Xiaojiang Peng 等AAAI 2025 · 被引用 1 次
- Looking 3D: Anomaly Detection with 2D-3D AlignmentAnkan Bhunia, Changjian Li, Hakan BilenCVPR 2024
- Learning Geometric-Aware Properties in 2D Representation Using Lightweight CAD Models, or Zero Real 3D PairsPattaramanee Arsomngern, Sarana Nutanong, Supasorn SuwajanakornCVPR 2023
它引用的顶会 Paper15
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 被引用 1,553 次
- Contrastive learning of global and local features for medical image segmentation with limited annotationsKrishna Chaitanya, Ertunc Erdil, Neerav Karani, Ender KonukogluNeurIPS 2020 · 被引用 714 次
- SoftTriple Loss: Deep Metric Learning Without Triplet SamplingQi Qian, Lei Shang, Baigui Sun, Juhua Hu 等ICCV 2019 · 被引用 419 次
相关 Paper
- Cross-Domain 3D Model Retrieval Based On Contrastive Learning And Label PropagationDan Song, Yue Yang, Weizhi Nie, Xuanya Li 等ACM MM 2022 · 被引用 6 次
- Hard Example Generation by Texture Synthesis for Cross-domain Shape Similarity LearningHuan Fu, Shunming Li, Rongfei Jia, Mingming Gong 等NeurIPS 2020 · 被引用 20 次
- C3CMR: Cross-Modality Cross-Instance Contrastive Learning for Cross-Media RetrievalJunsheng Wang, Tiantian Gong, Zhixiong Zeng, Changchang Sun 等ACM MM 2022 · 被引用 12 次
- Doodle to Object: Practical Zero-Shot Sketch-Based 3D Shape RetrievalBingrui Wang, Yuan ZhouAAAI 2023 · 被引用 3 次
- Unleashing Potential of Unsupervised Pre-Training with Intra-Identity Regularization for Person Re-IdentificationZizheng Yang, Xin Jin, Kecheng Zheng, Feng ZhaoCVPR 2022 · 被引用 30 次
