Unicom: Universal and Compact Representation Learning for Image Retrieval
Xiang An, Jiankang Deng, Kaicheng Yang, Jaiwei Li, Ziyong Feng, Jia Guo, Jing Yang, Tongliang Liu
Abstract
Modern image retrieval methods typically rely on fine-tuning pre-trained encoders to extract image-level descriptors. However, the most widely used models are pre-trained on ImageNet-1K with limited classes. The pre-trained feature representation is therefore not universal enough to generalize well to the diverse open-world classes. In this paper, we first cluster the large-scale LAION 400M dataset into one million pseudo classes based on the joint textual and visual features extracted by the CLIP model. Due to the confusion of label granularity, the automatically clustered dataset inevitably contains heavy inter-class conflict. To alleviate such conflict, we randomly select partial inter-class prototypes to construct the margin-based softmax loss. To further enhance the low-dimensional feature representation, we randomly select partial feature dimensions when calculating the similarities between embeddings and class-wise prototypes. The dual random partial selections are with respect to the class dimension and the feature dimension of the prototype matrix, making the classification conflict-robust and the feature embedding compact. Our method significantly outperforms state-of-the-art unsupervised and supervised image retrieval approaches on multiple benchmarks. The code and pre-trained models are released to facilitate future research https://github.com/deepglint/unicom .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ee4fa62-d916-4175-b5b9-bf5973c53189Cited by top-tier papers21
- ALIP: Adaptive Language-Image Pre-training with Synthetic CaptionKaicheng Yang, Jiankang Deng, Xiang An, Jiawei Li et al.ICCV 2023 · 93 citations
- Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text RetrievalHailang Huang, Zhijie Nie, Ziqiao Wang, Ziyu ShangAAAI 2024 · 47 citations
- CLIP-CID: Efficient CLIP Distillation via Cluster-Instance DiscriminationKaicheng Yang, Tiancheng Gu, Xiang An, Haiqiang Jiang et al.AAAI 2025 · 26 citations
- Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and NegativesZhangchi Feng, Richong Zhang, Zhijie NieACM MM 2024 · 14 citations
- Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMsZhiyu Pan, Yizheng Wu, Jiashen Hua, Junyi Feng et al.ICLR 2026 · 11 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Image Clustering via the Principle of Rate Reduction in the Age of Pretrained ModelsTianzhe Chu, Shengbang Tong, Tianjiao Ding, Xili Dai et al.ICLR 2024 · 22 citations
- Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression LearningChenyu Yang, Xizhou Zhu, Jinguo Zhu, Weijie Su et al.NeurIPS 2024 · 10 citations
- Reproducible Scaling Laws for Contrastive Language-Image LearningMehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman et al.CVPR 2023
- Identifying Interpretable Subspaces in Image RepresentationsNeha Mukund Kalibhat, Shweta Bhardwaj, C. Bayan Bruss, Hamed Firooz et al.ICML 2023 · 41 citations
- Open-Set Fine-Grained Retrieval via Prompting Vision-Language EvaluatorShijie Wang, Jianlong Chang, Haojie Li, Zhihui Wang et al.CVPR 2023
