Joint Representation Learning and Novel Category Discovery on Single- and Multi-modal Data
Xuhui Jia, Kai Han, Yukun Zhu, Bradley Green
Abstract
This paper studies the problem of novel category discovery on single- and multi-modal data with labels from different but relevant categories. We present a generic, end-to-end framework to jointly learn a reliable representation and assign clusters to unlabelled data. To avoid over-fitting the learnt embedding to labelled data, we take inspiration from self-supervised representation learning by noise-contrastive estimation and extend it to jointly handle labelled and unlabelled data. In particular, we propose using category discrimination on labelled data and cross-modal discrimination on multi-modal data to augment instance discrimination used in conventional contrastive learning approaches. We further employ Winner-Take-All (WTA) hashing algorithm on the shared representation space to generate pairwise pseudo labels for unlabelled data to better predict cluster assignments. We thoroughly evaluate our framework on large-scale multi-modal video benchmarks Kinetics-400 and VGG-Sound, and image benchmarks CIFAR10, CIFAR100 and ImageNet, obtaining state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers32
- A Unified Objective for Novel Class DiscoveryEnrico Fini, Enver Sangineto, Stéphane Lathuilière, Zhun Zhong et al.ICCV 2021 · 248 citations
- Generalized Category DiscoverySagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanCVPR 2022 · 194 citations
- Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge DistillationBingchen Zhao, Kai HanNeurIPS 2021 · 161 citations
- Learning Semi-supervised Gaussian Mixture Models for Generalized Category DiscoveryBingchen Zhao, Xin Wen, Kai HanICCV 2023 · 109 citations
- Grow and Merge: A Unified Framework for Continuous Categories DiscoveryXinwei Zhang, Jianwen Jiang, Yutong Feng, Zhi-Fan Wu et al.NeurIPS 2022 · 57 citations
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic et al.NeurIPS 2020 · 423 citations
Related papers
- Neighborhood Contrastive Learning for Novel Class DiscoveryZhun Zhong, Enrico Fini, Subhankar Roy, Zhiming Luo et al.CVPR 2021
- Automatically Discovering and Learning New Visual Categories with Ranking StatisticsKai Han, Sylvestre-Alvise Rebuffi, Sébastien Ehrhardt, Andrea Vedaldi et al.ICLR 2020 · 222 citations
- CURE: Consistency-under-Unified Semantic Regularization for Generalized Category DiscoveryYuwei Bian, Shidong Wang, Haofeng ZhangICML 2026
- Audio-Visual Instance Discrimination with Cross-Modal AgreementPedro Morgado, Nuno Vasconcelos, Ishan MisraCVPR 2021
- Dynamic Conceptional Contrastive Learning for Generalized Category DiscoveryNan Pu, Zhun Zhong, Nicu SebeCVPR 2023
