Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
Wei He, Xianghan Meng, Zhiyuan Huang, Xianbiao Qi, Rong Xiao, Chun-Guang Li
Abstract
Generalized Category Discovery (GCD) aims to identify both known and unknown categories, with only partial labels given for the known categories, posing a challenging open-set recognition problem. State-of-the-art approaches for GCD are usually built on multi-modality representation learning, which pays heavily attention to inter-modality alignment rather than intra-modality alignment. In this paper, we propose a novel and effective multi-modal representation learning approach for GCD via Semi-Supervised Rate Reduction, called SSR 2 -GCD, to learn cross-modality representations with desired underlying structures via properly harnessing intra-modality alignment. Moreover, to boost knowledge transfer, we integrate the information from prompt candidates by leveraging the inter-modal alignment offered by Vision Language Models. We conduct extensive experiments on generic and fine-grained benchmark datasets, demonstrating superior performance of the proposed approach and verifying the importance of intramodality alignment. The code is available at: https: //github.com/hewei98/SSR2-GCD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a2fd6a1-63ff-4fc0-ad45-968c7771352cBuilds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate ReductionYaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song et al.NeurIPS 2020 · 265 citations
- Zero-Shot Composed Image Retrieval with Textual InversionAlberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del BimboICCV 2023 · 214 citations
- Generalized Category DiscoverySagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanCVPR 2022 · 194 citations
Related papers
- GenDis: Generative-Discriminative Dual-View Co-Training for Generalized Category DiscoveryXi Chen, Chuan Qin, Jinpeng Li, Shasha Hu et al.ACL 2026
- Learning Semi-supervised Gaussian Mixture Models for Generalized Category DiscoveryBingchen Zhao, Xin Wen, Kai HanICCV 2023 · 109 citations
- A Unified Knowledge Transfer Network for Generalized Category DiscoveryWenkai Shi, Wenbin An, Feng Tian, Yan Chen et al.AAAI 2024 · 10 citations
- Transfer and Alignment Network for Generalized Category DiscoveryWenbin An, Feng Tian, Wenkai Shi, Yan Chen et al.AAAI 2024 · 17 citations
- ALLGCD: Leveraging All Unlabeled Data for Generalized Category DiscoveryXinzi Cao, Ke Chen, Feidiao Yang, Xiawu Zheng et al.ICCV 2025 · 2 citations
