SpectralGCD: Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
Lorenzo Caselli, Marco Mistretta, Simone Magistri, Andrew D. Bagdanov
Abstract
Generalized Category Discovery (GCD) aims to identify novel categories in unlabeled data while leveraging a small labeled subset of known classes. Training a parametric classifier solely on image features often leads to overfitting to old classes, and recent multimodal approaches improve performance by incorporating textual information. However, they treat modalities independently and incur high computational cost. We propose SpectralGCD, an efficient and effective multimodal approach to GCD that uses CLIP cross-modal image-concept similarities as a unified cross-modal representation. Each image is expressed as a mixture over semantic concepts from a large task-agnostic dictionary, which anchors learning to explicit semantics and reduces reliance on spurious visual cues. To maintain the semantic quality of representations learned by an efficient student, we introduce Spectral Filtering which exploits a cross-modal covariance matrix over the softmaxed similarities measured by a strong teacher model to automatically retain only relevant concepts from the dictionary. Forward and reverse knowledge distillation from the same teacher ensures that the cross-modal representations of the student remain both semantically sufficient and well-aligned. Across six benchmarks, SpectralGCD delivers accuracy comparable to or significantly superior to state-of-the-art methods at a fraction of the computational cost. The code is publicly available at: https://github.com/miccunifi/SpectralGCD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e496b5f-c473-43ef-b194-4f784dd55d40Cited by top-tier papers1
Ask how each one uses itBuilds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Open-Set Recognition: A Good Closed-Set Classifier is All You NeedSagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanICLR 2022 · 594 citations
- Data Filtering NetworksAlex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt et al.ICLR 2024 · 251 citations
- A Unified Objective for Novel Class DiscoveryEnrico Fini, Enver Sangineto, Stéphane Lathuilière, Zhun Zhong et al.ICCV 2021 · 248 citations
Related papers
- Prior-Constrained Association Learning for Fine-Grained Generalized Category DiscoveryMenglin Wang, Zhun Zhong, Xiaojin GongAAAI 2025 · 4 citations
- GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category DiscoveryEnguang Wang, Zhimao Peng, Zhengyuan Xie, Fei Yang et al.CVPR 2025
- DebGCD: Debiased Learning with Distribution Guidance for Generalized Category DiscoveryYuanpei Liu, Kai HanICLR 2025
- Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category DiscoveryWei He, Xianghan Meng, Zhiyuan Huang, Xianbiao Qi et al.CVPR 2026
- Solving the Catastrophic Forgetting Problem in Generalized Category DiscoveryXinzi Cao, Xiawu Zheng, Guanhong Wang, Weijiang Yu et al.CVPR 2024
