Identifying Latent Concepts and Structures for Generalized Category Discovery
Boyang Dai, Chaoqi Chen, Yizhou Yu
Abstract
Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings. However, current approaches primarily focus on designing clustering objectives, often overlooking a critical bottleneck: standard vision backbones yield high-rank, entangled token representations that are ill-suited for unsupervised discovery of latent concepts and structures. In this paper, we propose Compositional Primitive Fields (CPF-GCD), a novel representation learning framework that reshapes the feature space to make such latent structure identifiable by enforcing a low-rank compositional organization. Our core hypothesis is that all categories, whether known or novel, can be expressed as compositions and spatial arrangements of a finite set of learnable visual primitives that capture reusable concepts. CPF instantiates this geometric constraint via a spatial field mechanism. Inserted between the backbone and the head, it rewrites noisy patch tokens through low-rank primitive mixtures, effectively decomposing images into reusable atomic parts and their spatial layouts. By explicitly modeling the spatial distribution of primitives, CPF enables novel categories to emerge naturally as new activation patterns over a shared vocabulary. This shifts the focus of representation from merely partitioning global embeddings to constructing a structured and separable primitive field. Extensive experiments demonstrate that CPF serves as a generic, plug-and-play module that consistently boosts performance across diverse GCD baselines, validating that identifying and leveraging low-rank compositional structure is a crucial inductive bias for open-world recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed6239dd-0b84-4f78-8f1a-b85e18cb5363Builds on24
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
Related papers
- CoGe-GCD: Reframing Generalized Category Discovery with Compositional GeneralizationLuyao Tang, Jiewei Zheng, Kunze Huang, Chaoqi Chen et al.ICML 2026
- Dissecting Generalized Category Discovery: Multiplex Consensus under Self-DeconstructionLuyao Tang, Kunze Huang, Chaoqi Chen, Yuxuan Yuan et al.ICCV 2025 · 3 citations
- Hyperbolic Category DiscoveryYuanpei Liu, Zhenqi He, Kai HanCVPR 2025
- PartCo: Part-Level Correspondence Priors Enhance Category DiscoveryFernando Julio Cendra, Kai HanICML 2026 · 2 citations
- Federated Generalized Category DiscoveryNan Pu, Wenjing Li, Xingyuan Ji, Yalan Qin et al.CVPR 2024 · 11 citations
