CG-SSL: Concept-Guided Self-Supervised Learning
Sara Atito, Josef Kittler, Imran Razzak, Muhammad Awais
Abstract
Humans understand visual scenes by first capturing a global impression and then refining this understanding into distinct, object-like components. Inspired by this process, we introduce C oncept-G uided S elf-S upervised L earning (CG-SSL), a novel framework that brings structure and interpretability to representation learning through a curriculum of three training phases: (1) global scene encoding, (2) discovery of visual concepts via tokenised cross-attention, and (3) alignment of these concepts across views. Unlike traditional SSL methods, which simply enforce similarity between multiple augmented views of the same image, CG-SSL accounts for the fact that these views may highlight different parts of an object or scene. To address this, our method establishes explicit correspondences between views and aligns the representations of meaningful image regions. At its core, CG-SSL augments standard SSL with a lightweight decoder that learns and refines concept tokens via cross-attention with patch features. The concept tokens are trained using masked concept distillation and a feature-space reconstruction objective. A final alignment stage enforces view consistency by geometrically matching concept regions under heavy augmentation, enabling more compact, robust, and disentangled representations of scene regions. Across multiple backbone sizes, CG-SSL achieves state-of-the-art results on image segmentation benchmarks using k - NN and linear probes, substantially outperforming prior methods and approaching, or even surpassing, the performance of leading SSL models trained on over 100 × more data. Code and pretrained models will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb105445-76c1-42c7-afe0-70fb60f01ebdCited by top-tier papers1
Ask how each one uses itBuilds on20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- Unsupervised Object-Level Representation Learning from Scene ImagesJiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong et al.NeurIPS 2021 · 93 citations
- Visual Concepts TokenizationTao Yang, Yuwang Wang, Yan Lu, Nanning ZhengNeurIPS 2022 · 19 citations
- ACSeg: Adaptive Conceptualization for Unsupervised Semantic SegmentationKehan Li, Zhennan Wang, Zesen Cheng, Runyi Yu et al.CVPR 2023
- Decoupled Global-Local Alignment for Improving Compositional UnderstandingXiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu et al.ACM MM 2025 · 4 citations
- Learning Sparse Visual Representations via Spatial-Semantic FactorizationTheodore Z. Zhao, Sid Kiblawi, Jianwei Yang, Naoto Usuyama et al.ICML 2026
