Learning Sparse Visual Representations via Spatial-Semantic Factorization
Theodore Z. Zhao, Sid Kiblawi, Jianwei Yang, Naoto Usuyama, Reuben Tan, Noel Codella, Tristan Naumann, Hoifung Poon, Mu Wei
Abstract
Self-supervised learning (SSL) faces a fundamental conflict between semantic understanding and image reconstruction. High-level semantic SSL (e.g., DINO) relies on global tokens that are forced to be location-invariant for augmentation alignment, a process that inherently discards the spatial coordinates required for reconstruction. Conversely, generative SSL (e.g., MAE) preserves dense feature grids for reconstruction but fails to produce high-level abstractions. We introduce STELLAR, a framework that resolves this tension by factorizing visual features into a low-rank product of semantic concepts and their spatial distributions. This disentanglement allows us to perform DINO-style augmentation alignment on the semantic tokens while maintaining the precise spatial mapping in the localization matrix necessary for pixel-level reconstruction. We demonstrate that as few as 16 sparse tokens under this factorized form are sufficient to simultaneously support high-quality reconstruction (2.60 FID) and match the semantic performance of dense backbones (79.10% ImageNet accuracy). Our results highlight STELLAR as a versatile sparse representation that bridges the gap between discriminative and generative vision by strategically separating semantic identity from spatial geometry. Code available at https://aka.ms/stellar .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e6e5b9f-329e-494e-9c44-38f675162233Cited by top-tier papers1
Ask how each one uses itBuilds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- CG-SSL: Concept-Guided Self-Supervised LearningSara Atito, Josef Kittler, Imran Razzak, Muhammad AwaisNeurIPS 2025 · 1 citation
- You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised LearningThéo Moutakanni, Maxime Oquab, Marc Szafraniec, Maria Vakalopoulou et al.NeurIPS 2024 · 23 citations
- Learning Where to Learn in Cross-View Self-Supervised LearningLang Huang, Shan You, Mingkai Zheng, Fei Wang et al.CVPR 2022 · 43 citations
- Self-Supervised Learning Based on Transformed Image Reconstruction for Equivariance-Coherent Feature RepresentationQin Wang, Alessio Quercia, Benjamin Bruns, Abigail Morrison et al.AAAI 2026 · 2 citations
- A Theoretical Analysis of Self-Supervised Learning for Vision TransformersYu Huang, Zixin Wen, Yuejie Chi, Yingbin LiangICLR 2025
