Learning Sparse Visual Representations via Spatial-Semantic Factorization
Theodore Z. Zhao, Sid Kiblawi, Jianwei Yang, Naoto Usuyama, Reuben Tan, Noel Codella, Tristan Naumann, Hoifung Poon, Mu Wei
摘要
Self-supervised learning (SSL) faces a fundamental conflict between semantic understanding and image reconstruction. High-level semantic SSL (e.g., DINO) relies on global tokens that are forced to be location-invariant for augmentation alignment, a process that inherently discards the spatial coordinates required for reconstruction. Conversely, generative SSL (e.g., MAE) preserves dense feature grids for reconstruction but fails to produce high-level abstractions. We introduce STELLAR, a framework that resolves this tension by factorizing visual features into a low-rank product of semantic concepts and their spatial distributions. This disentanglement allows us to perform DINO-style augmentation alignment on the semantic tokens while maintaining the precise spatial mapping in the localization matrix necessary for pixel-level reconstruction. We demonstrate that as few as 16 sparse tokens under this factorized form are sufficient to simultaneously support high-quality reconstruction (2.60 FID) and match the semantic performance of dense backbones (79.10% ImageNet accuracy). Our results highlight STELLAR as a versatile sparse representation that bridges the gap between discriminative and generative vision by strategically separating semantic identity from spatial geometry. Code available at https://aka.ms/stellar .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
相关 Paper
- CG-SSL: Concept-Guided Self-Supervised LearningSara Atito, Josef Kittler, Imran Razzak, Muhammad AwaisNeurIPS 2025 · 被引用 1 次
- You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised LearningThéo Moutakanni, Maxime Oquab, Marc Szafraniec, Maria Vakalopoulou 等NeurIPS 2024 · 被引用 23 次
- Learning Where to Learn in Cross-View Self-Supervised LearningLang Huang, Shan You, Mingkai Zheng, Fei Wang 等CVPR 2022 · 被引用 43 次
- Self-Supervised Learning Based on Transformed Image Reconstruction for Equivariance-Coherent Feature RepresentationQin Wang, Alessio Quercia, Benjamin Bruns, Abigail Morrison 等AAAI 2026 · 被引用 2 次
- A Theoretical Analysis of Self-Supervised Learning for Vision TransformersYu Huang, Zixin Wen, Yuejie Chi, Yingbin LiangICLR 2025
