Sparse autoencoders reveal selective remapping of visual concepts during adaptation
Hyesu Lim, Jinho Choi, Jaegul Choo, Steffen Schneider
Abstract
Adapting foundation models for specific purposes has become a standard approach to build machine learning systems for downstream applications. Yet, it is an open question which mechanisms take place during adaptation. Here we develop a new Sparse Autoencoder (SAE) for the CLIP vision transformer, named PatchSAE, to extract interpretable concepts at granular levels (e.g., shape, color, or semantics of an object) and their patch-wise spatial attributions. We explore how these concepts influence the model output in downstream image classification tasks and investigate how recent state-of-the-art prompt-based adaptation techniques change the association of model inputs to these concepts. While activations of concepts slightly change between adapted and non-adapted models, we find that the majority of gains on common adaptation tasks can be explained with the existing concepts already present in the non-adapted foundation model. This work provides a concrete framework to train and use SAEs for Vision Transformers and provides insights into explaining adaptation mechanisms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language ModelsMateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge J. Belongie et al.NeurIPS 2025 · 79 citations
- On the Limits of Sparse Autoencoders: A Theoretical Framework and Reweighted RemedyJingyi Cui, Qi Zhang, Yifei Wang, Yisen WangICLR 2026 · 16 citations
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetShufan Shen, Junshu Sun, Qingming Huang, Shuhui WangNeurIPS 2025 · 13 citations
- DNA: Uncovering Universal Latent Forgery KnowledgeJingtong Dou, Chuancheng Shi, Anqi Yi, Shiming Guo et al.ICML 2026 · 8 citations
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf InteractionsHubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer et al.NeurIPS 2025 · 6 citations
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
Related papers
- Inside the Visual Mind: Neuroscience-Motivated Concept Circuits for Interpreting and Steering Vision TransformersTang Li, Yanlin Chen, Mengmeng Ma, Xi PengICML 2026
- Discovering and Steering Interpretable Concepts in Large Generative Music ModelsNikhil Singh, Manuel Cherep, Pattie MaesICLR 2026 · 17 citations
- Interpreting CLIP with Hierarchical Sparse AutoencodersVladimir Zaigrajew, Hubert Baniecki, Przemyslaw BiecekICML 2025
- Residual Stream Analysis with Multi-Layer SAEsTim Lawson, Lucy Farnik, Conor J. Houghton, Laurence AitchisonICLR 2025
- Language Models Can Explain Visual Features via SteeringJavier Ferrando, Enrique Lopez-Cuena, Pablo Agustin Martin-Torres, Daniel Hinjos et al.CVPR 2026 · 2 citations
