Deciphering 'What' and 'Where' Visual Pathways from Spectral Clustering of Layer-Distributed Neural Representations
Xiao Zhang, David Yunis, Michael Maire
Abstract
We present an approach for analyzing grouping information contained within a neural network's activations, per-mitting extraction of spatial layout and semantic segmentation from the behavior of large pre-trained vision models. Unlike prior work, our method conducts a who lis tic analysis of a network's activation state, leveraging features from all layers and obviating the need to guess which part of the model contains relevant information. Motivated by classic spectral clustering, we formulate this analysis in terms of an optimization objective involving a set of affinity matrices, each formed by comparing features within a different layer. Solving this optimization problem using gradient descent allows our technique to scale from single images to dataset-level analysis, including, in the latter, both intra-and inter-image relationships. Analyzing a pre-trained generative transformer provides insight into the computational strategy learned by such models. Equating affinity with key-query similarity across attention layers yields eigenvectors encoding scene spatial layout, whereas defining affinity by value vector similarity yields eigenvectors encoding object identity. This result suggests that key and query vectors co-ordinate attentional information flow according to spatial proximity (a ‘where’ pathway), while value vectors refine a semantic category representation (a ‘what’ pathway).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c7ef78a-6b36-49db-a42a-99a9b98b1bd9Cited by top-tier papers6
- Making Training-Free Diffusion Segmentors Scale with the Generative PowerBenyuan Meng, Qianqian Xu, Zitai Wang, Xiaochun Cao et al.CVPR 2026 · 2 citations
- SCCS: Deep Neural Spectral Clustering for Self-Supervised Subcellular Structure SegmentationJimao Jiang, Diya Sun, Tianbing Wang, Yuru PeiAAAI 2025 · 1 citation
- Residual Connections Harm Generative Representation LearningXiao Zhang, Ruoxi Jiang, William Gao, Rebecca Willett et al.CVPR 2026
- Nested Diffusion Models Using Hierarchical Latent PriorsXiao Zhang, Ruoxi Jiang, Rebecca Willett, Michael MaireCVPR 2025
- Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation ModelMostofa Rafid Uddin, H. M. Shadman Tabib, Thanh-Huy Nguyen, Kashish Gandhi et al.CVPR 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Dissecting Query-Key Interaction in Vision TransformersXu Pan, Aaron Philip, Ziqian Xie, Odelia SchwartzNeurIPS 2024 · 18 citations
- Analyzing Vision Transformers for Image Classification in Class Embedding SpaceMartina G. Vilas, Timothy Schaumlöffel, Gemma RoigNeurIPS 2023 · 43 citations
- Learning Neural Eigenfunctions for Unsupervised Semantic SegmentationZhijie Deng, Yucen LuoICCV 2023 · 7 citations
- Native Segmentation Vision TransformersGuillem Brasó, Aljosa Osep, Laura Leal-TaixéNeurIPS 2025 · 2 citations
- Decomposing Query-Key Feature Interactions Using Contrastive CovariancesAndrew Lee, Yonatan Belinkov, Fernanda Viégas, Martin WattenbergICML 2026 · 1 citation
