Deciphering 'What' and 'Where' Visual Pathways from Spectral Clustering of Layer-Distributed Neural Representations
Xiao Zhang, David Yunis, Michael Maire
摘要
We present an approach for analyzing grouping information contained within a neural network's activations, per-mitting extraction of spatial layout and semantic segmentation from the behavior of large pre-trained vision models. Unlike prior work, our method conducts a who lis tic analysis of a network's activation state, leveraging features from all layers and obviating the need to guess which part of the model contains relevant information. Motivated by classic spectral clustering, we formulate this analysis in terms of an optimization objective involving a set of affinity matrices, each formed by comparing features within a different layer. Solving this optimization problem using gradient descent allows our technique to scale from single images to dataset-level analysis, including, in the latter, both intra-and inter-image relationships. Analyzing a pre-trained generative transformer provides insight into the computational strategy learned by such models. Equating affinity with key-query similarity across attention layers yields eigenvectors encoding scene spatial layout, whereas defining affinity by value vector similarity yields eigenvectors encoding object identity. This result suggests that key and query vectors co-ordinate attentional information flow according to spatial proximity (a ‘where’ pathway), while value vectors refine a semantic category representation (a ‘what’ pathway).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Making Training-Free Diffusion Segmentors Scale with the Generative PowerBenyuan Meng, Qianqian Xu, Zitai Wang, Xiaochun Cao 等CVPR 2026 · 被引用 2 次
- SCCS: Deep Neural Spectral Clustering for Self-Supervised Subcellular Structure SegmentationJimao Jiang, Diya Sun, Tianbing Wang, Yuru PeiAAAI 2025 · 被引用 1 次
- Residual Connections Harm Generative Representation LearningXiao Zhang, Ruoxi Jiang, William Gao, Rebecca Willett 等CVPR 2026
- Nested Diffusion Models Using Hierarchical Latent PriorsXiao Zhang, Ruoxi Jiang, Rebecca Willett, Michael MaireCVPR 2025
- Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation ModelMostofa Rafid Uddin, H. M. Shadman Tabib, Thanh-Huy Nguyen, Kashish Gandhi 等CVPR 2026
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Dissecting Query-Key Interaction in Vision TransformersXu Pan, Aaron Philip, Ziqian Xie, Odelia SchwartzNeurIPS 2024 · 被引用 18 次
- Analyzing Vision Transformers for Image Classification in Class Embedding SpaceMartina G. Vilas, Timothy Schaumlöffel, Gemma RoigNeurIPS 2023 · 被引用 43 次
- Learning Neural Eigenfunctions for Unsupervised Semantic SegmentationZhijie Deng, Yucen LuoICCV 2023 · 被引用 7 次
- Native Segmentation Vision TransformersGuillem Brasó, Aljosa Osep, Laura Leal-TaixéNeurIPS 2025 · 被引用 2 次
- Decomposing Query-Key Feature Interactions Using Contrastive CovariancesAndrew Lee, Yonatan Belinkov, Fernanda Viégas, Martin WattenbergICML 2026 · 被引用 1 次
