Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
Thomas Fel, Binxu Wang, Michael A. Lepori, Matthew Kowal, Andrew Lee, Randall Balestriero, Sonia Joseph, Ekdeep Singh Lubana, Talia Konkle, Demba E. Ba, Martin Wattenberg
摘要
DINOv2 sees the world well enough to guide robots and segment images, but we still do not know what it sees. We conduct the first comprehensive analysis of DINOv2’s representational structure using overcomplete dictionary learning, extracting over 32,000 visual concepts in what constitutes the largest interpretability demonstration for any vision foundation model to date. This method provides the backbone of our study, which unfolds in three parts.
In the first part, we analyze how different downstream tasks recruit concepts from our learned dictionary, revealing functional specialization: classification exploits “Elsewhere” concepts that fire everywhere except on target objects, implementing learned negations; segmentation relies exclusively on boundary detectors forming coherent subspaces; depth estimation draws on three distinct monocular cue families matching visual neuroscience principles.
Turning to concept geometry and statistics, we find the learned dictionary deviates from ideal near-orthogonal (Grassmannian) structure, exhibiting higher coherence than random baselines. Concept atoms are not aligned with the neuron basis, confirming distributed encoding. We discover antipodal concept pairs that encode opposite semantics (e.g., “white shirt” vs “black shirt”), creating signed semantic axes. Separately, we identify concepts that activate exclusively on register tokens, revealing these encode global scene properties like motion blur and illumination. Across layers, positional information collapses toward a 2D sheet, yet within single images token geometry remains smooth and clustered even after position is removed, putting into question a purely sparse-coding view of representation.
To resolve this paradox, we advance a different view: tokens are formed by combining convex mixtures of a few archetypes (e.g., a rabbit among animals, brown among colors, fluffy among textures). Multi-head attention directly implements this construction, with activations behaving like sums of convex regions. In this picture, concepts are expressed by proximity to landmarks and by regions—not by unbounded linear directions. We call this the Minkowski Representation Hypothesis (MRH), and we examine its empirical signals and consequences for how we study, steer, and interpret vision-transformer representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Priors in time: Missing inductive biases for language model interpretabilityEkdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur, Valérie Costa 等ICLR 2026 · 被引用 19 次
- Circuit Mechanisms for Spatial Relation Generation in Diffusion TransformersBinxu Wang, Jingxuan Fan, Xu PanCVPR 2026 · 被引用 4 次
- Language Models Can Explain Visual Features via SteeringJavier Ferrando, Enrique Lopez-Cuena, Pablo Agustin Martin-Torres, Daniel Hinjos 等CVPR 2026 · 被引用 2 次
- Necessary Conditions for Compositional Generalization of Embedding ModelsArnas Uselis, Andrea Dittadi, Seong Joon OhICML 2026
- From Measurement to Mitigation: Quantifying and Reducing Identity Leakage in Image Representation Encoders with Linear Subspace RemovalDaniel George, Charles Yeh, Daniel Lee, Yifei ZhangCVPR 2026
它引用的顶会 Paper45
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
相关 Paper
- A Concept-Based Explainability Framework for Large Multimodal ModelsJayneel Parekh, Pegah Khayatan, Mustafa Shukor, Alasdair Newson 等NeurIPS 2024 · 被引用 48 次
- Understanding Video Transformers via Universal Concept DiscoveryMatthew Kowal, Achal Dave, Rares Ambrus, Adrien Gaidon 等CVPR 2024
- Register and [CLS] tokens induce a decoupling of local and global features in large ViTsAlexander Lappe, Martin A. GieseNeurIPS 2025 · 被引用 9 次
- Disentangling the Factors of Convergence between Brains and DINOv3Joséphine Raugel, Marc Szafraniec, Huy V. Vo, Camille Couprie 等ICLR 2026
- Beyond Scalars: Concept-Based Alignment Analysis in Vision TransformersJohanna Vielhaben, Dilyara Bareeva, Jim Berend, Wojciech Samek 等NeurIPS 2025 · 被引用 11 次
