The Geometry of Representational Failures in Vision Language Models
Daniele Savietto, Declan Campbell, André Panisson, Marco Nurisso, Giovanni Petri, Jonathan Cohen, Alan Perotti
摘要
Vision-Language Models (VLMs) exhibit puzzling failures in multi-object visual tasks, such as hallucinating non-existent elements or failing to identify the most similar objects among distractions. While these errors mirror human cognitive constraints, such as the "Binding Problem'', the internal mechanisms driving them in artificial systems remain poorly understood. Here, we propose a mechanistic insight by analyzing the representational geometry of open-weight VLMs (Qwen, InternVL, Gemma), comparing methodologies to distill "concept vectors'' - latent directions encoding visual concepts. We validate our concept vectors via steering interventions that reliably manipulate model behavior in both simplified and naturalistic vision tasks (e.g., forcing the model to perceive a red flower as blue). We observe that the geometric overlap between these vectors strongly correlates with specific error patterns, offering a grounded quantitative framework to understand how internal representations shape model behavior and drive visual failures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemDeclan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata 等NeurIPS 2024 · 被引用 101 次
- Emergence of Separable Manifolds in Deep Language RepresentationsJonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson 等ICML 2020 · 被引用 52 次
- Visual symbolic mechanisms: Emergent symbol processing in Vision Language ModelsRim Assouel, Declan Iain Campbell, Yoshua Bengio, Taylor Whittington WebbICLR 2026 · 被引用 14 次
相关 Paper
- Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMsZhining Liu, Ziyi Chen, Hui Liu, Chen Luo 等ICLR 2026 · 被引用 47 次
- Textual Supervision Enhances Geospatial Representations in Vision-Language ModelsMarcelo Sartori Locatelli, Fernando Tonucci, Jea Kwon, Luiz Felipe Vecchietti 等ICML 2026
- How can embedding models bind concepts?Arnas Uselis, Darina Koishigarina, Seong Joon OhICML 2026
- How do Language Models Bind Entities in Context?Jiahai Feng, Jacob SteinhardtICLR 2024 · 被引用 81 次
- Injection Without Distortion: Geometrically Constrained Knowledge Enhancement for Vision-Language ModelsZhongze Wu, Xiu Su, Feng Yang, Shan You 等AAAI 2026
