Visual symbolic mechanisms: Emergent symbol processing in Vision Language Models
Rim Assouel, Declan Iain Campbell, Yoshua Bengio, Taylor Whittington Webb
摘要
To accurately process a visual scene, observers must bind features together to represent individual objects. This capacity is necessary, for instance, to distinguish an image containing a red square and a blue circle from an image containing a blue square and a red circle. Recent work has found that language models solve this ‘binding problem’ via a set of symbol-like, content-independent indices, but it is unclear whether similar mechanisms are employed by Vision Language Models (VLM). This question is especially relevant, given the persistent failures of VLMs on tasks that require binding. Here, we identify a previously unknown set of emergent symbolic mechanisms that support binding specifically in VLMs, via a content-independent, spatial indexing scheme. Moreover, we find that binding errors, when they occur, can be traced directly to failures in these mechanisms. Taken together, these results shed light on the mechanisms that support symbol-like processing in VLMs, and suggest possible avenues for reducing the number of binding failures exhibited by these models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The Geometry of Representational Failures in Vision Language ModelsDaniele Savietto, Declan Campbell, André Panisson, Marco Nurisso 等ICML 2026 · 被引用 5 次
- Uncovering Grounding IDs: How External Cues Shape Multi-Modal BindingAmirmohammad Izadi, Hosein Hasani, Fatemeh Askari, Mobin Bagherian 等ICML 2026 · 被引用 3 次
- Necessary Conditions for Compositional Generalization of Embedding ModelsArnas Uselis, Andrea Dittadi, Seong Joon OhICML 2026
- Formalizing the Binding ProblemLianghuan Huang, Yihao Li, Saeed Salehi, Yingshan Chang 等ICML 2026
- How can embedding models bind concepts?Arnas Uselis, Darina Koishigarina, Seong Joon OhICML 2026
它引用的顶会 Paper15
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller 等ICLR 2024 · 被引用 229 次
- Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemDeclan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata 等NeurIPS 2024 · 被引用 101 次
- Break It Down: Evidence for Structural Compositionality in Neural NetworksMichael A. Lepori, Thomas Serre, Ellie PavlickNeurIPS 2023 · 被引用 66 次
相关 Paper
- How do Language Models Bind Entities in Context?Jiahai Feng, Jacob SteinhardtICLR 2024 · 被引用 81 次
- Seeing to Generalize: How Visual Data Corrects Binding ShortcutsNicolas Buzeta, Felipe del Rio, Cristian Hinostroza, Denis Parra 等ICML 2026 · 被引用 1 次
- Linear Mechanisms for Spatiotemporal Reasoning in Vision Language ModelsRaphaela Kang, Hongqiao Chen, Georgia Gkioxari, Pietro PeronaICLR 2026 · 被引用 11 次
- Towards Interpreting Visual Information Processing in Vision-Language ModelsClement Neo, Luke Ong, Philip Torr, Mor Geva 等ICLR 2025
- The Mechanistic Emergence of Symbol Grounding in Language ModelsShuyu Wu, Ziqiao Ma, Xiaoxi Luo, Yidong Huang 等ICML 2026 · 被引用 4 次
