Visual symbolic mechanisms: Emergent symbol processing in Vision Language Models
Rim Assouel, Declan Iain Campbell, Yoshua Bengio, Taylor Whittington Webb
Abstract
To accurately process a visual scene, observers must bind features together to represent individual objects. This capacity is necessary, for instance, to distinguish an image containing a red square and a blue circle from an image containing a blue square and a red circle. Recent work has found that language models solve this ‘binding problem’ via a set of symbol-like, content-independent indices, but it is unclear whether similar mechanisms are employed by Vision Language Models (VLM). This question is especially relevant, given the persistent failures of VLMs on tasks that require binding. Here, we identify a previously unknown set of emergent symbolic mechanisms that support binding specifically in VLMs, via a content-independent, spatial indexing scheme. Moreover, we find that binding errors, when they occur, can be traced directly to failures in these mechanisms. Taken together, these results shed light on the mechanisms that support symbol-like processing in VLMs, and suggest possible avenues for reducing the number of binding failures exhibited by these models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- The Geometry of Representational Failures in Vision Language ModelsDaniele Savietto, Declan Campbell, André Panisson, Marco Nurisso et al.ICML 2026 · 5 citations
- Uncovering Grounding IDs: How External Cues Shape Multi-Modal BindingAmirmohammad Izadi, Hosein Hasani, Fatemeh Askari, Mobin Bagherian et al.ICML 2026 · 3 citations
- Necessary Conditions for Compositional Generalization of Embedding ModelsArnas Uselis, Andrea Dittadi, Seong Joon OhICML 2026
- Formalizing the Binding ProblemLianghuan Huang, Yihao Li, Saeed Salehi, Yingshan Chang et al.ICML 2026
- How can embedding models bind concepts?Arnas Uselis, Darina Koishigarina, Seong Joon OhICML 2026
Builds on15
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller et al.ICLR 2024 · 229 citations
- Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemDeclan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata et al.NeurIPS 2024 · 101 citations
- Break It Down: Evidence for Structural Compositionality in Neural NetworksMichael A. Lepori, Thomas Serre, Ellie PavlickNeurIPS 2023 · 66 citations
Related papers
- How do Language Models Bind Entities in Context?Jiahai Feng, Jacob SteinhardtICLR 2024 · 81 citations
- Seeing to Generalize: How Visual Data Corrects Binding ShortcutsNicolas Buzeta, Felipe del Rio, Cristian Hinostroza, Denis Parra et al.ICML 2026 · 1 citation
- Linear Mechanisms for Spatiotemporal Reasoning in Vision Language ModelsRaphaela Kang, Hongqiao Chen, Georgia Gkioxari, Pietro PeronaICLR 2026 · 11 citations
- Towards Interpreting Visual Information Processing in Vision-Language ModelsClement Neo, Luke Ong, Philip Torr, Mor Geva et al.ICLR 2025
- The Mechanistic Emergence of Symbol Grounding in Language ModelsShuyu Wu, Ziqiao Ma, Xiaoxi Luo, Yidong Huang et al.ICML 2026 · 4 citations
