Textural or Textual: How Vision-Language Models Read Text in Images
Hanzhang Wang, Qingyuan Ma
Abstract
Typographic attacks are often attributed to the ability of multimodal pre-trained models to fuse textual semantics into visual representations, yet the mechanisms and locus of such interference remain unclear. We examine whether such models genuinely encode textual semantics or primarily rely on texture-based visual features. To disentangle orthographic form from meaning, we introduce the ToT dataset, which includes controlled word pairs that either share semantics with distinct appearances (synonyms) or share appearance with differing semantics (paronyms). A layerwise analysis of Intrinsic Dimension (ID) reveals that early layers exhibit competing dynamics between orthographic and semantic representations. In later layers, semantic accuracy increases as ID decreases, but this improvement largely stems from orthographic disambiguation. Notably, clear semantic differentiation emerges only in the final block, challenging the common assumption that semantic understanding is progressively constructed across depth. These findings reveal how current vision-language models construct text representations through texture-dependent processes, prompting a reconsideration of the gap between visual perception and semantic understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e37711f7-b351-405e-9cc3-e971705cee1aCited by top-tier papers1
Ask how each one uses itBuilds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
- Demystifying CLIP DataHu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang et al.ICLR 2024 · 249 citations
- Patching open-vocabulary models by interpolating weightsGabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song et al.NeurIPS 2022 · 230 citations
Related papers
- Exploring Intrinsic Dimension for Vision-Language Model PruningHanzhang Wang, Jiawen Zhang, Qingyuan MaICML 2024 · 5 citations
- Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language ModelsZoe Wanying He, Sean Trott, Meenakshi KhoslaEMNLP 2025 · 2 citations
- Token Embeddings Alignment for Cross-Modal RetrievalChen-Wei Xie, Jianmin Wu, Yun Zheng, Pan Pan et al.ACM MM 2022 · 18 citations
- Parts of Speech-Grounded Subspaces in Vision-Language ModelsJames Oldfield, Christos Tzelepis, Yannis Panagakis, Mihalis Nicolaou et al.NeurIPS 2023 · 13 citations
- Are Vision-Language Transformers Learning Multimodal Representations? A Probing PerspectiveEmmanuelle Salin, Badreddine Farah, Stéphane Ayache, Benoît FavreAAAI 2022 · 49 citations
