Textural or Textual: How Vision-Language Models Read Text in Images
Hanzhang Wang, Qingyuan Ma
摘要
Typographic attacks are often attributed to the ability of multimodal pre-trained models to fuse textual semantics into visual representations, yet the mechanisms and locus of such interference remain unclear. We examine whether such models genuinely encode textual semantics or primarily rely on texture-based visual features. To disentangle orthographic form from meaning, we introduce the ToT dataset, which includes controlled word pairs that either share semantics with distinct appearances (synonyms) or share appearance with differing semantics (paronyms). A layerwise analysis of Intrinsic Dimension (ID) reveals that early layers exhibit competing dynamics between orthographic and semantic representations. In later layers, semantic accuracy increases as ID decreases, but this improvement largely stems from orthographic disambiguation. Notably, clear semantic differentiation emerges only in the final block, challenging the common assumption that semantic understanding is progressively constructed across depth. These findings reveal how current vision-language models construct text representations through texture-dependent processes, prompting a reconsideration of the gap between visual perception and semantic understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum 等ICLR 2021 · 被引用 381 次
- Demystifying CLIP DataHu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang 等ICLR 2024 · 被引用 249 次
- Patching open-vocabulary models by interpolating weightsGabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song 等NeurIPS 2022 · 被引用 230 次
相关 Paper
- Exploring Intrinsic Dimension for Vision-Language Model PruningHanzhang Wang, Jiawen Zhang, Qingyuan MaICML 2024 · 被引用 5 次
- Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language ModelsZoe Wanying He, Sean Trott, Meenakshi KhoslaEMNLP 2025 · 被引用 2 次
- Token Embeddings Alignment for Cross-Modal RetrievalChen-Wei Xie, Jianmin Wu, Yun Zheng, Pan Pan 等ACM MM 2022 · 被引用 18 次
- Parts of Speech-Grounded Subspaces in Vision-Language ModelsJames Oldfield, Christos Tzelepis, Yannis Panagakis, Mihalis Nicolaou 等NeurIPS 2023 · 被引用 13 次
- Are Vision-Language Transformers Learning Multimodal Representations? A Probing PerspectiveEmmanuelle Salin, Badreddine Farah, Stéphane Ayache, Benoît FavreAAAI 2022 · 被引用 49 次
