Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment
Rohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency
摘要
Despite recent progress towards scaling up multimodal vision-language models, these models are still known to struggle on compositional generalization benchmarks such as Winoground. We find that a critical component lacking from current vision-language models is relation-level alignment: the ability to match directional semantic relations in text (e.g., 'mug in grass') with spatial relationships in the image (e.g., the position of the mug relative to the grass). To tackle this problem, we show that relation alignment can be enforced by encouraging the language attention from 'mug' to 'grass' (capturing the semantic relation 'in') to match the visual attention from the mug to the grass (capturing the corresponding physical relation). Tokens and their corresponding objects are softly identified using a weighted mean of cross-modal attention. We prove that this notion of soft cross-modal equivalence is equivalent to enforcing congruence between vision and language attention matrices under a 'change of basis' provided by the cross-modal attention matrix. Intuitively, our approach projects visual attention into the language attention space to calculate its divergence from the actual language attention, and vice versa. We apply our Cross-modal Attention Congruence Regularization (CACR) loss to fine-tune UNITER and improve its Winoground Group score by 5.75 points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Encoding and Controlling Global Semantics for Long-form Video Question AnsweringThong Nguyen, Zhiyuan Hu, Xiaobao Wu, Cong-Duy Nguyen 等EMNLP 2024 · 被引用 3 次
- Progressive Compositionality in Text-to-Image Generative ModelsXu Han, Linghao Jin, Xiaofeng Liu, Paul Pu LiangICLR 2025
- OLIVE: Object Level In-Context Visual EmbeddingsTimothy Ossowski, Junjie HuACL 2024
- EventGPT: Event Stream Understanding with Multimodal Large Language ModelsShaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang 等CVPR 2025
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun 等AAAI 2021 · 被引用 414 次
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh 等CVPR 2022 · 被引用 179 次
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-trainingHongwei Xue, Yupan Huang, Bei Liu, Houwen Peng 等NeurIPS 2021 · 被引用 100 次
- Auto-Parsing Network for Image Captioning and Visual Question AnsweringXu Yang, Chongyang Gao, Hanwang Zhang, Jianfei CaiICCV 2021 · 被引用 45 次
相关 Paper
- Remodeling Semantic Relationships in Vision-Language Fine-TuningXiangyang Wu, Liu Liu, Baosheng Yu, Jiayan Qiu 等AAAI 2026
- Learning Relation Alignment for Calibrated Cross-modal RetrievalShuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men 等ACL 2021
- Variational Adapter for Cross-modal Similarity RepresentationWenZhang Wei, Zhipeng Gui, Dehua Peng, Tiandi Ye 等ICML 2026 · 被引用 1 次
- Conceptual and Syntactical Cross-modal Alignment with Cross-level Consistency for Image-Text MatchingPengpeng Zeng, Lianli Gao, Xinyu Lyu, Shuaiqi Jing 等ACM MM 2021 · 被引用 37 次
- Cross Modal Fine-grained Alignment via Granularity-aware and Region-uncertain ModelingJiale Liu, Haoming Zhou, Yishu Liu, Bingzhi Chen 等AAAI 2026
