Do Vision-and-Language Transformers Learn Grounded Predicate-Noun Dependencies?
Mitja Nikolaus, Emmanuelle Salin, Stéphane Ayache, Abdellah Fourtassi, Benoît Favre
摘要
Recent advances in vision-and-language modeling have seen the development of Transformer architectures that achieve remarkable performance on multimodal reasoning tasks. Yet, the exact capabilities of these black-box models are still poorly understood. While much of previous work has focused on studying their ability to learn meaning at the word-level, their ability to track syntactic dependencies between words has received less attention. We take a first step in closing this gap by creating a new multimodal task targeted at evaluating understanding of predicate-noun dependencies in a controlled setup. We evaluate a range of state-of-the-art models and find that their performance on the task varies considerably, with some models performing relatively well and others at chance level. In an effort to explain this variability, our analyses indicate that the quality (and not only sheer quantity) of pretraining data is essential. Additionally, the best performing models leverage fine-grained multimodal pretraining objectives in addition to the standard image-text matching objectives. This study highlights that targeted and controlled evaluations are a crucial step for a precise and rigorous test of the multimodal knowledge of vision-and-language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Measuring Progress in Fine-grained Vision-and-Language UnderstandingEmanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks 等ACL 2023 · 被引用 9 次
- MASS: Overcoming Language Bias in Image-Text MatchingJiwan Chung, Seungwon Lim, Sangkyu Lee, Youngjae YuAAAI 2025 · 被引用 1 次
- Weakly-Supervised Learning of Visual Relations in Multimodal PretrainingEmanuele Bugliarello, Aida Nematzadeh, Lisa Anne HendricksEMNLP 2023 · 被引用 1 次
- Extract Free Dense Misalignment from CLIPJeongYeon Nam, Jinbae Im, Wonjae Kim, Taeho KilAAAI 2025
它引用的顶会 Paper9
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh 等CVPR 2022 · 被引用 179 次
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language TechnologiesSunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian 等EMNLP 2021 · 被引用 113 次
- Visually Grounded Reasoning across Languages and CulturesFangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy 等EMNLP 2021 · 被引用 87 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
相关 Paper
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-trainingHongwei Xue, Yupan Huang, Bei Liu, Houwen Peng 等NeurIPS 2021 · 被引用 100 次
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalDivyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit ChopraICLR 2026 · 被引用 3 次
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- Are Vision-Language Transformers Learning Multimodal Representations? A Probing PerspectiveEmmanuelle Salin, Badreddine Farah, Stéphane Ayache, Benoît FavreAAAI 2022 · 被引用 49 次
- Investigating Compositional Challenges in Vision-Language Models for Visual GroundingYunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie 等CVPR 2024 · 被引用 4 次
