Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning
Shivaen Ramshetty, Gaurav Verma, Srijan Kumar
Abstract
The robustness of multimodal deep learning models to realistic changes in the input text is critical for their applicability to important tasks such as text-to-image retrieval and cross-modal entailment. To measure robustness, several existing approaches edit the text data, but do so without leveraging the cross-modal information present in multimodal data. Information from the visual modality, such as color, size, and shape, provide additional attributes that users can include in their inputs. Thus, we propose cross-modal attribute insertions as a realistic perturbation strategy for vision-and-language data that inserts visual attributes of the objects in the image into the corresponding text (e.g., "girl on a chair" → "little girl on a wooden chair"). Our proposed approach for cross-modal attribute insertions is modular, controllable, and task-agnostic. We find that augmenting input text using cross-modal insertions causes state-of-the-art approaches for text-to-image retrieval and cross-modal entailment to perform poorly, resulting in relative drops of ∼ 15% in MRR and ∼ 20% in F 1 score, respectively. Crowd-sourced annotations demonstrate that cross-modal insertions lead to higher quality augmentations for multimodal data than augmentations using text-only data, and are equivalent in quality to original examples. We release the code to encourage robustness evaluations of deep vision-and-language models: https://github.com/claws-lab/ multimodal-robustness-xmai .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Empirical Study of Training End-to-End Vision-and-Language TransformersZi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang et al.CVPR 2022 · 313 citations
- Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for TextNishtha Madaan, Inkit Padhi, Naveen Panwar, Diptikalyan SahaAAAI 2021 · 115 citations
- Tailor: Generating and Perturbing Text with Semantic ControlsAlexis Ross, Tongshuang Wu, Hao Peng, Matthew E. Peters et al.ACL 2022 · 85 citations
- Human-Adversarial Visual Question AnsweringSasha Sheng, Amanpreet Singh, Vedanuj Goswami, Jose Alberto Lopez Magana et al.NeurIPS 2021 · 81 citations
Related papers
- Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content DilutionsGaurav Verma, Vishwa Vinay, Ryan A. Rossi, Srijan KumarEMNLP 2022 · 9 citations
- VisMoDAI: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language ModelsHuanchen Wang, Wencheng Zhang, Zhiqiang Wang, Zhicong Lu et al.IEEE VIS 2025 · 1 citation
- Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalFeifei Zhang, Mingliang Xu, Qirong Mao, Changsheng XuACM MM 2020 · 39 citations
- On Evaluating the Robustness of Large Vision-Language Models via Untargeted Modality Alignment Breaking Adversarial AttackZhichao Li, Hongshan Yang, Zhibo Wang, Huiyu Xu et al.USENIX Security 2026
- Is the Modality Gap a Bug or a Feature? A Robustness PerspectiveRhea Chowers, Oshri Naparstek, Udi Barzelay, Yair WeissCVPR 2026 · 4 citations
