Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content Dilutions
Gaurav Verma, Vishwa Vinay, Ryan A. Rossi, Srijan Kumar
Abstract
As multimodal learning finds applications in a wide variety of high-stakes societal tasks, investigating their robustness becomes important. Existing work has focused on understanding the robustness of vision-and-language models to imperceptible variations on benchmark tasks. In this work, we investigate the robustness of multimodal classifiers to cross-modal dilutions – a plausible variation. We develop a model that, given a multimodal (image + text) input, generates additional dilution text that (a) maintains relevance and topical coherence with the image and existing text, and (b) when added to the original text, leads to misclassification of the multimodal input. Via experiments on Crisis Humanitarianism and Sentiment Detection tasks, we find that the performance of task-specific fusion-based multimodal classifiers drops by 23.3% and 22.5%, respectively, in the presence of dilutions generated by our model. Metric-based comparisons with several baselines and human evaluations indicate that our dilutions show higher relevance and topical coherence, while simultaneously being more effective at demonstrating the brittleness of the multimodal classifiers. Our work aims to highlight and encourage further research on the robustness of deep multimodal models to realistic variations, especially in human-facing societal applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language LearningShivaen Ramshetty, Gaurav Verma, Srijan KumarACL 2023 · 4 citations
- Learning the Visualness of Text Using Large Vision-Language ModelsGaurav Verma, Ryan A. Rossi, Christopher Tensmeyer, Jiuxiang Gu et al.EMNLP 2023 · 2 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami et al.NeurIPS 2020 · 1,022 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Human-Adversarial Visual Question AnsweringSasha Sheng, Amanpreet Singh, Vedanuj Goswami, Jose Alberto Lopez Magana et al.NeurIPS 2021 · 81 citations
- POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-trainingYizhe Zhang, Guoyin Wang, Chunyuan Li, Zhe Gan et al.EMNLP 2020 · 68 citations
Related papers
- Multimodal Categorization of Crisis Events in Social MediaMahdi Abavisani, Liwei Wu, Shengli Hu, Joel R. Tetreault et al.CVPR 2020
- Defending Multimodal Fusion Models Against Single-Source AdversariesKarren Yang, Wan-Yi Lin, Manash Barman, Filipe Condessa et al.CVPR 2021
- Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and FusionDaiqing Wu, Dongbao Yang, Yu Zhou, Can MaACM MM 2024 · 13 citations
- Improving Multimodal fusion via Mutual Dependency MaximisationPierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé ClavelEMNLP 2021 · 29 citations
- TranX-Adapter: Bridging Artifacts and Semantics within MLLMs for Robust AI-generated Image DetectionWenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu et al.ICML 2026 · 1 citation
