Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content Dilutions
Gaurav Verma, Vishwa Vinay, Ryan A. Rossi, Srijan Kumar
摘要
As multimodal learning finds applications in a wide variety of high-stakes societal tasks, investigating their robustness becomes important. Existing work has focused on understanding the robustness of vision-and-language models to imperceptible variations on benchmark tasks. In this work, we investigate the robustness of multimodal classifiers to cross-modal dilutions – a plausible variation. We develop a model that, given a multimodal (image + text) input, generates additional dilution text that (a) maintains relevance and topical coherence with the image and existing text, and (b) when added to the original text, leads to misclassification of the multimodal input. Via experiments on Crisis Humanitarianism and Sentiment Detection tasks, we find that the performance of task-specific fusion-based multimodal classifiers drops by 23.3% and 22.5%, respectively, in the presence of dilutions generated by our model. Metric-based comparisons with several baselines and human evaluations indicate that our dilutions show higher relevance and topical coherence, while simultaneously being more effective at demonstrating the brittleness of the multimodal classifiers. Our work aims to highlight and encourage further research on the robustness of deep multimodal models to realistic variations, especially in human-facing societal applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language LearningShivaen Ramshetty, Gaurav Verma, Srijan KumarACL 2023 · 被引用 4 次
- Learning the Visualness of Text Using Large Vision-Language ModelsGaurav Verma, Ryan A. Rossi, Christopher Tensmeyer, Jiuxiang Gu 等EMNLP 2023 · 被引用 2 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Human-Adversarial Visual Question AnsweringSasha Sheng, Amanpreet Singh, Vedanuj Goswami, Jose Alberto Lopez Magana 等NeurIPS 2021 · 被引用 81 次
- POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-trainingYizhe Zhang, Guoyin Wang, Chunyuan Li, Zhe Gan 等EMNLP 2020 · 被引用 68 次
相关 Paper
- Multimodal Categorization of Crisis Events in Social MediaMahdi Abavisani, Liwei Wu, Shengli Hu, Joel R. Tetreault 等CVPR 2020
- Defending Multimodal Fusion Models Against Single-Source AdversariesKarren Yang, Wan-Yi Lin, Manash Barman, Filipe Condessa 等CVPR 2021
- Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and FusionDaiqing Wu, Dongbao Yang, Yu Zhou, Can MaACM MM 2024 · 被引用 13 次
- Improving Multimodal fusion via Mutual Dependency MaximisationPierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé ClavelEMNLP 2021 · 被引用 29 次
- TranX-Adapter: Bridging Artifacts and Semantics within MLLMs for Robust AI-generated Image DetectionWenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu 等ICML 2026 · 被引用 1 次
