Are Multimodal Transformers Robust to Missing Modality?
Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine, Xi Peng
摘要
Multimodal data collected from the real world are often imperfect due to missing modalities. Therefore multimodal models that are robust against modal-incomplete data are highly preferred. Recently, Transformer models have shown great success in processing multimodal data. However, existing work has been limited to either architecture designs or pre-training strategies; whether Transformer models are naturally robust against missing-modal data has rarely been investigated. In this paper, we present the first-of-its-kind work to comprehensively investigate the behavior of Transformers in the presence of modal-incomplete data. Unsurprising, we find Transformer models are sensitive to missing modalities while different modal fusion strategies will significantly affect the robustness. What surprised us is that the optimal fusion strategy is dataset dependent even for the same Transformer model; there does not exist a universal strategy that works in general cases. Based on these findings, we propose a principle method to improve the robustness of Transformer models byautomatically searching for an optimal fusion strategy regarding input data. Experimental validations on three benchmarks support the superior performance of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper60
- Rethinking Vision Transformers for MobileNet Size and SpeedYanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis 等ICCV 2023 · 被引用 300 次
- Boosting Multi-modal Model Performance with Adaptive Gradient ModulationHong Li, Xingyu Li, Pengbo Hu, Yinuo Lei 等ICCV 2023 · 被引用 84 次
- Towards Good Practices for Missing Modality Robust Action RecognitionSangmin Woo, Sumin Lee, Yeonju Park, Muhammad Adi Nugroho 等AAAI 2023 · 被引用 80 次
- Multimodal Patient Representation Learning with Missing Modalities and LabelsZhenbang Wu, Anant Dadu, Nicholas J. Tustison, Brian B. Avants 等ICLR 2024 · 被引用 38 次
- Deep Correlated Prompting for Visual Recognition with Missing ModalitiesLianyu Hu, Tongkai Shi, Wei Feng, Fanhua Shang 等NeurIPS 2024 · 被引用 37 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
相关 Paper
- Defending Multimodal Fusion Models Against Single-Source AdversariesKarren Yang, Wan-Yi Lin, Manash Barman, Filipe Condessa 等CVPR 2021
- Tag-assisted Multimodal Sentiment Analysis under Uncertain Missing ModalitiesJiandian Zeng, Tianyi Liu, Jiantao ZhouSIGIR 2022 · 被引用 84 次
- Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing ModalitiesJinming Zhao, Ruichen Li, Qin JinACL 2021
- Transformer-based Feature Reconstruction Network for Robust Multimodal Sentiment AnalysisZiqi Yuan, Wei Li, Hua Xu, Wenmeng YuACM MM 2021 · 被引用 186 次
- Robust Modality-Incomplete Anomaly Detection: A Modality-Instructive Framework with BenchmarkBingchen Miao, Wenqiao Zhang, Juncheng Li, Wangyu Wu 等ACM MM 2025 · 被引用 1 次
