Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection
Ankan Mullick, Saransh Sharma, Abhik Jana, Pawan Goyal
摘要
The rise of multimodal data, integrating text, audio, and visuals, has created new opportunities for studying multimodal tasks such as intent detection.This work investigates the effectiveness of Large Language Models (LLMs) and non-LLMs, including text-only and multimodal models, in the multimodal intent detection task.Our study reveals that Mistral-7B, a text-only LLM, outperforms most competitive multimodal models by approximately 9% on MIntRec-1 and 4% on MIntRec2.0dataset.This performance advantage comes from a strong textual bias in these datasets, where over 90% of the samples require textual input, either alone or in combination with other modalities, for correct classification.We confirm the modality bias of these datasets via human evaluation, too.Next, we propose a framework to debias the datasets, and upon debiasing, more than 70% of the samples in MIntRec-1 and more than 50% in MIntRec2.0get removed, resulting in significant performance degradation across all models, with smaller multimodal fusion models being the most affected with an accuracy drop of over 50 -60%.Further, we analyze the context-specific relevance of different modalities through empirical analysis.Our findings highlight the challenges posed by modality bias in multimodal intent datasets and emphasize the need for unbiased datasets to evaluate multimodal models effectively.We release both the code and the dataset used for this work. 1Model M-1 M-2.0 [Approach] Acc F1 Acc F1 BERT [A] 70.8 67.4 57.1 49.3 M-7B [B] 82.9 82.5 65.2 64.4 L2-7B [B] 79.3 79.5 57.1 56.7 Q-7B [B] 72.4 64.6 61.6 62.4 L3-8B [B] 77.3 77.1 61.3 59.9 L2-13B [B] 80.7 80.4 54.8 54.3 Mult [C] 71.5 67.9 58.4 51.5 MAG [C] 72.7 68.6 58.2 49.4 MISA [C] 71.8 69.1 57.8 51.9 SDIF [C] 72.8 71.6 58.6 52.5 ClaudeT [D] 57.7 56.4 39.7 38.1 ClaudeV [D] 59.1 55.9 40.9 39.2 GPT-4T [D] 60.4 59.8 42.1 41.8 GPT-4V [D] 59.5 58.8 41.8 41.4 VLLaMA [D] 23.6 23.4 11.8 11.0 VLLaVA [D] 32.4 30.5 18.8 17.5 VChatGPT [D] 33.3 33.0 18.7 16.8
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
相关 Paper
- MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in ConversationsHanlei Zhang, Xin Wang, Hua Xu, Qianrui Zhou 等ICLR 2024 · 被引用 29 次
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou 等ACM MM 2022 · 被引用 66 次
- Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language ModelsJean Park, Kuk Jin Jang, Basam Alasaly, Sriharsha Mopidevi 等AAAI 2025 · 被引用 21 次
- Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face DetectorXiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu 等CVPR 2025
- Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal SettingsMd Messal Monem Miah, Adrita Anika, Xi Shi, Ruihong HuangACL 2025
