Lune

EMNLP2025Top-tier venue

Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection

Ankan Mullick, Saransh Sharma, Abhik Jana, Pawan Goyal

2025Year

Abstract

The rise of multimodal data, integrating text, audio, and visuals, has created new opportunities for studying multimodal tasks such as intent detection.This work investigates the effectiveness of Large Language Models (LLMs) and non-LLMs, including text-only and multimodal models, in the multimodal intent detection task.Our study reveals that Mistral-7B, a text-only LLM, outperforms most competitive multimodal models by approximately 9% on MIntRec-1 and 4% on MIntRec2.0dataset.This performance advantage comes from a strong textual bias in these datasets, where over 90% of the samples require textual input, either alone or in combination with other modalities, for correct classification.We confirm the modality bias of these datasets via human evaluation, too.Next, we propose a framework to debias the datasets, and upon debiasing, more than 70% of the samples in MIntRec-1 and more than 50% in MIntRec2.0get removed, resulting in significant performance degradation across all models, with smaller multimodal fusion models being the most affected with an accuracy drop of over 50 -60%.Further, we analyze the context-specific relevance of different modalities through empirical analysis.Our findings highlight the challenges posed by modality bias in multimodal intent datasets and emphasize the need for unbiased datasets to evaluate multimodal models effectively.We release both the code and the dataset used for this work. 1Model M-1 M-2.0 [Approach] Acc F1 Acc F1 BERT [A] 70.8 67.4 57.1 49.3 M-7B [B] 82.9 82.5 65.2 64.4 L2-7B [B] 79.3 79.5 57.1 56.7 Q-7B [B] 72.4 64.6 61.6 62.4 L3-8B [B] 77.3 77.1 61.3 59.9 L2-13B [B] 80.7 80.4 54.8 54.3 Mult [C] 71.5 67.9 58.4 51.5 MAG [C] 72.7 68.6 58.2 49.4 MISA [C] 71.8 69.1 57.8 51.9 SDIF [C] 72.8 71.6 58.6 52.5 ClaudeT [D] 57.7 56.4 39.7 38.1 ClaudeV [D] 59.1 55.9 40.9 39.2 GPT-4T [D] 60.4 59.8 42.1 41.8 GPT-4V [D] 59.5 58.8 41.8 41.4 VLLaMA [D] 23.6 23.4 11.8 11.0 VLLaVA [D] 32.4 30.5 18.8 17.5 VChatGPT [D] 33.3 33.0 18.7 16.8

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines