Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
Jean Park, Kuk Jin Jang, Basam Alasaly, Sriharsha Mopidevi, Andrew Zolensky, Eric Eaton, Insup Lee, Kevin B. Johnson
Abstract
Multimodal large language models (MLLMs) can simultaneously process visual, textual, and auditory data, capturing insights that complement human analysis. However, existing video question-answering (VidQA) benchmarks and datasets often exhibit a bias toward a single modality, despite the goal of requiring advanced reasoning skills that integrate diverse modalities to answer the queries.
In this work, we introduce the modality importance score (MIS) to identify such bias. It is designed to assess which modality embeds the necessary information to answer the question. Additionally, we propose an innovative method using state-of-the-art MLLMs to estimate the modality importance, which can serve as a proxy for human judgments of modality perception. With this MIS, we demonstrate the presence of unimodal bias and the scarcity of genuinely multimodal questions in existing datasets. We further validate the modality importance score with multiple ablation studies to evaluate the performance of MLLMs on permuted feature sets. Our results indicate that current models do not effectively integrate information due to modality imbalance in existing datasets. Our proposed MLLM-derived MIS can guide the curation of modality-balanced datasets that advance multimodal learning and enhance MLLMs' capabilities to understand and utilize synergistic relations across modalities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d60b54f-eadf-4e2f-98ea-39617b641f75Cited by top-tier papers8
- Multi-speaker Attention Alignment for Multimodal Social InteractionLiangyang Ouyang, Yifei Huang, Mingfang Zhang, Caixin Kang et al.CVPR 2026 · 8 citations
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXingrui Wang, Jiang Liu, Chao Huang, Xiaodong Yu et al.ICLR 2026 · 4 citations
- Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMsJingze Wu, Quan Zhang, Hongfei Suo, Zeqiang Cai et al.CVPR 2026 · 2 citations
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language ModelsHengzhuang Li, Xinsong Zhang, QIMING PENG, Bin Luo et al.CVPR 2026 · 2 citations
- PharmaQA: Prompt-Based Molecular Representation Learning via Pharmacophore-Oriented Question AnsweringChengwei Ai, Qiaozhen Meng, Mengwei Sun, Ruihan Dong et al.AAAI 2026
Builds on11
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao et al.AAAI 2020 · 129 citations
- AVQA: A Dataset for Audio-Visual Question Answering on VideosPinci Yang, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 60 citations
- Perceptual Score: What Data Modalities Does Your Model Perceive?Itai Gat, Idan Schwartz, Alexander G. SchwingNeurIPS 2021 · 56 citations
- Self-supervised Pre-training and Contrastive Representation Learning for Multiple-choice Video QASeonhoon Kim, Seohyeong Jeong, Eunbyul Kim, Inho Kang et al.AAAI 2021 · 44 citations
- Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringDohwan Ko, Ji Soo Lee, Woo-Young Kang, Byungseok Roh et al.EMNLP 2023 · 30 citations
Related papers
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalDivyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit ChopraICLR 2026 · 3 citations
- How Can Objects Help Video-Language Understanding?Zitian Tang, Shijie Wang, Junho Cho, Jaewook Yoo et al.ICCV 2025 · 8 citations
- Text Takes Over: A Study of Modality Bias in Multimodal Intent DetectionAnkan Mullick, Saransh Sharma, Abhik Jana, Pawan GoyalEMNLP 2025
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li et al.CVPR 2025
- Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response TheoryShunki Uebayashi, Kento Masui, Kyohei Atarashi, Han Bao et al.ICLR 2026 · 1 citation
