Perceptual Score: What Data Modalities Does Your Model Perceive?
Itai Gat, Idan Schwartz, Alexander G. Schwing
Abstract
Machine learning advances in the last decade have relied significantly on large-scale datasets that continue to grow in size. Increasingly, those datasets also contain different data modalities. However, large multi-modal datasets are hard to annotate, and annotations may contain biases that we are often unaware of. Deep-net-based classifiers, in turn, are prone to exploit those biases and to find shortcuts. To study and quantify this concern, we introduce the perceptual score, a metric that assesses the degree to which a model relies on the different subsets of the input features, i.e., modalities. Using the perceptual score, we find a surprisingly consistent trend across four popular datasets: recent, more accurate state-of-the-art multi-modal models for visual question-answering or visual dialog tend to perceive the visual data less than their predecessors. This trend is concerning as answers are hence increasingly inferred from textual cues only. Using the perceptual score also helps to analyze model biases by decomposing the score into data subset contributions. We hope to spur a discussion on the perceptiveness of multi-modal models and also hope to encourage the community working on multi-modal classifiers to start quantifying perceptiveness via the proposed perceptual score.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b05dd88-94c5-4d04-9fa2-c329d6194c25Cited by top-tier papers14
- Boosting Multi-modal Model Performance with Adaptive Gradient ModulationHong Li, Xingyu Li, Pengbo Hu, Yinuo Lei et al.ICCV 2023 · 84 citations
- Optimizing Relevance Maps of Vision Transformers Improves RobustnessHila Chefer, Idan Schwartz, Lior WolfNeurIPS 2022 · 55 citations
- Quantifying and Enhancing Multi-modal Robustness with Modality PreferenceZequn Yang, Yake Wei, Ce Liang, Di HuICLR 2024 · 27 citations
- Layer Collaboration in the Forward-Forward AlgorithmGuy Lorberbom, Itai Gat, Yossi Adi, Alexander G. Schwing et al.AAAI 2024 · 22 citations
- Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language ModelsJean Park, Kuk Jin Jang, Basam Alasaly, Sriharsha Mopidevi et al.AAAI 2025 · 21 citations
Builds on3
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional EntropiesItai Gat, Idan Schwartz, Alexander G. Schwing, Tamir HazanNeurIPS 2020 · 111 citations
- Counterfactual Samples Synthesizing for Robust Visual Question AnsweringLong Chen, Xin Yan, Jun Xiao, Hanwang Zhang et al.CVPR 2020
Related papers
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalDivyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit ChopraICLR 2026 · 3 citations
- MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-JudgeSua Lee, Sanghee Park, Jinbae ImACL 2026 · 1 citation
- Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward ModelingSeojeong Park, Jiho Choi, Junyong Kang, Seonho Lee et al.ICML 2026 · 1 citation
- Modality-Balanced Models for Visual DialogueHyounghun Kim, Hao Tan, Mohit BansalAAAI 2020 · 29 citations
- MultiModalQA: complex question answering over text, tables and imagesAlon Talmor, Ori Yoran, Amnon Catav, Dan Lahav et al.ICLR 2021 · 229 citations
