Mitigating Language Bias of LMMs in Social Intelligence Understanding with Virtual Counterfactual Calibration
Peng Chen, Xiao-Yu Guo, Yuan-Fang Li, Xiaowang Zhang, Zhiyong Feng
Abstract
Social intelligence is essential for understanding complex human expressions and social interactions. While large multimodal models (LMMs) have demonstrated remarkable performance in social intelligence question answering (SIQA), they are still inclined to generate responses relying on language priors and ignoring the relevant context due to the dominant prevalence of text-based data in the pretraining stage. To interpret the aforementioned language bias of LMMs, we employ a structure causal model and posit that counterfactual reasoning can mitigate the bias by avoiding spurious correlations between LMMs' internal commonsense knowledge and the given context. However, it is costly and challenging to construct multimodal counterfactual samples. To tackle the above challenges, we propose an output Distribution Calibration network with Virtual Counterfactual (DCVC) data augmentation framework. DCVC devises a novel output distribution calibration network to mitigate the impact of negative language biases while preserving beneficial priors. Perturbations are introduced to the output distributions of LMMs to simulate the distribution shifts from counterfactual manipulations of the context, which is employed to construct counterfactual augmented data virtually. Experiments on multiple datasets demonstrate the effectiveness and generalizability of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af2df570-6f19-44c7-a5a4-786c05df13f9Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun et al.AAAI 2021 · 414 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- HiTeA: Hierarchical Temporal-Aware Video-Language Pre-trainingQinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu et al.ICCV 2023 · 102 citations
- FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningJing Zhou, Yanan Zheng, Jie Tang, Li Jian et al.ACL 2022 · 91 citations
Related papers
- Counterfactual Vision and Language LearningEhsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi et al.CVPR 2020
- Multi-Level Counterfactual Contrast for Visual Commonsense ReasoningXi Zhang, Feifei Zhang, Changsheng XuACM MM 2021 · 22 citations
- Counterfactual VQA: A Cause-Effect Look at Language BiasYulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu et al.CVPR 2021
- Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention CausalityGuanyu Zhou, Yibo Yan, Xin Zou, Kun Wang et al.ICLR 2025
- Social Debiasing for Fair Multi-Modal LLMsHarry Cheng, Yangyang Guo, Qing Guo, Ming-Hsuan Yang et al.ICCV 2025 · 1 citation
