Beyond Fixed Biases: Decoding the Role of Reasoning Uncertainty in MLLM Modality Conflicts
Zhuoran Zhang, Tengyue Wang, Xilin Gong, Yang Shi, Haotian Wang, Di Wang, Lijie Hu
Abstract
Multimodal Large Language Models (MLLMs) must resolve conflicts when modalities provide contradictory information, a behavior we term "modality following". We propose a framework that decomposes this behavior into case-specific relative preference uncertainty and stable inherent preference. Across diverse MLLMs and benchmarks, the probability of following a modality consistently decreases as its relative preference uncertainty increases, a trend robust to alternative uncertainty indices. This regularity defines a "balance point'' where modality preferences are evenly matched, offering a capability-disentangled measure of modality bias. Layer-wise probing further shows that ambiguous cases near the balance point trigger middle-to-late-layer "concept oscillations," where top predictions vacillate between modality-supported answers. Finally, we demonstrate the framework's utility for preference steering through Supervised Fine-Tuning (SFT). We find that data efficiency is governed by preference uncertainty: training on easy samples (where one modality dominates) fails to generalize, whereas targeting the identified ``boundary cases" is essential for robust preference alignment and suppressing internal vacillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5cbe38f-8dc1-48ae-9921-2366680cdea3Builds on16
- Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal TransformersStella Frank, Emanuele Bugliarello, Desmond ElliottEMNLP 2021 · 36 citations
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive BenchmarkYang Shi, Yuhao Dong, Yue Ding, Yuran Wang et al.CVPR 2026 · 35 citations
- EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit IdentificationLin Zhang, Wenshuo Dong, Zhuoran Zhang, Shu Yang et al.NeurIPS 2025 · 26 citations
- Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric PerspectivesShaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima et al.ICCV 2025 · 25 citations
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-trainingJunlin Han, Shengbang Tong, David Fan, Yufan Ren et al.ICLR 2026 · 25 citations
Related papers
- Evaluating and Steering Modality Preferences in Multi-modal LLMsYu Zhang, Jinlong Ma, Yongshuai Hou, Xuefeng Bai et al.ICML 2026
- CoMMIT: Coordinated Multimodal Instruction TuningXintong Li, Junda Wu, Tong Yu, Rui Wang et al.EMNLP 2025
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao et al.ACL 2025
- ESTJ: Enhancing Structured Tendency Judgment in Hybrid-Modal Table UnderstandingShu-Xun Yang, Xian-Ling Mao, Heyan HuangACM MM 2025
- Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading ScenariosYunkai Dang, Mengxi Gao, Yibo Yan, Xin Zou et al.EMNLP 2025 · 1 citation
