Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
Yucheng Wang, Yifan Hou, Aydin Javadov, Mubashara Akhtar, Mrinmaya Sachan
摘要
Multimodal large language models (MLLMs) promise enhanced reasoning by integrating diverse inputs such as text, vision, and audio. Yet, despite their perceptual strengths, their ability to reason across modalities remains underexplored, with conflicting reports on whether additional modalities help or harm performance. These inconsistencies stem from a lack of controlled evaluation frameworks and analysis of models' internals to isolate when and why modality interactions support or undermine reasoning. We address this gap through a logic-based evaluation framework that categorizes multimodal reasoning into six interaction patterns, varying how factual information is distributed across modalities and logically combined. Empirically, additional modalities enhance reasoning only when they provide independent and sufficient reasoning paths, while redundant or chained entailment support in extra modalities often hurts performance. In addition, models recognize cross-modal facts reliably and always reason on text effectively. Moreover, reasoning is degraded in three systematic ways: weaker modalities drag down overall performance, conflicts bias preference toward certain modalities, and joint signals from different modalities fail to be integrated effectively. Therefore, we identify two core failures: task-composition bottleneck, where recognition and reasoning cannot be jointly executed in one pass, and fusion bottleneck, where early integration introduces bias. For further investigation, we find that attention patterns fail to encode fact usefulness, but a simple two-step prompting (recognize then reason) restores performance, confirming the task-composition bottleneck. Moreover, modality identity remains recoverable in early layers, and softening attention in early fusion improves reasoning, highlighting biased fusion as another failure mode. In general, our findings show that integration, not perception, is the main barrier to multimodal reasoning, suggesting composition-aware training and early fusion control as promising directions. 1 * Equal contribution 1 Our code and data are publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual SimulationsLinjie Li, Mahtab Bigverdi, Jiawei Gu, Zixian Ma 等ICLR 2026 · 被引用 22 次
- MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction ExpertsHaofei Yu, Zhengyang Qi, Lawrence Jang, Russ Salakhutdinov 等EMNLP 2024 · 被引用 11 次
- Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning BenchmarkYunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li 等ICML 2025
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li 等CVPR 2025
相关 Paper
- HAVE-Bench: Hierarchical Audio-Visual Evaluation from Perception to InteractionZhong Muyan, Erfei Cui, Sen Xing, Weiyun Wang 等CVPR 2026
- Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information DecompositionWanlong Fang, Tianle Zhang, Wen Tao, Alvin ChanICML 2026 · 被引用 17 次
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu 等ICML 2026
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMsZheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li 等AAAI 2026 · 被引用 2 次
- Bring Reason to Vision: Understanding Perception and Reasoning through Model MergingShiqi Chen, Jinghan Zhang, Tongyao Zhu, Wei Liu 等ICML 2025
