Deciphering Cross-Modal Alignment in Large Vision-Language Models Via Modality Integration Rate
Qidong Huang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Jiaqi Wang, Weiming Zhang, Nenghai Yu
摘要
The early stage of multi-modal pre-training plays a pivotal role in aligning two modalities for Large Vision-Language Models (LVLMs), while evaluating its training quality usually requires the costly supervised fine-tuning (SFT) stage to verify the downstream benchmark scores. Loss, perplexity, and in-context evaluation results are commonly used pre-training metrics for Large Language Models (LLMs), while we observed that these metrics are less indicative when quantifying the pre-trained LVLMs. Due to the lack of proper metrics, the research of LVLMs in the multi-modal fusion stage is hindered greatly, including the training data choice, efficient module design, etc. In this paper, we first present Modality Integration Rate (MIR), an effective, robust, and generalized metric to indicate the multimodal alignment quality of LVLMs without SFT. This metric evaluates LVLM pre-training from the inter-modal distribution distance perspective, which is 1) Effective to represent the fusion quality and show a positive relation with the benchmark performance after SFT, 2) Robust toward different training/evaluation data, and 3) Generalize across training configurations and architecture choices. Complementing MIR, we further propose learnable Modality Calibration (MoCa), a lightweight module to narrow the modality gap at each language model layer during training. A series of experiments are conducted to explore the effectiveness of MIR and MoCa, demonstrating that MIR is highly indicative about training data selection, training strategy schedule, and architecture design to improve pre-training. The code is at: shikiw/Modality-Integration-Rate.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SineProject: Machine Unlearning for Stable Vision-Language AlignmentArpit Garg, Hemanth Saratchandran, Simon LuceyCVPR 2026 · 被引用 2 次
- One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs HallucinationZhan Fa, Yue Duan, Jian Zhang, Lei Qi 等CVPR 2026 · 被引用 2 次
- Deep Pre-Alignment for VLMsTianyu Yu, Kechen Fang, Zihao Wan, Kaidong Zhang 等ICML 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- Learning Relation Alignment for Calibrated Cross-modal RetrievalShuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men 等ACL 2021
- Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein DistanceMuyang Li, Yucheng Liu, Jianbo Ma, Elliot Osborne 等CVPR 2026 · 被引用 2 次
- MoRA: Missing Modality Low-Rank Adaptation for Visual RecognitionShu Zhao, Nilesh A. Ahuja, Tan Yu, Tianyi Shen 等ICLR 2026 · 被引用 5 次
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal EmbeddingsHaonan Chen, Hong Liu, Yuping Luo, Liang Wang 等ACL 2026 · 被引用 20 次
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information LossBozhou Li, Xinda Xue, Sihan Yang, Yang Shi 等ICLR 2026 · 被引用 5 次
