MBQ: Modality-Balanced Quantization for Large Vision-Language Models
Shiyao Li, Yingchun Hu, Xuefei Ning, Xihui Liu, Ke Hong, Xiaotao Jia, Xiuhong Li, Yaqi Yan, Pei Ran, Guohao Dai, Shengen Yan, Huazhong Yang, Yu Wang
摘要
Vision-Language Models (VLMs) have enabled a variety of real-world applications. The large parameter size of VLMs brings large memory and computation overhead which poses significant challenges for deployment. Post-Training Quantization (PTQ) is an effective technique to reduce the memory and computation overhead. Existing PTQ methods mainly focus on large language models (LLMs), without considering the differences across other modalities. In this paper, we discover that there is a significant difference in sensitivity between language and vision tokens in large VLMs. Therefore, treating tokens from different modalities equally, as in existing PTQ methods, may over-emphasize the insensitive modalities, leading to significant accuracy loss. To deal with the above issue, we propose a simple yet effective method, Modality-Balanced Quantization (MBQ), for large VLMs. Specifically, MBQ incorporates the different sensitivities across modalities during the calibration process to minimize the reconstruction loss for better quantization parameters. Extensive experiments show that MBQ can significantly improve task accuracy by up to 4.4% and 11.6% under W3A16 and W4A8 quantization for 7B to 70B VLMs, compared to SOTA baselines. Additionally, we implement a W3A16 GPU kernel that fuses the dequantization and GEMV operators, achieving a 1.4→ speedup on LLaVAonevision-7B on the RTX 4090. The code is available at https://github.com/thu-nics/MBQ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation ModelsTianchen Zhao, Ke Hong, Xinhao Yang, Xuefeng Xiao 等NeurIPS 2025 · 被引用 19 次
- IAG: Input-aware Backdoor Attack on VLM-based Visual GroundingJunxian Li, Beining Xu, Simin Chen, Jiatong Li 等CVPR 2026 · 被引用 13 次
- MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static QuantizationJiangyong Yu, Sifan Zhou, Dawei Yang, Shuoyu Li 等ACM MM 2025 · 被引用 11 次
- Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated GradientsZiwei Xiang, Fanhu Zeng, Hongjian Fang, Rui-Qi Wang 等CVPR 2026 · 被引用 7 次
- Gated Relational Alignment via Confidence-based Distillation for Efficient VLMsYanlong Chen, Amir Habibian, Luca Benini, Yawei LiICML 2026 · 被引用 5 次
它引用的顶会 Paper22
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
相关 Paper
- ImpQuant: Fine-Grained Importance-Aware Quantization for Large Vision-Language ModelsJundong Zhou, Tianao Cai, Yujie Huang, Xinbing Wang 等ICML 2026 · 被引用 2 次
- VLM-PTQ: Efficient Post-Training Quantization for Large Vision-Language ModelsJuncan Deng, Kejie HuangCVPR 2026 · 被引用 2 次
- Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models QuantizationChenwei Jia, Baoting Li, Xuchong Zhang, Mingzhuo Wei 等CVPR 2026 · 被引用 3 次
- VEQ: Modality-Adaptive Quantization for MoE Vision-Language ModelsGuangshuo Qin, Zhiteng Li, Zheng Chen, Weihang Zhang 等ICML 2026 · 被引用 1 次
- Q-VLM: Post-training Quantization for Large Vision-Language ModelsChangyuan Wang, Ziwei Wang, Xiuwei Xu, Yansong Tang 等NeurIPS 2024 · 被引用 57 次
