VLM-PTQ: Efficient Post-Training Quantization for Large Vision-Language Models
Juncan Deng, Kejie Huang
摘要
Post-training quantization (PTQ) serves as a vital technique for efficiently compressing large-scale models, with weight-compensation methods such as GPTQ (symmetric calibration) and GPTAQ (asymmetric calibration) showing remarkable success. However, directly applying these methods to Vision-Language Models (VLMs) exposes two notable shortcomings: 1) the standard rounding-to-nearest (RTN) method is suboptimal for the asymmetric objective, failing to account for residual-induced shifts in the optimal quantization target; and 2) all input channels are processed uniformly across modalities, overlooking the distinct information densities. In this paper, we introduce VLM-PTQ, an asymmetric post-training quantization framework for VLMs. First, we derive a closed-form correction term that shifts the quantization target, which explicitly accounts for the output residual and the corresponding inverse Hessian column, yielding a better local optimum than RTN. Second, we propose a modality-aware quantization that differentiates channel importance between vision and language tokens, allowing the quantizer to pre-compute better quantization parameters through a lightweight search. Our method extends weight-compensation methods with minimal overhead while achieving significant performance improvements in low-bit scenarios. Extensive experiments demonstrate that VLM-PTQ achieves competitive results compared to existing methods, effectively compressing models from 1B to 72B parameters on a single GPU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models QuantizationChenwei Jia, Baoting Li, Xuchong Zhang, Mingzhuo Wei 等CVPR 2026 · 被引用 3 次
- MBQ: Modality-Balanced Quantization for Large Vision-Language ModelsShiyao Li, Yingchun Hu, Xuefei Ning, Xihui Liu 等CVPR 2025
- ImpQuant: Fine-Grained Importance-Aware Quantization for Large Vision-Language ModelsJundong Zhou, Tianao Cai, Yujie Huang, Xinbing Wang 等ICML 2026 · 被引用 2 次
- Rethinking Residual Errors in Compensation-based LLM QuantizationShuaiting Li, Juncan Deng, Kedong Xu, Rongtao Deng 等ICLR 2026 · 被引用 4 次
- GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric CalibrationYuhang Li, Ruokai Yin, Donghyun Lee, Shiting Xiao 等ICML 2025
