VLM-PTQ: Efficient Post-Training Quantization for Large Vision-Language Models
Juncan Deng, Kejie Huang
Abstract
Post-training quantization (PTQ) serves as a vital technique for efficiently compressing large-scale models, with weight-compensation methods such as GPTQ (symmetric calibration) and GPTAQ (asymmetric calibration) showing remarkable success. However, directly applying these methods to Vision-Language Models (VLMs) exposes two notable shortcomings: 1) the standard rounding-to-nearest (RTN) method is suboptimal for the asymmetric objective, failing to account for residual-induced shifts in the optimal quantization target; and 2) all input channels are processed uniformly across modalities, overlooking the distinct information densities. In this paper, we introduce VLM-PTQ, an asymmetric post-training quantization framework for VLMs. First, we derive a closed-form correction term that shifts the quantization target, which explicitly accounts for the output residual and the corresponding inverse Hessian column, yielding a better local optimum than RTN. Second, we propose a modality-aware quantization that differentiates channel importance between vision and language tokens, allowing the quantizer to pre-compute better quantization parameters through a lightweight search. Our method extends weight-compensation methods with minimal overhead while achieving significant performance improvements in low-bit scenarios. Extensive experiments demonstrate that VLM-PTQ achieves competitive results compared to existing methods, effectively compressing models from 1B to 72B parameters on a single GPU.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b991303a-a236-445a-92d9-194dff516df2Builds on22
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models QuantizationChenwei Jia, Baoting Li, Xuchong Zhang, Mingzhuo Wei et al.CVPR 2026 · 3 citations
- MBQ: Modality-Balanced Quantization for Large Vision-Language ModelsShiyao Li, Yingchun Hu, Xuefei Ning, Xihui Liu et al.CVPR 2025
- ImpQuant: Fine-Grained Importance-Aware Quantization for Large Vision-Language ModelsJundong Zhou, Tianao Cai, Yujie Huang, Xinbing Wang et al.ICML 2026 · 2 citations
- Rethinking Residual Errors in Compensation-based LLM QuantizationShuaiting Li, Juncan Deng, Kedong Xu, Rongtao Deng et al.ICLR 2026 · 4 citations
- GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric CalibrationYuhang Li, Ruokai Yin, Donghyun Lee, Shiting Xiao et al.ICML 2025
