Lune

CVPR2026顶会

VLM-PTQ: Efficient Post-Training Quantization for Large Vision-Language Models

Juncan Deng, Kejie Huang

出版方
2026年份
2被引次数

摘要

Post-training quantization (PTQ) serves as a vital technique for efficiently compressing large-scale models, with weight-compensation methods such as GPTQ (symmetric calibration) and GPTAQ (asymmetric calibration) showing remarkable success. However, directly applying these methods to Vision-Language Models (VLMs) exposes two notable shortcomings: 1) the standard rounding-to-nearest (RTN) method is suboptimal for the asymmetric objective, failing to account for residual-induced shifts in the optimal quantization target; and 2) all input channels are processed uniformly across modalities, overlooking the distinct information densities. In this paper, we introduce VLM-PTQ, an asymmetric post-training quantization framework for VLMs. First, we derive a closed-form correction term that shifts the quantization target, which explicitly accounts for the output residual and the corresponding inverse Hessian column, yielding a better local optimum than RTN. Second, we propose a modality-aware quantization that differentiates channel importance between vision and language tokens, allowing the quantizer to pre-compute better quantization parameters through a lightweight search. Our method extends weight-compensation methods with minimal overhead while achieving significant performance improvements in low-bit scenarios. Extensive experiments demonstrate that VLM-PTQ achieves competitive results compared to existing methods, effectively compressing models from 1B to 72B parameters on a single GPU.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext b991303a-a236-445a-92d9-194dff516df2

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖