AQuant: Repurposing CODEC for VLM Acceleration via Adaptive Quantization
Zhuoran Song, Chunyu Qi, Jian Weng, Xiaoyao Liang, Haibing Guan
Abstract
Vision-Language Models (VLMs) have reached the forefront of accuracy in various vision understanding tasks. Despite their remarkable success, the computing costs of VLMs scale significantly with the high image resolutions or the increasing number of video frames that need to be processed, posing substantial challenges for deployment to real-time applications. Although specialized quantization accelerators have been developed, they may not be the optimal solutions due to their neglect of the inherent data similarity within VLMs. Additionally, the use of floating-point units for floating-point to integer conversion introduces non-negligible hardware overhead. This paper introduces Adaptive Quantization (AQuant), an algorithm-hardware co-design framework that repurposes the CODEC to accelerate VLM inference in an end-to-end and unified manner. AQuant leverages the inherent similarities in visual tokens, exploiting them for differential value (delta) generation, which is well-suited for dynamic quantization due to its narrower distribution. To eliminate the expensive floating-point similarity detection, AQuant integrates an exponent-based similarity detection operation. On the hardware side, we enhance the video CODEC's capabilities to efficiently implement exponentsimilarity detection and adaptive quantization. The framework also incorporates a Neural Processing Unit (NPU) with mixedprecision support, which collaborates closely with the CODEC to translate algorithmic savings into real speedup. Experimental results show that AQuant achieves speedups of , and over state-of-the-art accelerators, such as LLM.265, CMC, and Xavier AGX GPU, with negligible accuracy loss.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 301db832-4efb-401b-ab38-9eef4b0e644aRelated papers
- DuoQ: A DSP Utilization-aware and Outlier-free Quantization for FPGA-based LLMs AccelerationZhuoquan Yu, Huidong Ji, Yue Cao, Junfu Wu et al.DAC 2025 · 1 citation
- MBQ: Modality-Balanced Quantization for Large Vision-Language ModelsShiyao Li, Yingchun Hu, Xuefei Ning, Xihui Liu et al.CVPR 2025
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language ModelsYutong Wang, Haiyu Wang, Sai Qian ZhangNeurIPS 2025 · 16 citations
- VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference AccelerationDezhan Tu, Danylo Vashchilenko, Yuzhe Lu, Panpan XuICLR 2025
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language ModelsChiyue Wei, Cong Guo, Junyao Zhang, Haoxuan Shan et al.HPCA 2026 · 2 citations
