Efficient Low-Bit Quantization with Adaptive Scales for Multi-Task Co-Training
Boyu Liu, Haoyu Huang, Linlin Yang, Yanjing Li, Guodong Guo, Xianbin Cao, Baochang Zhang
Abstract
Co-training can achieve parameter-efficient multi-task models but remains unexplored for quantization-aware training. Our investigation shows that directly introducing co-training into existing quantization-aware training (QAT) methods results in significant performance degradation. Our experimental study identifies that the primary issue with existing QAT methods stems from the inadequate activation quantization scales for the co-training framework. To address this issue, we propose Task-Specific Scales Quantization for Multi-Task Co-Training (TSQ-MTC) to tackle mismatched quantization scales. Specifically, a task-specific learnable multi-scale activation quantizer (TLMAQ) is incorporated to enrich the representational ability of shared features for different tasks. Additionally, we find that in the deeper layers of the Transformer model, the quantized network suffers from information distortion within the attention quantizer. A structurebased layer-by-layer distillation (SLLD) is then introduced to ensure that the quantized features effectively preserve the information from their full-precision counterparts. Our extensive experiments in two co-training scenarios demonstrate the effectiveness and versatility of TSQ-MTC. In particular, we successfully achieve a 4-bit quantized low-level visual foundation model based on IPT, which attains a PSNR comparable to the full-precision model while offering a 7.99× compression ratio in the ×4 super-resolution task on the Set5 benchmark. † Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PMQ-VE: Progressive Multi-Frame Quantization for Video EnhancementZhanfeng Feng, Long Peng, Xin Di, Yong Guo et al.NeurIPS 2025 · 17 citations
- SURGE: Surrogate Gradient Adaptation in Binary Neural NetworksHaoyu Huang, Boyu Liu, Linlin Yang, Yanjing Li et al.ICML 2026
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 354 citations
- Accurate Post Training Quantization With Small Calibration SetsItay Hubara, Yury Nahshan, Yair Hanani, Ron Banner et al.ICML 2021 · 238 citations
- Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerYanjing Li, Sheng Xu, Baochang Zhang, Xianbin Cao et al.NeurIPS 2022 · 185 citations
Related papers
- Mixa-Q: Revisiting Activation Sparsity for Vision Transformers From a Mixed-Precision Quantization PerspectiveWeitian Wang, Shubham Rai, Cecilia De la Parra, Akash KumarICCV 2025 · 3 citations
- 2DQuant: Low-bit Post-Training Quantization for Image Super-ResolutionKai Liu, Haotong Qin, Yong Guo, Xin Yuan et al.NeurIPS 2024 · 24 citations
- Quantized Feature Distillation for Network QuantizationKe Zhu, Yin-Yin He, Jianxin WuAAAI 2023 · 21 citations
- Task Vector Quantization for Memory-Efficient Model MergingYoungeun Kim, Seunghwan Lee, Aecheon Jung, Bogon Ryu et al.ICCV 2025 · 8 citations
- QSCA: Quantization with Self-Compensating Auxiliary for Monocular Depth EstimationJincheol Yang, Jaemin Choi, Matti Zinke, Suk-Ju KangNeurIPS 2025 · 1 citation
