SAQ-SAM: Semantically-Aligned Quantization for Segment Anything Model
Jing Zhang, Zhikai Li, Chengzhi Hu, Xuewen Liu, Qingyi Gu
摘要
Segment Anything Model (SAM) exhibits remarkable zeroshot segmentation capability; however, its prohibitive computational costs make edge deployment challenging. Although post-training quantization (PTQ) offers a promising compression solution, existing methods yield unsatisfactory results when applied to SAM, owing to its specialized model components and promptable workflow: (i) The mask decoder's attention exhibits extreme activation outliers, and we find that aggressive clipping (even 100×), without smoothing or isolation, is effective in suppressing outliers while maintaining performance. Unfortunately, traditional distributionbased metrics (e.g., MSE) fail to provide such large-scale clipping. (ii) Existing quantization reconstruction methods neglect semantic interactivity of SAM, leading to misalignment between image feature and prompt intention. To address the above issues, we propose SAQ-SAM in this paper, which boosts PTQ for SAM from the perspective of semantic alignment. Specifically, we propose Perceptual-Consistency Clipping, which exploits attention focus overlap to promote aggressive clipping while preserving semantic capabilities. Furthermore, we propose Prompt-Aware Reconstruction, which incorporates image-prompt interactions by leveraging crossattention in mask decoder, thus facilitating alignment in both distribution and semantic. Moreover, to ensure the interaction efficiency, we design a layer-skipping strategy for image tokens in encoder. Extensive experiments are conducted on various SAM sizes and tasks, including instance segmentation, oriented object detection, and semantic segmentation, and the results show that our method consistently exhibits advantages. For example, when quantizing SAM-B to 4-bit, SAQ-SAM achieves 11.7% higher mAP than the baseline in instance segmentation task. Code is available at https://github.com/jingjing0419/SAQ-SAM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory RetrievalJing Zhang, Zhikai Li, Xuewen Liu, Qingyi GuICLR 2026 · 被引用 5 次
- PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation ModelsXuewen Liu, Zhikai Li, Jing Zhang, Mengjuan Chen 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper16
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang 等ICLR 2023 · 被引用 753 次
相关 Paper
- PTQ4SAM: Post-Training Quantization for Segment AnythingChengtao Lv, Hong Chen, Jinyang Guo, Yifu Ding 等CVPR 2024 · 被引用 22 次
- CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything ModelHouji Wen, Jiangyong Yu, Dawei Yang, Jun LiCVPR 2026 · 被引用 2 次
- TinySAM: Pushing the Envelope for Efficient Segment Anything ModelHan Shu, Wenshuo Li, Yehui Tang, Yiman Zhang 等AAAI 2025 · 被引用 57 次
- EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment AnythingYunyang Xiong, Bala Varadarajan, Lemeng Wu, Xiaoyu Xiang 等CVPR 2024 · 被引用 185 次
- Personalize Segment Anything Model with One ShotRenrui Zhang, Zhengkai Jiang, Ziyu Guo, Shilin Yan 等ICLR 2024 · 被引用 333 次
