Vector Quantization with Self-Attention for Quality-Independent Representation Learning
Zhou Yang, Weisheng Dong, Xin Li, Mengluan Huang, Yulin Sun, Guangming Shi
Abstract
Recently, the robustness of deep neural networks has drawn extensive attention due to the potential distribution shift between training and testing data (e.g., deep models trained on high-quality images are sensitive to corruption during testing). Many researchers attempt to make the model learn invariant representations from multiple corrupted data through data augmentation or image-pairbased feature distillation to improve the robustness. Inspired by sparse representation in image restoration, we opt to address this issue by learning image-quality-independent feature representation in a simple plug-and-play manner, that is, to introduce discrete vector quantization (VQ) to remove redundancy in recognition models. Specifically, we first add a codebook module to the network to quantize deep features. Then we concatenate them and design a self-attention module to enhance the representation. During training, we enforce the quantization of features from clean and corrupted images in the same discrete embedding space so that an invariant quality-independent feature representation can be learned to improve the recognition robustness of low-quality images. Qualitative and quantitative experimental results show that our method achieved this goal effectively, leading to a new state-of-the-art result of 43.1 % mCE on ImageNet-C with ResNet50 as the backbone. On other robustness benchmark datasets, such as ImageNet-R, our method also has an accuracy improvement of almost 2%. The source code is available at https://see . xidian.edu.cn/faculty/wsdong/Projects/ VQSA.htm
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30300d93-bc2a-438e-b51a-38430aef8248Cited by top-tier papers6
- Moment Quantization for Video Temporal GroundingXiaolong Sun, Le Wang, Sanping Zhou, Liushuai Shi et al.ICCV 2025 · 2 citations
- UniDxMD: Towards Unified Representation for Cross-Modal Unsupervised Domain Adaptation in 3D Semantic SegmentationZhengyin Liang, Hui Yin, Min Liang, Qianqian Du et al.ICCV 2025 · 2 citations
- Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring ClassesZhou Yang, Mingtao Feng, Tao Huang, Fangfang Wu et al.CVPR 2025
- Controllable Blur Data Augmentation Using 3D-Aware Motion EstimationInsoo Kim, Hana Lee, Hyong-Euk Lee, Jinwoo ShinICLR 2025
- You Always Recognize Me (YARM): Robust Texture Synthesis Against Multi-View CorruptionWeihang Ran, Wei Yuan, Yinqiang ZhengICML 2025
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
Related papers
- Visual Recognition-Driven Image Restoration for Multiple Degradation with Intrinsic Semantics RecoveryZizheng Yang, Jie Huang, Jiahao Chang, Man Zhou et al.CVPR 2023
- Deep Degradation Prior for Low-Quality Image ClassificationYang Wang, Yang Cao, Zheng-Jun Zha, Jing Zhang et al.CVPR 2020
- Quality-Agnostic Image Recognition via Invertible DecoderInsoo Kim, Seungju Han, Ji-Won Baek, Seong-Jin Park et al.CVPR 2021
- Ada-DQA: Adaptive Diverse Quality-aware Feature Acquisition for Video Quality AssessmentHongbo Liu, Mingda Wu, Kun Yuan, Ming Sun et al.ACM MM 2023 · 18 citations
- Discrete Representations Strengthen Vision Transformer RobustnessChengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl Vondrick et al.ICLR 2022 · 47 citations
