Vector Quantization with Self-Attention for Quality-Independent Representation Learning
Zhou Yang, Weisheng Dong, Xin Li, Mengluan Huang, Yulin Sun, Guangming Shi
摘要
Recently, the robustness of deep neural networks has drawn extensive attention due to the potential distribution shift between training and testing data (e.g., deep models trained on high-quality images are sensitive to corruption during testing). Many researchers attempt to make the model learn invariant representations from multiple corrupted data through data augmentation or image-pairbased feature distillation to improve the robustness. Inspired by sparse representation in image restoration, we opt to address this issue by learning image-quality-independent feature representation in a simple plug-and-play manner, that is, to introduce discrete vector quantization (VQ) to remove redundancy in recognition models. Specifically, we first add a codebook module to the network to quantize deep features. Then we concatenate them and design a self-attention module to enhance the representation. During training, we enforce the quantization of features from clean and corrupted images in the same discrete embedding space so that an invariant quality-independent feature representation can be learned to improve the recognition robustness of low-quality images. Qualitative and quantitative experimental results show that our method achieved this goal effectively, leading to a new state-of-the-art result of 43.1 % mCE on ImageNet-C with ResNet50 as the backbone. On other robustness benchmark datasets, such as ImageNet-R, our method also has an accuracy improvement of almost 2%. The source code is available at https://see . xidian.edu.cn/faculty/wsdong/Projects/ VQSA.htm
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Moment Quantization for Video Temporal GroundingXiaolong Sun, Le Wang, Sanping Zhou, Liushuai Shi 等ICCV 2025 · 被引用 2 次
- UniDxMD: Towards Unified Representation for Cross-Modal Unsupervised Domain Adaptation in 3D Semantic SegmentationZhengyin Liang, Hui Yin, Min Liang, Qianqian Du 等ICCV 2025 · 被引用 2 次
- Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring ClassesZhou Yang, Mingtao Feng, Tao Huang, Fangfang Wu 等CVPR 2025
- Controllable Blur Data Augmentation Using 3D-Aware Motion EstimationInsoo Kim, Hana Lee, Hyong-Euk Lee, Jinwoo ShinICLR 2025
- You Always Recognize Me (YARM): Robust Texture Synthesis Against Multi-View CorruptionWeihang Ran, Wei Yuan, Yinqiang ZhengICML 2025
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph 等ICLR 2020 · 被引用 1,572 次
相关 Paper
- Visual Recognition-Driven Image Restoration for Multiple Degradation with Intrinsic Semantics RecoveryZizheng Yang, Jie Huang, Jiahao Chang, Man Zhou 等CVPR 2023
- Deep Degradation Prior for Low-Quality Image ClassificationYang Wang, Yang Cao, Zheng-Jun Zha, Jing Zhang 等CVPR 2020
- Quality-Agnostic Image Recognition via Invertible DecoderInsoo Kim, Seungju Han, Ji-Won Baek, Seong-Jin Park 等CVPR 2021
- Ada-DQA: Adaptive Diverse Quality-aware Feature Acquisition for Video Quality AssessmentHongbo Liu, Mingda Wu, Kun Yuan, Ming Sun 等ACM MM 2023 · 被引用 18 次
- Discrete Representations Strengthen Vision Transformer RobustnessChengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl Vondrick 等ICLR 2022 · 被引用 47 次
