MixPro: Data Augmentation with MaskMix and Progressive Attention Labeling for Vision Transformer
Qihao Zhao, Yangyu Huang, Wei Hu, Fan Zhang, Jun Liu
Abstract
The recently proposed data augmentation TransMix employs attention labels to help visual transformers (ViT) achieve better robustness and performance. However, TransMix is deficient in two aspects: 1) The image cropping method of TransMix may not be suitable for ViTs. 2) At the early stage of training, the model produces unreliable attention maps. TransMix uses unreliable attention maps to compute mixed attention labels that can affect the model. To address the aforementioned issues, we propose MaskMix and Progressive Attention Labeling (PAL) in image and label space, respectively. In detail, from the perspective of image space, we design MaskMix, which mixes two images based on a patch-like grid mask. In particular, the size of each mask patch is adjustable and is a multiple of the image patch size, which ensures each image patch comes from only one image and contains more global contents. From the perspective of label space, we design PAL, which utilizes a progressive factor to dynamically re-weight the attention weights of the mixed attention label. Finally, we combine MaskMix and Progressive Attention Labeling as our new data augmentation method, named MixPro. The experimental results show that our method can improve various ViT-based models at scales on ImageNet classification (73.8% top-1 accuracy based on DeiT-T for 300 epochs). After being pre-trained with MixPro on ImageNet, the ViT-based models also demonstrate better transferability to semantic segmentation, object detection, and instance segmentation. Furthermore, compared to TransMix, MixPro also shows stronger robustness on several benchmarks. The code is available at https://github.com/fistyee/MixPro.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1cf40c3-9e04-4c75-be65-eb5b9780ffd5Cited by top-tier papers6
- MDCS: More Diverse Experts with Consistency Self-distillation for Long-tailed RecognitionQihao Zhao, Chen Jiang, Wei Hu, Fan Zhang et al.ICCV 2023 · 35 citations
- LTGC: Long-Tail Recognition via Leveraging LLMs-Driven Generated ContentQihao Zhao, Yalun Dai, Hao Li, Wei Hu et al.CVPR 2024 · 22 citations
- Adversarial AutoMixupHuafeng Qin, Xin Jin, Yun Jiang, Mounîm A. El-Yacoubi et al.ICLR 2024 · 19 citations
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal UnderstandingXin Jin, Siyuan Li, Siyong Jian, Kai Yu et al.ICLR 2026 · 14 citations
- Diffusemix: Label-Preserving Data Augmentation with Diffusion ModelsKhawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood, Karthik NandakumarCVPR 2024
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
Related papers
- TransMix: Attend to Mix for Vision TransformersJieneng Chen, Shuyang Sun, Ju He, Philip H. S. Torr et al.CVPR 2022
- SMMix: Self-Motivated Image Mixing for Vision TransformersMengzhao Chen, Mingbao Lin, Zhihang Lin, Yuxin Zhang et al.ICCV 2023 · 15 citations
- Token-Label Alignment for Vision TransformersHan Xiao, Wenzhao Zheng, Zheng Zhu, Jie Zhou et al.ICCV 2023 · 5 citations
- TokenMixup: Efficient Attention-guided Token-level Data Augmentation for TransformersHyeong Kyu Choi, Joonmyung Choi, Hyunwoo J. KimNeurIPS 2022 · 49 citations
- Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision TransformersBum Jun Kim, Sang Woo KimAAAI 2025 · 2 citations
