TransMix: Attend to Mix for Vision Transformers
Jieneng Chen, Shuyang Sun, Ju He, Philip H. S. Torr, Alan L. Yuille, Song Bai
Abstract
Mixup-based augmentation has been found to be effective for generalizing models during training, especially for Vision Transformers (ViTs) since they can easily overfit. However, previous mixup-based methods have an underlying prior knowledge that the linearly interpolated ratio of targets should be kept the same as the ratio proposed in input interpolation. This may lead to a strange phenomenon that sometimes there is no valid object in the mixed image due to the random process in augmentation but there is still response in the label space. To bridge such gap between the input and label spaces, we propose TransMix, which mixes labels based on the attention maps of Vision Transformers. The confidence of the label will be larger if the corresponding input image is weighted higher by the attention map. TransMix is embarrassingly simple and can be implemented in just a few lines of code without introducing any extra parameters and FLOPs to ViT-based models. Experimental results show that our method can consistently improve various ViT-based models at scales on ImageNet classification. After pre-trained with TransMix on ImageNet, the ViT-based models also demonstrate better transferability to semantic segmentation, object detection and instance segmentation. TransMix also exhibits to be more robust when evaluating on 4 different benchmarks. Code is publicly available at https://github.com/Beckschen/TransMix .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55f27350-2393-4531-99d2-cf341d313900Cited by top-tier papers27
- C-Mixup: Improving Generalization in RegressionHuaxiu Yao, Yiping Wang, Linjun Zhang, James Y. Zou et al.NeurIPS 2022 · 106 citations
- Enhancing Minority Classes by Mixing: An Adaptative Optimal Transport Approach for Long-tailed ClassificationJintong Gao, He Zhao, Zhuo Li, Dandan GuoNeurIPS 2023 · 64 citations
- Efficient Large Multi-modal Models via Visual Context CompressionJieneng Chen, Luoxin Ye, Ju He, Zhaoyang Wang et al.NeurIPS 2024 · 49 citations
- TransFace: Calibrating Transformer Training for Face Recognition from a Data-Centric PerspectiveJun Dan, Yang Liu, Haoyu Xie, Jiankang Deng et al.ICCV 2023 · 36 citations
- Harnessing Hard Mixed Samples with Decoupled RegularizerZicheng Liu, Siyuan Li, Ge Wang, Lirong Wu et al.NeurIPS 2023 · 28 citations
Builds on33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- MixPro: Data Augmentation with MaskMix and Progressive Attention Labeling for Vision TransformerQihao Zhao, Yangyu Huang, Wei Hu, Fan Zhang et al.ICLR 2023 · 3 citations
- SMMix: Self-Motivated Image Mixing for Vision TransformersMengzhao Chen, Mingbao Lin, Zhihang Lin, Yuxin Zhang et al.ICCV 2023 · 15 citations
- Token-Label Alignment for Vision TransformersHan Xiao, Wenzhao Zheng, Zheng Zhu, Jie Zhou et al.ICCV 2023 · 5 citations
- ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision TransformersHanwen Cao, Haobo Lu, Xiaosen Wang, Kun HeICCV 2025 · 5 citations
- All Tokens Matter: Token Labeling for Training Better Vision TransformersZihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou et al.NeurIPS 2021 · 252 citations
