TransMix: Attend to Mix for Vision Transformers
Jieneng Chen, Shuyang Sun, Ju He, Philip H. S. Torr, Alan L. Yuille, Song Bai
摘要
Mixup-based augmentation has been found to be effective for generalizing models during training, especially for Vision Transformers (ViTs) since they can easily overfit. However, previous mixup-based methods have an underlying prior knowledge that the linearly interpolated ratio of targets should be kept the same as the ratio proposed in input interpolation. This may lead to a strange phenomenon that sometimes there is no valid object in the mixed image due to the random process in augmentation but there is still response in the label space. To bridge such gap between the input and label spaces, we propose TransMix, which mixes labels based on the attention maps of Vision Transformers. The confidence of the label will be larger if the corresponding input image is weighted higher by the attention map. TransMix is embarrassingly simple and can be implemented in just a few lines of code without introducing any extra parameters and FLOPs to ViT-based models. Experimental results show that our method can consistently improve various ViT-based models at scales on ImageNet classification. After pre-trained with TransMix on ImageNet, the ViT-based models also demonstrate better transferability to semantic segmentation, object detection and instance segmentation. TransMix also exhibits to be more robust when evaluating on 4 different benchmarks. Code is publicly available at https://github.com/Beckschen/TransMix .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- C-Mixup: Improving Generalization in RegressionHuaxiu Yao, Yiping Wang, Linjun Zhang, James Y. Zou 等NeurIPS 2022 · 被引用 106 次
- Enhancing Minority Classes by Mixing: An Adaptative Optimal Transport Approach for Long-tailed ClassificationJintong Gao, He Zhao, Zhuo Li, Dandan GuoNeurIPS 2023 · 被引用 64 次
- Efficient Large Multi-modal Models via Visual Context CompressionJieneng Chen, Luoxin Ye, Ju He, Zhaoyang Wang 等NeurIPS 2024 · 被引用 49 次
- TransFace: Calibrating Transformer Training for Face Recognition from a Data-Centric PerspectiveJun Dan, Yang Liu, Haoyu Xie, Jiankang Deng 等ICCV 2023 · 被引用 36 次
- Harnessing Hard Mixed Samples with Decoupled RegularizerZicheng Liu, Siyuan Li, Ge Wang, Lirong Wu 等NeurIPS 2023 · 被引用 28 次
它引用的顶会 Paper33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- MixPro: Data Augmentation with MaskMix and Progressive Attention Labeling for Vision TransformerQihao Zhao, Yangyu Huang, Wei Hu, Fan Zhang 等ICLR 2023 · 被引用 3 次
- SMMix: Self-Motivated Image Mixing for Vision TransformersMengzhao Chen, Mingbao Lin, Zhihang Lin, Yuxin Zhang 等ICCV 2023 · 被引用 15 次
- Token-Label Alignment for Vision TransformersHan Xiao, Wenzhao Zheng, Zheng Zhu, Jie Zhou 等ICCV 2023 · 被引用 5 次
- ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision TransformersHanwen Cao, Haobo Lu, Xiaosen Wang, Kun HeICCV 2025 · 被引用 5 次
- All Tokens Matter: Token Labeling for Training Better Vision TransformersZihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou 等NeurIPS 2021 · 被引用 252 次
