Patch-Mix Transformer for Unsupervised Domain Adaptation: A Game Perspective
Jinjing Zhu, Haotian Bai, Lin Wang
摘要
Endeavors have been recently made to leverage the vision transformer (ViT) for the challenging unsupervised domain adaptation (UDA) task. They typically adopt the cross-attention in ViT for direct domain alignment. However, as the performance of cross-attention highly relies on the quality of pseudo labels for targeted samples, it becomes less effective when the domain gap becomes large. We solve this problem from a game theory's perspective with the proposed model dubbed as PMTrans, which bridges source and target domains with an intermediate domain. Specifically, we propose a novel ViT-based module called PatchMix that effectively builds up the intermediate domain, i.e., probability distribution, by learning to sample patches from both domains based on the game-theoretical models. This way, it learns to mix the patches from the source and target domains to maximize the cross entropy (CE), while exploiting two semi-supervised mixup losses in the feature and label spaces to minimize it. As such, we interpret the process of UDA as a min-max CE game with three players, including the feature extractor, classifier, and PatchMix, to find the Nash Equilibria. Moreover, we leverage attention maps from ViT to re-weight the label of each patch by its importance, making it possible to obtain more domain-discriminative feature representations. We conduct extensive experiments on four benchmark datasets, and the results show that PMTrans significantly surpasses the ViT-based and CNNbased SoTA methods by +3.6% on Office-Home, +1.4% on Office-31, and +17.7% on DomainNet, respectively. https: //vlis2022.github.io/cvpr23/PMTrans This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Agile Multi-Source-Free Domain AdaptationXinyao Li, Jingjing Li, Fengling Li, Lei Zhu 等AAAI 2024 · 被引用 23 次
- HiGDA: Hierarchical Graph of Nodes to Learn Local-to-Global Topology for Semi-Supervised Domain AdaptationBa Hung Ngo, Doanh C. Bui, Nhat-Tuong Do-Tran, Tae Jong ChoiAAAI 2025 · 被引用 20 次
- GoodSAM: Bridging Domain and Capacity Gaps via Segment Anything Model for Distortion-Aware Panoramic Semantic SegmentationWeiming Zhang, Yexin Liu, Xu Zheng, Lin WangCVPR 2024 · 被引用 14 次
- On -Divergence Principled Domain Adaptation: An Improved FrameworkZiqiao Wang, Yongyi MaoNeurIPS 2024 · 被引用 13 次
- Style Adaptation and Uncertainty Estimation for Multi-Source Blended-Target Domain AdaptationYuwu Lu, Haoyu Huang, Xue HuNeurIPS 2024 · 被引用 12 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
相关 Paper
- CDTrans: Cross-domain Transformer for Unsupervised Domain AdaptationTongkun Xu, Weihua Chen, Pichao Wang, Fan Wang 等ICLR 2022 · 被引用 293 次
- Adapting Self-Supervised Vision Transformers by Probing Attention-Conditioned Masking ConsistencyViraj Prabhu, Sriram Yenamandra, Aaditya Singh, Judy HoffmanNeurIPS 2022 · 被引用 17 次
- Energy-based Self-Training and Normalization for Unsupervised Domain AdaptationSamitha Herath, Basura Fernando, Ehsan Abbasnejad, Munawar Hayat 等ICCV 2023 · 被引用 10 次
- TransMix: Attend to Mix for Vision TransformersJieneng Chen, Shuyang Sun, Ju He, Philip H. S. Torr 等CVPR 2022
- PiPa: Pixel- and Patch-wise Self-supervised Learning for Domain Adaptative Semantic SegmentationMu Chen, Zhedong Zheng, Yi Yang, Tat-Seng ChuaACM MM 2023 · 被引用 65 次
