ALOFT: A Lightweight MLP-Like Architecture with Dynamic Low-Frequency Transform for Domain Generalization
Jintao Guo, Na Wang, Lei Qi, Yinghuan Shi
Abstract
Domain generalization (DG) aims to learn a model that generalizes well to unseen target domains utilizing multiple source domains without re-training. Most existing DG works are based on convolutional neural networks (CNNs). However, the local operation of the convolution kernel makes the model focus too much on local representations (e.g., texture), which inherently causes the model more prone to overfit to the source domains and hampers its generalization ability. Recently, several MLP-based methods have achieved promising results in supervised learning tasks by learning global interactions among different patches of the image. Inspired by this, in this paper, we first analyze the difference between CNN and MLP methods in DG and find that MLP methods exhibit a better generalization ability because they can better capture the global representations (e.g., structure) than CNN methods. Then, based on a recent lightweight MLP method, we obtain a strong baseline that outperforms most state-of-theart CNN-based methods. The baseline can learn global structure representations with a filter to suppress structureirrelevant information in the frequency space. Moreover, we propose a dynAmic LOw-Frequency spectrum Transform (ALOFT) that can perturb local texture features while preserving global structure features, thus enabling the filter to remove structure-irrelevant information sufficiently. Extensive experiments on four benchmarks have demonstrated that our method can achieve great performance improvement with a small number of parameters compared to SOTA CNN-based DG methods. Our code is available at https://github.com/lingeringlight/ALOFT/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9dd56d18-ba19-4a80-ab72-3b4e05c3cbc8Cited by top-tier papers15
- DomainDrop: Suppressing Domain-Sensitive Channels for Domain GeneralizationJintao Guo, Lei Qi, Yinghuan ShiICCV 2023 · 47 citations
- Generalizable Decision Boundaries: Dualistic Meta-Learning for Open Set Domain GeneralizationXiran Wang, Jian Zhang, Lei Qi, Yinghuan ShiICCV 2023 · 39 citations
- DomainAdaptor: A Novel Approach to Test-time AdaptationJian Zhang, Lei Qi, Yinghuan Shi, Yang GaoICCV 2023 · 31 citations
- Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain SchedulerKunyu Peng, Di Wen, Kailun Yang, Ao Luo et al.NeurIPS 2024 · 20 citations
- Reasoning-Driven Multimodal LLM for Domain GeneralizationZhipeng Xu, Zilong Wang, Xinyang Jiang, Dongsheng Li et al.ICLR 2026 · 11 citations
Builds on42
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
Related papers
- Deep Frequency Filtering for Domain GeneralizationShiqi Lin, Zhizheng Zhang, Zhipeng Huang, Yan Lu et al.CVPR 2023
- Adaptive Texture Filtering for Single-Domain Generalized SegmentationXinhui Li, Mingjia Li, Yaxing Wang, Chuan-Xian Ren et al.AAAI 2023 · 9 citations
- START: A Generalized State Space Model with Saliency-Driven Token-Aware TransformationJintao Guo, Lei Qi, Yinghuan Shi, Yang GaoNeurIPS 2024 · 6 citations
- Domain Generalization by Learning and Removing Domain-specific FeaturesYu Ding, Lei Wang, Bin Liang, Shuming Liang et al.NeurIPS 2022 · 75 citations
- Learning Transferrable and Interpretable Representations for Domain GeneralizationZhekai Du, Jingjing Li, Ke Lu, Lei Zhu et al.ACM MM 2021 · 11 citations
