Flatness-Aware Minimization for Domain Generalization
Xingxuan Zhang, Renzhe Xu, Han Yu, Yancheng Dong, Pengfei Tian, Peng Cui
Abstract
Domain generalization (DG) seeks to learn robust models that generalize well under unknown distribution shifts. As a critical aspect of DG, optimizer selection has not been explored in depth. Currently, most DG methods follow the widely used benchmark, DomainBed, and utilize Adam as the default optimizer for all datasets. However, we reveal that Adam is not necessarily the optimal choice for the majority of current DG methods and datasets. Based on the perspective of loss landscape flatness, we propose a novel approach, Flatness-Aware Minimization for Domain Generalization (FAD), which can efficiently optimize both zeroth-order and first-order flatness simultaneously for DG. We provide theoretical analyses of the FAD's out-ofdistribution (OOD) generalization error and convergence. Our experimental results demonstrate the superiority of FAD on various DG datasets. Additionally, we confirm that FAD is capable of discovering flatter optima in comparison to other zeroth-order and first-order flatness-aware optimization methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0806a25-7ee3-49d5-b207-bcd5fb8979adCited by top-tier papers22
- Cross-modal Representation Flattening for Multi-modal Domain GeneralizationYunfeng Fan, Wenchao Xu, Haozhao Wang, Song GuoNeurIPS 2024 · 21 citations
- Unknown Domain Inconsistency Minimization for Domain GeneralizationSeungjae Shin, HeeSun Bae, Byeonghu Na, Yoon-Yeong Kim et al.ICLR 2024 · 10 citations
- Asymptotic Unbiased Sample Sampling to Speed Up Sharpness-Aware MinimizationJiaxin Deng, Junbiao Pang, Baochang Zhang, Guodong GuoAAAI 2025 · 5 citations
- Adversarial Data Augmentation for Single Domain Generalization via Lyapunov Exponent-Guided OptimizationZuyu Zhang, Ning Chen, Yongshan Liu, Qinghua Zhang et al.ICCV 2025 · 2 citations
- Improving Sharpness-Aware Minimization by LookaheadRunsheng Yu, Youzhi Zhang, James T. KwokICML 2024 · 1 citation
Builds on49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
Related papers
- Navigating the Flatlands: Dual Adaptive Sharpness-Aware Minimization for Domain GeneralizationJunwen He, Yang He, Lebing Zheng, Zirui Yin et al.ICML 2026
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho et al.NeurIPS 2021 · 630 citations
- Seeking Consistent Flat Minima for Better Domain Generalization via Refining Loss LandscapesAodi Li, Liansheng Zhuang, Xiao Long, Minghong Yao et al.CVPR 2025
- Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Hao Zou et al.CVPR 2023
- Sharpness-Aware Minimization Enhances Feature Quality via Balanced LearningJacob Mitchell Springer, Vaishnavh Nagarajan, Aditi RaghunathanICLR 2024 · 13 citations
