Gradient-Guided Annealing for Domain Generalization
Aristotelis Ballas, Christos Diou
Abstract
Figure 1. (left) Decision boundaries of a 4 th -degree polynomial logistic regression model with 2D input. In this example, feature x1 is class-specific and x2 is domain-specific, while color represents classes and shapes represent domains. The samples with solid red and green colors are included in the training data, whereas the fainted samples are part of the hidden held-out test set. As a result, domain shift is represented by a change in x2. Although the classifier should only infer based on x1, traditional gradient descent leads to overfitting (top-left). The proposed method, GGA (bottom-left), introduces an annealing process that depends on gradient agreement, leading to models that generalize well to new, unobserved target domains. (right) Schematics of the parameter updates of ERM (top-right) and GGA (bottom-right). Parameters updated via ERM are driven by gradient conflict, whereas GGA searches for a point where gradients align before continuing descending towards a minima.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Reasoning-Driven Multimodal LLM for Domain GeneralizationZhipeng Xu, Zilong Wang, Xinyang Jiang, Dongsheng Li et al.ICLR 2026 · 11 citations
- Noise-Aware Generalization: Robustness to In-Domain Noise and Out-of-Domain GeneralizationSiqi Wang, Aoming Liu, Bryan A. PlummerICLR 2026 · 3 citations
- Rethinking Out-of-Distribution Detection and Generalization with Collective Behavior DynamicsZhenbin Wang, Lei Zhang, Wei Huang, Zhao Zhang et al.NeurIPS 2025 · 2 citations
- Anomaly-Related Residual Fields for Cross-domain Anomaly DetectionKewei Gao, Jiayi Xie, Zhengda Shen, Weijun Qin et al.CVPR 2026
- HamiPose: Hamiltonian Optimization for Unsupervised Domain Adaptive Pose EstimationJiawen Li, Fei Jiang, Dandan Zhu, Aimin ZhouCVPR 2026
Builds on21
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
Related papers
- Federated Unsupervised Domain Generalization Using Global and Local Alignment of GradientsFarhad Pourpanah, Mahdiyar Molahasani, Milad Soltany, Michael A. Greenspan et al.AAAI 2025 · 10 citations
- Domain Generalization via Gradient SurgeryLucas Mansilla, Rodrigo Echeveste, Diego H. Milone, Enzo FerranteICCV 2021 · 98 citations
- Gradient Distribution Alignment Certificates Better Adversarial Domain AdaptationZhiqiang Gao, Shufei Zhang, Kaizhu Huang, Qiufeng Wang et al.ICCV 2021 · 56 citations
- One-Step Generalization Ratio Guided Optimization for Domain GeneralizationSumin Cho, Dongwon Kim, Kwangsu KimICML 2025
- Towards Understanding GD with Hard and Conjugate Pseudo-labels for Test-Time AdaptationJun-Kun Wang, Andre WibisonoICLR 2023 · 2 citations
