Loss Function Learning for Domain Generalization by Implicit Gradient
Boyan Gao, Henry Gouk, Yongxin Yang, Timothy M. Hospedales
Abstract
Generalising robustly to distribution shift is a major challenge that is pervasive across most realworld applications of machine learning. A recent study highlighted that many advanced algorithms proposed to tackle such domain generalisation (DG) fail to outperform a properly tuned empirical risk minimisation (ERM) baseline. We take a different approach, and explore the impact of the ERM loss function on out-of-domain generalisation. In particular, we introduce a novel metalearning approach to loss function search based on implicit gradient. This enables us to discover a general purpose parametric loss function that provides a drop-in replacement for cross-entropy. Our loss can be used in standard training pipelines to efficiently train robust models using any neural architecture on new datasets. The results show that it clearly surpasses cross-entropy, enables simple ERM to outperform some more complicated prior DG methods, and provides excellent performance across a variety of DG benchmarks. Furthermore, unlike most existing DG approaches, our setup applies to the most practical setting of single-source domain generalisation, on which we show significant improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Domain Generalization Guided by Gradient Signal to Noise Ratio of ParametersMateusz Michalkiewicz, Masoud Faraki, Xiang Yu, Manmohan Chandraker et al.ICCV 2023 · 9 citations
- Optimization Inspired Few-Shot Adaptation for Large Language ModelsBoyan Gao, Xin Wang, Yibo Yang, David A. CliftonNeurIPS 2025 · 3 citations
- Automated Loss function Search for Class-imbalanced Node ClassificationXinyu Guo, Kai Wu, Xiaoyu Zhang, Jing LiuICML 2024 · 2 citations
- DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated LearningSikai Bai, Jie Zhang, Song Guo, Shuaicheng Li et al.CVPR 2024
- EnsLoss: Stochastic Calibrated Loss Ensembles for Preventing Overfitting in ClassificationBen DaiICML 2025
Builds on12
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
- Domain Generalization with MixStyleKaiyang Zhou, Yongxin Yang, Yu Qiao, Tao XiangICLR 2021 · 986 citations
Related papers
- An Empirical Investigation of Domain Generalization with Empirical Risk MinimizersRamakrishna Vedantam, David Lopez-Paz, David J. SchwabNeurIPS 2021 · 49 citations
- Distribution Shift Is Key to Learning Invariant PredictionHong Zheng, Fei TengAAAI 2026
- Coping with Label Shift via Distributionally Robust OptimisationJingzhao Zhang, Aditya Krishna Menon, Andreas Veit, Srinadh Bhojanapalli et al.ICLR 2021 · 79 citations
- Task-Robust Model-Agnostic Meta-LearningLiam Collins, Aryan Mokhtari, Sanjay ShakkottaiNeurIPS 2020 · 66 citations
- Probable Domain Generalization via Quantile Risk MinimizationCian Eastwood, Alexander Robey, Shashank Singh, Julius von Kügelgen et al.NeurIPS 2022 · 99 citations
