Maximum Likelihood Estimation is All You Need for Well-Specified Covariate Shift
Jiawei Ge, Shange Tang, Jianqing Fan, Cong Ma, Chi Jin
Abstract
A key challenge of modern machine learning systems is to achieve Out-of-Distribution (OOD) generalization -- generalizing to target data whose distribution differs from that of source data. Despite its significant importance, the fundamental question of ``what are the most effective algorithms for OOD generalization'' remains open even under the standard setting of covariate shift. This paper addresses this fundamental question by proving that, surprisingly, classical Maximum Likelihood Estimation (MLE) purely using source data (without any modification) achieves the minimax optimality for covariate shift under the well-specified setting. That is, no algorithm performs better than MLE in this setting (up to a constant factor), justifying MLE is all you need. Our result holds for a very rich class of parametric models, and does not require any boundedness condition on the density ratio. We illustrate the wide applicability of our framework by instantiating it to three concrete examples -- linear regression, logistic regression, and phase retrieval. This paper further complement the study by proving that, under the misspecified setting, MLE is no longer the optimal choice, whereas Maximum Weighted Likelihood Estimator (MWLE) emerges as minimax optimal in certain scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15ac10f5-582e-4dd3-af88-e272e0e6282cCited by top-tier papers4
- High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit RegularizationYihang Chen, Fanghui Liu, Taiji Suzuki, Volkan CevherICML 2024 · 5 citations
- When Shift Happens - Confounding Is to BlameAbbavaram Gowtham Reddy, Celia Rubio-Madrigal, Rebekka Burkholz, Krikamol MuandetICLR 2026 · 5 citations
- Mixed-Sample SGD: an End-to-end Analysis of Supervised Transfer LearningYuyang Deng, Samory KpotufeNeurIPS 2025 · 1 citation
- Benign Overfitting in Out-of-Distribution Generalization of Linear ModelsShange Tang, Jiayun Wu, Jianqing Fan, Chi JinICLR 2025
Builds on4
- The Risks of Invariant Risk MinimizationElan Rosenfeld, Pradeep Kumar Ravikumar, Andrej RisteskiICLR 2021 · 356 citations
- Near-Optimal Linear Regression under Distribution ShiftQi Lei, Wei Hu, Jason D. LeeICML 2021 · 45 citations
- A new similarity measure for covariate shift with applications to nonparametric regressionReese Pathak, Cong Ma, Martin J. WainwrightICML 2022 · 40 citations
- Minimax Lower Bounds for Transfer Learning with Linear and One-hidden Layer Neural NetworksSeyed Mohammadreza Mousavi Kalan, Zalan Fabian, Salman Avestimehr, Mahdi SoltanolkotabiNeurIPS 2020 · 37 citations
Related papers
- Towards a Unified Analysis of Kernel-based Methods Under Covariate ShiftXingdong Feng, Xin He, Caixing Wang, Chao Wang et al.NeurIPS 2023 · 17 citations
- A Theoretical Analysis on Independence-driven Importance Weighting for Covariate-shift GeneralizationRenzhe Xu, Xingxuan Zhang, Zheyan Shen, Tong Zhang et al.ICML 2022 · 36 citations
- Open Set Label Shift with Test Time Out-of-Distribution ReferenceChangkun Ye, Russell Tsuchida, Lars Petersson, Nick BarnesCVPR 2025
- Feed Two Birds with One Scone: Exploiting Wild Data for Both Out-of-Distribution Generalization and DetectionHaoyue Bai, Gregory Canal, Xuefeng Du, Jeongyeol Kwon et al.ICML 2023 · 67 citations
- Double-Weighting for Covariate Shift AdaptationJosé Ignacio Segovia-Martín, Santiago Mazuelas, Anqi LiuICML 2023 · 9 citations
