An Information-theoretic Approach to Distribution Shifts
Marco Federici, Ryota Tomioka, Patrick Forré
Abstract
Safely deploying machine learning models to the real world is often a challenging process. Models trained with data obtained from a specific geographic location tend to fail when queried with data obtained elsewhere, agents trained in a simulation can struggle to adapt when deployed in the real world or novel environments, and neural networks that are fit to a subset of the population might carry some selection bias into their decision process. In this work, we describe the problem of data shift from a novel information-theoretic perspective by (i) identifying and describing the different sources of error, (ii) comparing some of the most promising objectives explored in the recent domain generalization, and fair classification literature. From our theoretical analysis and empirical evaluation, we conclude that the model selection procedure needs to be guided by careful considerations regarding the observed data, the factors used for correction, and the structure of the data-generating process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7d15bb2-f036-448e-8d8d-afa718d444f2Cited by top-tier papers9
- Handling Distribution Shifts on Graphs: An Invariance PerspectiveQitian Wu, Hengrui Zhang, Junchi Yan, David WipfICLR 2022 · 261 citations
- Environment-Aware Dynamic Graph Learning for Out-of-Distribution GeneralizationHaonan Yuan, Qingyun Sun, Xingcheng Fu, Ziwei Zhang et al.NeurIPS 2023 · 54 citations
- Generalizing to Evolving Domains with Latent Structure-Aware Sequential AutoencoderTiexin Qin, Shiqi Wang, Haoliang LiICML 2022 · 34 citations
- When Shift Happens - Confounding Is to BlameAbbavaram Gowtham Reddy, Celia Rubio-Madrigal, Rebekka Burkholz, Krikamol MuandetICLR 2026 · 5 citations
- Identifiability Results for Multimodal Contrastive LearningImant Daunhawer, Alice Bizeul, Emanuele Palumbo, Alexander Marx et al.ICLR 2023 · 4 citations
Builds on9
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic RetinopathyEmma Beede, Elizabeth Elliott Baylor, Fred Hersch, Anna Iurchenko et al.CHI 2020 · 589 citations
- Domain Generalization using Causal MatchingDivyat Mahajan, Shruti Tople, Amit SharmaICML 2021 · 399 citations
Related papers
- A Practical Upper Bound on Selection Bias Effects in Medical Prediction ModelsKara Liu, Maggie Wang, Russ B. AltmanKDD 2026
- Model Transferability with Responsive Decision SubjectsYatong Chen, Zeyu Tang, Kun Zhang, Yang LiuICML 2023 · 11 citations
- Change is Hard: A Closer Look at Subpopulation ShiftYuzhe Yang, Haoran Zhang, Dina Katabi, Marzyeh GhassemiICML 2023 · 149 citations
- Fairness and Accuracy under Domain GeneralizationThai-Hoang Pham, Xueru Zhang, Ping ZhangICLR 2023 · 2 citations
- Fairness Guarantees under Demographic ShiftStephen Giguere, Blossom Metevier, Bruno Castro da Silva, Yuriy Brun et al.ICLR 2022 · 57 citations
