Modeling the Data-Generating Process is Necessary for Out-of-Distribution Generalization
Jivat Neet Kaur, Emre Kiciman, Amit Sharma
摘要
Recent empirical studies on domain generalization (DG) have shown that DG algorithms that perform well on some distribution shifts fail on others, and no state-of-the-art DG algorithm performs consistently well on all shifts. Moreover, real-world data often has multiple distribution shifts over different attributes; hence we introduce multi-attribute distribution shift datasets and find that the accuracy of existing DG algorithms falls even further. To explain these results, we provide a formal characterization of generalization under multi-attribute shifts using a canonical causal graph. Based on the relationship between spurious attributes and the classification label, we obtain realizations of the canonical causal graph that characterize common distribution shifts and show that each shift entails different independence constraints over observed variables. As a result, we prove that any algorithm based on a single, fixed constraint cannot work well across all shifts, providing theoretical evidence for mixed empirical results on DG algorithms. Based on this insight, we develop Causally Adaptive Constraint Minimization (CACM), an algorithm that uses knowledge about the data-generating process to adaptively identify and apply the correct independence constraints for regularization. Results on fully synthetic, MNIST, small NORB, and Waterbirds datasets, covering binary and multi-valued attributes and labels, show that adaptive dataset-dependent constraints lead to the highest accuracy on unseen domains whereas incorrect constraints fail to do so. Our results demonstrate the importance of modeling the causal relationships inherent in the data-generating process. INTRODUCTION To perform reliably in real world settings, machine learning models must be robust to distribution shifts -where the training distribution differs from the test distribution. Given data from multiple domains that share a common optimal predictor, the domain generalization (DG) task (Wang et al., 2021; Zhou et al., 2021) encapsulates this challenge by evaluating accuracy on an unseen domain. Recent empirical studies of DG algorithms (Wiles et al., 2022; Ye et al., 2022) have characterized different kinds of distribution shifts across domains. Using MNIST as an example, a diversity shift is when domains are created either by adding new values of a spurious attribute like rotation (e.g., Rotated-MNIST dataset (Ghifary et al., 2015; Piratla et al., 2020) ) whereas a correlation shift is when domains exhibit different values of correlation between the class label and a spurious attribute like color (e.g., Colored-MNIST (Arjovsky et al., 2019)). Partly because advances in representation learning for DG (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskMaya Okawa, Ekdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2023 · 被引用 113 次
- Mechanistic Mode ConnectivityEkdeep Singh Lubana, Eric J. Bigelow, Robert P. Dick, David Scott Krueger 等ICML 2023 · 被引用 57 次
- Joint Learning of Label and Environment Causal Independence for Graph Out-of-Distribution GeneralizationShurui Gui, Meng Liu, Xiner Li, Youzhi Luo 等NeurIPS 2023 · 被引用 54 次
- Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept SpaceCore Francisco Park, Maya Okawa, Andrew Lee, Ekdeep Singh Lubana 等NeurIPS 2024 · 被引用 39 次
- Do causal predictors generalize better to new domains?Vivian Y. Nastl, Moritz HardtNeurIPS 2024 · 被引用 21 次
它引用的顶会 Paper13
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang 等ICML 2021 · 被引用 1,163 次
相关 Paper
- Causality Inspired Representation Learning for Domain GeneralizationFangrui Lv, Jian Liang, Shuang Li, Bin Zang 等CVPR 2022 · 被引用 190 次
- Causal Structure-guided Distributionally Robust Optimization under Domain ShiftsSeonggyeom Kim, Eunjung Choi, Dong-Kyu ChaeKDD 2026
- Probable Domain Generalization via Quantile Risk MinimizationCian Eastwood, Alexander Robey, Shashank Singh, Julius von Kügelgen 等NeurIPS 2022 · 被引用 99 次
- Algorithmic Fairness Generalization under Covariate and Dependence Shifts SimultaneouslyChen Zhao, Kai Jiang, Xintao Wu, Haoliang Wang 等KDD 2024 · 被引用 6 次
- Seeing Through the Shift: Causality-Inspired Robust Generalized Category DiscoveryWei Feng, Yiwen Jiang, Sijin Zhou, Zhuang Qi 等CVPR 2026 · 被引用 2 次
