Achievable distributional robustness when the robust risk is only partially identified
Julia Kostin, Nicola Gnecco, Fanny Yang
Abstract
In safety-critical applications, machine learning models should generalize well under worst-case distribution shifts, that is, have a small robust risk. Invariance-based algorithms can provably take advantage of structural assumptions on the shifts when the training distributions are heterogeneous enough to identify the robust risk. However, in practice, such identifiability conditions are rarely satisfied -- a scenario so far underexplored in the theoretical literature. In this paper, we aim to fill the gap and propose to study the more general setting when the robust risk is only partially identifiable. In particular, we introduce the worst-case robust risk as a new measure of robustness that is always well-defined regardless of identifiability. Its minimum corresponds to an algorithm-independent (population) minimax quantity that measures the best achievable robustness under partial identifiability. While these concepts can be defined more broadly, in this paper we introduce and derive them explicitly for a linear model for concreteness of the presentation. First, we show that existing robustness methods are provably suboptimal in the partially identifiable case. We then evaluate these methods and the minimizer of the (empirical) worst-case robust risk on real-world gene expression data and find a similar trend: the test error of existing robustness methods grows increasingly suboptimal as the fraction of data from unseen environments increases, whereas accounting for partial identifiability allows for better generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5f3ea64-72b9-4734-9c69-a517bded6c87Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution GeneralizationKartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet et al.NeurIPS 2021 · 372 citations
- Gradient Matching for Domain GeneralizationYuge Shi, Jeffrey Seely, Philip H. S. Torr, Siddharth Narayanaswamy et al.ICLR 2022 · 358 citations
Related papers
- Learning Optimal Features via Partial InvarianceMoulik Choraria, Ibtihal Ferwana, Ankur Mani, Lav R. VarshneyAAAI 2023 · 3 citations
- Provably Invariant Learning without Domain InformationXiaoyu Tan, Lin Yong, Shengyu Zhu, Chao Qu et al.ICML 2023 · 24 citations
- Learning Adversarially Robust Representations via Worst-Case Mutual Information MaximizationSicheng Zhu, Xiao Zhang, David EvansICML 2020 · 30 citations
- Robust Generalization despite Distribution Shift via Minimum Discriminating InformationTobias Sutter, Andreas Krause, Daniel KuhnNeurIPS 2021 · 13 citations
- Adaptive Risk Minimization: Learning to Adapt to Domain ShiftMarvin Zhang, Henrik Marklund, Nikita Dhawan, Abhishek Gupta et al.NeurIPS 2021 · 284 citations
