Distributionally Robust Models with Parametric Likelihood Ratios
Paul Michel, Tatsunori Hashimoto, Graham Neubig
摘要
As machine learning models are deployed ever more broadly, it becomes increasingly important that they are not only able to perform well on their training distribution, but also yield accurate predictions when confronted with distribution shift. The Distributionally Robust Optimization (DRO) framework proposes to address this issue by training models to minimize their expected risk under a collection of distributions, to imitate test-time shifts. This is most commonly achieved by instance-level re-weighting of the training objective to emulate the likelihood ratio with possible test distributions, which allows for estimating their empirical risk via importance sampling (assuming that they are subpopulations of the training distribution). However, re-weighting schemes in the literature are usually limited due to the difficulty of keeping the optimization problem tractable and the complexity of enforcing normalization constraints. In this paper, we show that three simple ideas -- mini-batch level normalization, a KL penalty and simultaneous gradient updates -- allow us to train models with DRO using a broader class of parametric likelihood ratios. In a series of experiments on both image and text classification benchmarks, we find that models trained with the resulting parametric adversaries are consistently more robust to subpopulation shifts when compared to other DRO approaches, and that the method performs reliably well with little hyper-parameter tuning. Code to reproduce our experiments can be found at https://github.com/pmichel31415/P-DRO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Understanding Contrastive Learning via Distributionally Robust OptimizationJunkang Wu, Jiawei Chen, Jiancan Wu, Wentao Shi 等NeurIPS 2023 · 被引用 55 次
- UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware MixupZongbo Han, Zhipeng Liang, Fan Yang, Liu Liu 等NeurIPS 2022 · 被引用 53 次
- Beyond Uniform Sampling: Offline Reinforcement Learning with Imbalanced DatasetsZhang-Wei Hong, Aviral Kumar, Sathwik Karnik, Abhishek Bhandwaldar 等NeurIPS 2023 · 被引用 34 次
- Learning from Heterogeneity: Generalizing Dynamic Facial Expression Recognition via Distributionally Robust OptimizationFeng-Qi Cui, Anyang Tong, Jinyang Huang, Jie Zhang 等ACM MM 2025 · 被引用 10 次
- Factored DRO: Factored Distributionally Robust Policies for Contextual BanditsTong Mu, Yash Chandak, Tatsunori B. Hashimoto, Emma BrunskillNeurIPS 2022 · 被引用 8 次
它引用的顶会 Paper10
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- Large-Scale Methods for Distributionally Robust OptimizationDaniel Levy, Yair Carmon, John C. Duchi, Aaron SidfordNeurIPS 2020 · 被引用 281 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Examining and Combating Spurious Features under Distribution ShiftChunting Zhou, Xuezhe Ma, Paul Michel, Graham NeubigICML 2021 · 被引用 78 次
相关 Paper
- Doubly Robust Instance-Reweighted Adversarial TrainingDaouda Sow, Sen Lin, Zhangyang Wang, Yingbin LiangICLR 2024 · 被引用 2 次
- Distributionally Robust Post-hoc Classifiers under Prior ShiftsJiaheng Wei, Harikrishna Narasimhan, Ehsan Amid, Wen-Sheng Chu 等ICLR 2023
- Understanding Why Generalized Reweighting Does Not Improve Over ERMRuntian Zhai, Chen Dan, J. Zico Kolter, Pradeep Kumar RavikumarICLR 2023 · 被引用 6 次
- Coping with Label Shift via Distributionally Robust OptimisationJingzhao Zhang, Aditya Krishna Menon, Andreas Veit, Srinadh Bhojanapalli 等ICLR 2021 · 被引用 79 次
- Non-convex Distributionally Robust Optimization: Non-asymptotic AnalysisJikai Jin, Bohang Zhang, Haiyang Wang, Liwei WangNeurIPS 2021 · 被引用 65 次
