In-N-Out: Pre-Training and Self-Training using Auxiliary Information for Out-of-Distribution Robustness
Sang Michael Xie, Ananya Kumar, Robbie Jones, Fereshte Khani, Tengyu Ma, Percy Liang
摘要
Consider a prediction setting with few in-distribution labeled examples and many unlabeled examples both in-and out-of-distribution (OOD). The goal is to learn a model which performs well both in-distribution and OOD. In these settings, auxiliary information is often cheaply available for every input. How should we best leverage this auxiliary information for the prediction task? Empirically across three image and time-series datasets, and theoretically in a multi-task linear regression setting, we show that (i) using auxiliary information as input features improves in-distribution error but can hurt OOD error; but (ii) using auxiliary information as outputs of auxiliary pre-training tasks improves OOD error. To get the best of both worlds, we introduce In-N-Out, which first trains a model with auxiliary inputs and uses it to pseudolabel all the in-distribution inputs, then pre-trains a model on OOD auxiliary outputs and fine-tunes this model with the pseudolabels (self-training). We show both theoretically and empirically that In-N-Out outperforms auxiliary inputs or outputs alone on both in-distribution and OOD error.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution GeneralizationKartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet 等NeurIPS 2021 · 被引用 372 次
- Fishr: Invariant Gradient Variances for Out-of-Distribution GeneralizationAlexandre Ramé, Corentin Dancette, Matthieu CordICML 2022 · 被引用 262 次
- Cycle Self-Training for Domain AdaptationHong Liu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 被引用 236 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Rethinking Pre-training and Self-trainingBarret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui 等NeurIPS 2020 · 被引用 755 次
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini 等NeurIPS 2020 · 被引用 731 次
- Understanding Self-Training for Gradual Domain AdaptationAnanya Kumar, Tengyu Ma, Percy LiangICML 2020 · 被引用 266 次
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 被引用 263 次
相关 Paper
- Improved OOD Generalization via Adversarial Training and PretraingMingyang Yi, Lu Hou, Jiacheng Sun, Lifeng Shang 等ICML 2021 · 被引用 99 次
- Strengthen Out-of-Distribution Detection Capability with Progressive Self-Knowledge DistillationYang Yang, Haonan XuICML 2025
- Diversified Outlier Exposure for Out-of-Distribution Detection via Informative ExtrapolationJianing Zhu, Yu Geng, Jiangchao Yao, Tongliang Liu 等NeurIPS 2023 · 被引用 54 次
- Out-of-distribution Detection Learning with Unreliable Out-of-distribution SourcesHaotian Zheng, Qizhou Wang, Zhen Fang, Xiaobo Xia 等NeurIPS 2023 · 被引用 53 次
- The Value of Out-of-Distribution DataAshwin De Silva, Rahul Ramesh, Carey E. Priebe, Pratik Chaudhari 等ICML 2023 · 被引用 16 次
