Omitted Variable Bias in Language Models Under Distribution Shift
Victoria Lin, Louis-Philippe Morency, Eli Ben-Michael
摘要
Despite their impressive performance on a wide variety of tasks, modern language models remain susceptible to distribution shifts, exhibiting brittle behavior when evaluated on data that differs in distribution from their training data. In this paper, we describe how distribution shifts in language models can be separated into observable and unobservable components, and we discuss how established approaches for dealing with distribution shift address only the former. Importantly, we identify that the resulting omitted variable bias from unobserved variables can compromise both evaluation and optimization in language models. To address this challenge, we introduce a framework that maps the strength of the omitted variables to bounds on the worst-case generalization performance of language models under distribution shift. In empirical experiments, we show that using these bounds directly in language model evaluation and optimization provides more principled measures of out-of-distribution performance, improves true out-of-distribution performance relative to standard distribution shift adjustment methods, and further enables inference about the strength of the omitted variables when target distribution labels are available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Change is Hard: A Closer Look at Subpopulation ShiftYuzhe Yang, Haoran Zhang, Dina Katabi, Marzyeh GhassemiICML 2023 · 被引用 149 次
- Doubly Robust Distributionally Robust Off-Policy Evaluation and LearningNathan Kallus, Xiaojie Mao, Kaiwen Wang, Zhengyuan ZhouICML 2022 · 被引用 39 次
- Does Reasoning Emerge? Examining the Probabilities of Causation in Large Language ModelsJavier González, Aditya V. NoriNeurIPS 2024 · 被引用 13 次
- Doubly Robust Counterfactual ClassificationKwangho Kim, Edward H. Kennedy, José R. ZubizarretaNeurIPS 2022 · 被引用 9 次
相关 Paper
- Robust Prompt Optimization for Large Language Models Against Distribution ShiftsMoxin Li, Wenjie Wang, Fuli Feng, Yixin Cao 等EMNLP 2023 · 被引用 5 次
- Is Fine-tuning Needed? Pre-trained Language Models Are Near Perfect for Out-of-Domain DetectionRheeya Uppaal, Junjie Hu, Yixuan LiACL 2023 · 被引用 9 次
- Can Large Language Models Learn Independent Causal Mechanisms?Gaël Gendron, Bao Trung Nguyen, Alex Yuxuan Peng, Michael J. Witbrock 等EMNLP 2024 · 被引用 2 次
- Types of Out-of-Distribution Texts and How to Detect ThemUdit Arora, William Huang, He HeEMNLP 2021
- Measuring Distribution Shift in User Prompts and Its Effects on LLM PerformanceParker Seegmiller, Sarah Masud PreumACL 2026 · 被引用 1 次
