Omitted Variable Bias in Language Models Under Distribution Shift
Victoria Lin, Louis-Philippe Morency, Eli Ben-Michael
Abstract
Despite their impressive performance on a wide variety of tasks, modern language models remain susceptible to distribution shifts, exhibiting brittle behavior when evaluated on data that differs in distribution from their training data. In this paper, we describe how distribution shifts in language models can be separated into observable and unobservable components, and we discuss how established approaches for dealing with distribution shift address only the former. Importantly, we identify that the resulting omitted variable bias from unobserved variables can compromise both evaluation and optimization in language models. To address this challenge, we introduce a framework that maps the strength of the omitted variables to bounds on the worst-case generalization performance of language models under distribution shift. In empirical experiments, we show that using these bounds directly in language model evaluation and optimization provides more principled measures of out-of-distribution performance, improves true out-of-distribution performance relative to standard distribution shift adjustment methods, and further enables inference about the strength of the omitted variables when target distribution labels are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ade44839-b341-4ca7-ad6a-3e3e325ad816Builds on9
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Change is Hard: A Closer Look at Subpopulation ShiftYuzhe Yang, Haoran Zhang, Dina Katabi, Marzyeh GhassemiICML 2023 · 149 citations
- Doubly Robust Distributionally Robust Off-Policy Evaluation and LearningNathan Kallus, Xiaojie Mao, Kaiwen Wang, Zhengyuan ZhouICML 2022 · 39 citations
- Does Reasoning Emerge? Examining the Probabilities of Causation in Large Language ModelsJavier González, Aditya V. NoriNeurIPS 2024 · 13 citations
- Doubly Robust Counterfactual ClassificationKwangho Kim, Edward H. Kennedy, José R. ZubizarretaNeurIPS 2022 · 9 citations
Related papers
- Robust Prompt Optimization for Large Language Models Against Distribution ShiftsMoxin Li, Wenjie Wang, Fuli Feng, Yixin Cao et al.EMNLP 2023 · 5 citations
- Is Fine-tuning Needed? Pre-trained Language Models Are Near Perfect for Out-of-Domain DetectionRheeya Uppaal, Junjie Hu, Yixuan LiACL 2023 · 9 citations
- Can Large Language Models Learn Independent Causal Mechanisms?Gaël Gendron, Bao Trung Nguyen, Alex Yuxuan Peng, Michael J. Witbrock et al.EMNLP 2024 · 2 citations
- Types of Out-of-Distribution Texts and How to Detect ThemUdit Arora, William Huang, He HeEMNLP 2021
- Measuring Distribution Shift in User Prompts and Its Effects on LLM PerformanceParker Seegmiller, Sarah Masud PreumACL 2026 · 1 citation
