Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and Future
Linyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu, Yidong Wang, Jingming Zhuo, Lingqiao Liu, Jindong Wang, Jennifer Foster, Yue Zhang
Abstract
Machine learning (ML) systems in natural language processing (NLP) face significant challenges in generalizing to out-of-distribution (OOD) data, where the test distribution differs from the training data distribution. This poses important questions about the robustness of NLP models and their high accuracy, which may be artificially inflated due to their underlying sensitivity to systematic biases. Despite these challenges, there is a lack of comprehensive surveys on the generalization challenge from an OOD perspective in natural language understanding. Therefore, this paper aims to fill this gap by presenting the first comprehensive review of recent progress, methods, and evaluations on this topic. We further discuss the challenges involved and potential future research directions. By providing convenient access to existing work, we hope this survey will encourage future research in this area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng et al.ICLR 2024 · 368 citations
- SLANG: New Concept Comprehension of Large Language ModelsLingrui Mei, Shenghua Liu, Yiwei Wang, Baolong Bi et al.EMNLP 2024 · 7 citations
- Robust Preference Alignment via Directional Neighborhood ConsensusRuochen Mao, Yuling Shi, Xiaodong Gu, Jiaheng WeiICLR 2026 · 2 citations
- The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language ModelsAdithya Bhaskar, Dan Friedman, Danqi ChenACL 2024 · 1 citation
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World ScenariosLisa Alazraki, Lihu Chen, Ana Brassard, Joe Stacey et al.ACL 2026 · 1 citation
Builds on62
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
Related papers
- Types of Out-of-Distribution Texts and How to Detect ThemUdit Arora, William Huang, He HeEMNLP 2021
- OODREB: Benchmarking State-of-the-Art Methods for Out-Of-Distribution Generalization on Relation ExtractionHaotian Chen, Houjing Guo, Bingsheng Chen, Xiangdong ZhouWWW 2024 · 1 citation
- Improved OOD Generalization via Adversarial Training and PretraingMingyang Yi, Lu Hou, Jiacheng Sun, Lifeng Shang et al.ICML 2021 · 99 citations
- IRM - when it works and when it doesn't: A test case of natural language inferenceYana Dranker, He He, Yonatan BelinkovNeurIPS 2021 · 22 citations
- Do-GOOD: Towards Distribution Shift Evaluation for Pre-Trained Visual Document Understanding ModelsJiabang He, Yi Hu, Lei Wang, Xing Xu et al.SIGIR 2023 · 4 citations
