The Effect of Natural Distribution Shift on Question Answering Models
John Miller, Karl Krauth, Benjamin Recht, Ludwig Schmidt
Abstract
We build four new test sets for the Stanford Question Answering Dataset (SQuAD) and evaluate the ability of question-answering systems to generalize to new data. Our first test set is from the original Wikipedia domain and measures the extent to which existing systems overfit the original test set. Despite several years of heavy test set re-use, we find no evidence of adaptive overfitting. The remaining three test sets are constructed from New York Times articles, Reddit posts, and Amazon product reviews and measure robustness to natural distribution shifts. Across a broad range of models, we observe average performance drops of 3.8, 14.0, and 17.4 F1 points, respectively. In contrast, a strong human baseline matches or exceeds the performance of SQuAD models on the original domain and exhibits little to no drop in new domains. Taken together, our results confirm the surprising resilience of the holdout method and emphasize the need to move towards evaluation metrics that incorporate robustness to natural distribution shifts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e65fa67-8120-41e9-940e-10955023577fCited by top-tier papers43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
Builds on3
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling et al.ICLR 2020 · 179 citations
- What do Models Learn from Question Answering Datasets?Priyanka Sen, Amir SaffariEMNLP 2020 · 40 citations
Related papers
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information ExtractionJi Qi, Chuchun Zhang, Xiaozhi Wang, Kaisheng Zeng et al.EMNLP 2023 · 2 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
- Harvesting and Refining Question-Answer Pairs for Unsupervised QAZhongli Li, Wenhui Wang, Li Dong, Furu Wei et al.ACL 2020 · 29 citations
- To Adapt or to Annotate: Challenges and Interventions for Domain Adaptation in Open-Domain Question AnsweringDheeru Dua, Emma Strubell, Sameer Singh, Pat VergaACL 2023 · 3 citations
