Do Question Answering Modeling Improvements Hold Across Benchmarks?
Nelson F. Liu, Tony Lee, Robin Jia, Percy Liang
Abstract
Do question answering (QA) modeling improvements (e.g., choice of architecture and training procedure) hold consistently across the diverse landscape of QA benchmarks? To study this question, we introduce the notion of concurrence—two benchmarks have high concurrence on a set of modeling approaches if they rank the modeling approaches similarly. We measure the concurrence between 32 QA benchmarks on a set of 20 diverse modeling approaches and find that human-constructed benchmarks have high concurrence amongst themselves, even if their passage and question distributions are very different. Surprisingly, even downsampled human-constructed benchmarks (i.e., collecting less data) and programmatically-generated benchmarks (e.g., cloze-formatted examples) have high concurrence with human-constructed benchmarks. These results indicate that, despite years of intense community focus on a small number of benchmarks, the modeling improvements studied hold broadly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- How Reliable is Language Model Micro-Benchmarking?Gregory Yauney, Shahzaib Saqib Warraich, Swabha SwayamdiptaICLR 2026 · 7 citations
- Collaborative Performance Prediction for Large Language ModelsQiyuan Zhang, Fuyuan Lyu, Xue Liu, Chen MaEMNLP 2024 · 1 citation
- Correct Looks Better: Pairwise Comparisons Reveal Accuracy RankingsMina Remeli, Moritz HardtICML 2026
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot et al.ICLR 2020 · 244 citations
- The Effect of Natural Distribution Shift on Question Answering ModelsJohn Miller, Karl Krauth, Benjamin Recht, Ludwig SchmidtICML 2020 · 158 citations
Related papers
- Examining the robustness of LLM evaluation to the distributional assumptions of benchmarksCharlotte Siska, Katerina Marazopoulou, Melissa Ailem, James BonoACL 2024
- Small Models Exhibit Limited Answer Consistency in Repetition Trials of the Multiple-Choice MMLU-Redux and MedQA BenchmarksClaudio S. Pinhanez, Paulo R. Cavalin, Cassia Sampaio Sanctos, Marcelo Carpinette GraveAAAI 2026
- Unveiling the Tapestry of Consistency in Large Vision-Language ModelsYuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan et al.NeurIPS 2024 · 27 citations
- Train-before-Test Harmonizes Language Model RankingsGuanhua Zhang, Ricardo Dominguez-Olmedo, Moritz HardtICLR 2026 · 13 citations
- QuanBench: Benchmarking Quantum Code Generation with Large Language ModelsXiaoyu Guo, Minggu Wang, Jianjun ZhaoASE 2025 · 5 citations
