What Question Answering can Learn from Trivia Nerds
Jordan L. Boyd-Graber, Benjamin Börschinger
Abstract
Question answering (QA)is not just building systems; this NLP subfield also creates and curates challenging question datasets that reveal the best systems. We argue that QA datasets-and QA leaderboards-closely resemble trivia tournaments: the questions agents-humans or machines-answer reveals a "winner". However, the research community has ignored the lessons from decades of the trivia community creating vibrant, fair, and effective QA competitions. After detailing problems with existing QA datasets, we outline several lessons that transfer to QA research: removing ambiguity, identifying better QA agents, and adjudicating disputes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d2296f8-83f3-45d7-b107-9bd5a5d22682Cited by top-tier papers5
- FiD-Light: Efficient and Effective Retrieval-Augmented Text GenerationSebastian Hofstätter, Jiecao Chen, Karthik Raman, Hamed ZamaniSIGIR 2023 · 47 citations
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- You Make me Feel like a Natural Question: Training QA Systems on Transformed Trivia QuestionsTasnim Kabir, Yoo Yeon Sung, Saptarashmi Bandyopadhyay, Hao Zou et al.EMNLP 2024
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
- Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor et al.ACL 2021
Builds on4
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 92 citations
- Build It, Break It, Fix It: Contesting Secure DevelopmentAndrew Ruef, Michael W. Hicks, James Parker, Dave Levin et al.CCS 2016 · 80 citations
- Interactive Machine Comprehension with Information Seeking AgentsXingdi Yuan, Jie Fu, Marc-Alexandre Côté, Yi Tay et al.ACL 2020 · 11 citations
Related papers
- What do Models Learn from Question Answering Datasets?Priyanka Sen, Amir SaffariEMNLP 2020 · 40 citations
- Do Question Answering Modeling Improvements Hold Across Benchmarks?Nelson F. Liu, Tony Lee, Robin Jia, Percy LiangACL 2023 · 1 citation
- How to Compare Things Properly? A Study of Argument Relevance in Comparative Question AnsweringIrina Nikishina, Saba Anwar, Nikolay Dolgov, Maria Manina et al.ACL 2025
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 51 citations
