What Question Answering can Learn from Trivia Nerds
Jordan L. Boyd-Graber, Benjamin Börschinger
2020年份
3被引次数
5顶会引用
摘要
Question answering (QA)is not just building systems; this NLP subfield also creates and curates challenging question datasets that reveal the best systems. We argue that QA datasets-and QA leaderboards-closely resemble trivia tournaments: the questions agents-humans or machines-answer reveals a "winner". However, the research community has ignored the lessons from decades of the trivia community creating vibrant, fair, and effective QA competitions. After detailing problems with existing QA datasets, we outline several lessons that transfer to QA research: removing ambiguity, identifying better QA agents, and adjudicating disputes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- FiD-Light: Efficient and Effective Retrieval-Augmented Text GenerationSebastian Hofstätter, Jiecao Chen, Karthik Raman, Hamed ZamaniSIGIR 2023 · 被引用 47 次
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe 等ACL 2024 · 被引用 30 次
- You Make me Feel like a Natural Question: Training QA Systems on Transformed Trivia QuestionsTasnim Kabir, Yoo Yeon Sung, Saptarashmi Bandyopadhyay, Hao Zou 等EMNLP 2024
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
- Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor 等ACL 2021
它引用的顶会 Paper4
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 被引用 162 次
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 被引用 92 次
- Build It, Break It, Fix It: Contesting Secure DevelopmentAndrew Ruef, Michael W. Hicks, James Parker, Dave Levin 等CCS 2016 · 被引用 80 次
- Interactive Machine Comprehension with Information Seeking AgentsXingdi Yuan, Jie Fu, Marc-Alexandre Côté, Yi Tay 等ACL 2020 · 被引用 11 次
相关 Paper
- What do Models Learn from Question Answering Datasets?Priyanka Sen, Amir SaffariEMNLP 2020 · 被引用 40 次
- Do Question Answering Modeling Improvements Hold Across Benchmarks?Nelson F. Liu, Tony Lee, Robin Jia, Percy LiangACL 2023 · 被引用 1 次
- How to Compare Things Properly? A Study of Argument Relevance in Comparative Question AnsweringIrina Nikishina, Saba Anwar, Nikolay Dolgov, Maria Manina 等ACL 2025
- SQUAB: Evaluating LLM robustness to Ambiguous and Unanswerable Questions in Semantic ParsingSimone Papicchio, Luca Cagliero, Paolo PapottiEMNLP 2025
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 被引用 51 次
