What Makes Reading Comprehension Questions Difficult?
Saku Sugawara, Nikita Nangia, Alex Warstadt, Samuel R. Bowman
摘要
For a natural language understanding benchmark to be useful in research, it has to consist of examples that are diverse and difficult enough to discriminate among current and near-future state-of-the-art systems. However, we do not yet know how best to select text sources to collect a variety of challenging examples. In this study, we crowdsource multiple-choice reading comprehension questions for passages taken from seven qualitatively distinct sources, analyzing what attributes of passages contribute to the difficulty and question types of the collected examples. To our surprise, we find that passage source, length, and readability measures do not significantly affect question difficulty. Through our manual annotation of seven reasoning types, we observe several trends between passage sources and reasoning types, e.g., logical reasoning is more often required in questions written for technical passages. These results suggest that when creating a new benchmark dataset, selecting a diverse set of passages can help ensure a diverse range of question types, but that passage difficulty need not be a priority.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Cascading Biases: Investigating the Effect of Heuristic Annotation Strategies on Data and ModelsChaitanya Malaviya, Sudeep Bhatia, Mark YatskarEMNLP 2022 · 被引用 3 次
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
它引用的顶会 Paper8
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 被引用 325 次
- Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real TasksAnna Rogers, Olga Kovaleva, Matthew Downey, Anna RumshiskyAAAI 2020 · 被引用 141 次
- TORQUE: A Reading Comprehension Dataset of Temporal Ordering QuestionsQiang Ning, Hao Wu, Rujun Han, Nanyun Peng 等EMNLP 2020 · 被引用 79 次
- What Can We Learn from Collective Human Opinions on Natural Language Inference Data?Yixin Nie, Xiang Zhou, Mohit BansalEMNLP 2020 · 被引用 77 次
相关 Paper
- Evaluating the Rationale Understanding of Critical Reasoning in Logical Reading ComprehensionAkira Kawabata, Saku SugawaraEMNLP 2023
- How Hard is this Test Set? NLI Characterization by Exploiting Training DynamicsAdrian Cosma, Stefan Ruseti, Mihai Dascalu, Cornelia CarageaEMNLP 2024
- Natural Language Inference in Context - Investigating Contextual Reasoning over Long TextsHanmeng Liu, Leyang Cui, Jian Liu, Yue ZhangAAAI 2021 · 被引用 57 次
- English Machine Reading Comprehension Datasets: A SurveyDaria Dzendzik, Jennifer Foster, Carl VogelEMNLP 2021 · 被引用 8 次
- What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt 等ACL 2021
