STARC: Structured Annotations for Reading Comprehension
Yevgeni Berzak, Jonathan Malmaud, Roger Levy
摘要
We present STARC (Structured Annotations for Reading Comprehension), a new annotation framework for assessing reading comprehension with multiple choice questions. Our framework introduces a principled structure for the answer choices and ties them to textual span annotations. The framework is implemented in OneStopQA, a new high-quality dataset for evaluation and analysis of reading comprehension in English. We use this dataset to demonstrate that STARC can be leveraged for a key new application for the development of SAT-like reading comprehension materials: automatic annotation quality probing via span ablation experiments. We further show that it enables in-depth analyses and comparisons between machine and human reading comprehension behavior, including error distributions and guessing ability. Our experiments also reveal that the standard multiple choice dataset in NLP, RACE (Lai et al., 2017) , is limited in its ability to measure reading comprehension. 47% of its questions can be guessed by machines without accessing the passage, and 18% are unanimously judged by humans as not having a unique correct answer. OneStopQA provides an alternative test set for reading comprehension which alleviates these shortcomings and has a substantially higher human ceiling performance. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Fine-Grained Prediction of Reading Comprehension from Eye MovementsOmer Shubi, Yoav Meiri, Cfir Avraham Hadar, Yevgeni BerzakEMNLP 2024 · 被引用 6 次
- Decoding Reading Goals from Eye MovementsOmer Shubi, Cfir Avraham Hadar, Yevgeni BerzakACL 2025 · 被引用 4 次
- Long-Tailed Question Answering in an Open WorldYi Dai, Hao Lang, Yinhe Zheng, Fei Huang 等ACL 2023 · 被引用 3 次
- Decoding Open-Ended Information Seeking Goals from Eye Movements in ReadingCfir Avraham Hadar, Omer Shubi, Yoav Meiri, Amit Heshes 等ICLR 2026 · 被引用 2 次
- SCOP: Evaluating the Comprehension Process of Large Language Models from a Cognitive ViewYongjie Xiao, Hongru Liang, Peixin Qin, Yao Zhang 等ACL 2025 · 被引用 1 次
相关 Paper
- WebSRC: A Dataset for Web-Based Structural Reading ComprehensionXingyu Chen, Zihan Zhao, Lu Chen, Jiabao Ji 等EMNLP 2021 · 被引用 42 次
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 被引用 92 次
- VisualMRC: Machine Reading Comprehension on Document ImagesRyota Tanaka, Kyosuke Nishida, Sen YoshidaAAAI 2021 · 被引用 201 次
- MMM: Multi-Stage Multi-Task Learning for Multi-Choice Reading ComprehensionDi Jin, Shuyang Gao, Jiun-Yu Kao, Tagyoung Chung 等AAAI 2020 · 被引用 72 次
- What do Models Learn from Question Answering Datasets?Priyanka Sen, Amir SaffariEMNLP 2020 · 被引用 40 次
