STARC: Structured Annotations for Reading Comprehension
Yevgeni Berzak, Jonathan Malmaud, Roger Levy
Abstract
We present STARC (Structured Annotations for Reading Comprehension), a new annotation framework for assessing reading comprehension with multiple choice questions. Our framework introduces a principled structure for the answer choices and ties them to textual span annotations. The framework is implemented in OneStopQA, a new high-quality dataset for evaluation and analysis of reading comprehension in English. We use this dataset to demonstrate that STARC can be leveraged for a key new application for the development of SAT-like reading comprehension materials: automatic annotation quality probing via span ablation experiments. We further show that it enables in-depth analyses and comparisons between machine and human reading comprehension behavior, including error distributions and guessing ability. Our experiments also reveal that the standard multiple choice dataset in NLP, RACE (Lai et al., 2017) , is limited in its ability to measure reading comprehension. 47% of its questions can be guessed by machines without accessing the passage, and 18% are unanimously judged by humans as not having a unique correct answer. OneStopQA provides an alternative test set for reading comprehension which alleviates these shortcomings and has a substantially higher human ceiling performance. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2dae4447-ecee-4ac1-9e73-d8779d56ec7cCited by top-tier papers8
- Fine-Grained Prediction of Reading Comprehension from Eye MovementsOmer Shubi, Yoav Meiri, Cfir Avraham Hadar, Yevgeni BerzakEMNLP 2024 · 6 citations
- Decoding Reading Goals from Eye MovementsOmer Shubi, Cfir Avraham Hadar, Yevgeni BerzakACL 2025 · 4 citations
- Long-Tailed Question Answering in an Open WorldYi Dai, Hao Lang, Yinhe Zheng, Fei Huang et al.ACL 2023 · 3 citations
- Decoding Open-Ended Information Seeking Goals from Eye Movements in ReadingCfir Avraham Hadar, Omer Shubi, Yoav Meiri, Amit Heshes et al.ICLR 2026 · 2 citations
- SCOP: Evaluating the Comprehension Process of Large Language Models from a Cognitive ViewYongjie Xiao, Hongru Liang, Peixin Qin, Yao Zhang et al.ACL 2025 · 1 citation
Related papers
- WebSRC: A Dataset for Web-Based Structural Reading ComprehensionXingyu Chen, Zihan Zhao, Lu Chen, Jiabao Ji et al.EMNLP 2021 · 42 citations
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 92 citations
- VisualMRC: Machine Reading Comprehension on Document ImagesRyota Tanaka, Kyosuke Nishida, Sen YoshidaAAAI 2021 · 201 citations
- MMM: Multi-Stage Multi-Task Learning for Multi-Choice Reading ComprehensionDi Jin, Shuyang Gao, Jiun-Yu Kao, Tagyoung Chung et al.AAAI 2020 · 72 citations
- What do Models Learn from Question Answering Datasets?Priyanka Sen, Amir SaffariEMNLP 2020 · 40 citations
