Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
Seonjeong Hwang, Hyounghun Kim, Gary Lee
摘要
Estimating the cognitive complexity of reading comprehension (RC) items is crucial for assessing item difficulty before it is administered to learners. Unlike syntactic and semantic features, such as passage length or semantic similarity between options, cognitive features that arise during answer reasoning are not readily extractable using existing NLP tools and have traditionally relied on human annotation. In this study, we examine whether large language models (LLMs) can estimate the cognitive complexity of RC items by focusing on two dimensions-Evidence Scope and Transformation Level-that indicate the degree of cognitive burden involved in reasoning about the answer. Our experimental results demonstrate that LLMs can approximate the cognitive complexity of items, indicating their potential as tools for prior difficulty analysis. Further analysis reveals a gap between LLMs' reasoning ability and their metacognitive awareness: even when they produce correct answers, they sometimes fail to correctly identify the features underlying their own reasoning process. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative ComprehensionYing Xu, Dakuo Wang, Mo Yu, Daniel Ritchie 等ACL 2022 · 被引用 131 次
- Hierarchical Deconstruction of LLM Reasoning: A Graph-Based Framework for Analyzing Knowledge UtilizationMiyoung Ko, Sue Hyun Park, Joonsuk Park, Minjoon SeoEMNLP 2024 · 被引用 3 次
相关 Paper
- A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasksElie Antoine, Frédéric Béchet, Géraldine Damnati, Philippe LanglaisEMNLP 2024 · 被引用 2 次
- NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity ClassesLizhou Fan, Wenyue Hua, Lingyao Li, Haoyang Ling 等ACL 2024 · 被引用 8 次
- SMART: Evaluating LLMs' Mathematical Reasoning via a Human Cognitive Process-Inspired BenchmarkYujie Hou, Mei Wang, Yaoyao Zhong, Ting Zhang 等ACL 2026
- Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?Neeladri Bhuiya, Viktor Schlegel, Stefan WinklerEMNLP 2024 · 被引用 2 次
- A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item GenerationSeonjeong Hwang, Jun Seo, Hyounghun Kim, Gary LeeACL 2026
