SCOP: Evaluating the Comprehension Process of Large Language Models from a Cognitive View
Yongjie Xiao, Hongru Liang, Peixin Qin, Yao Zhang, Wenqiang Lei
Abstract
Despite the great potential of large language models (LLMs) in machine comprehension, it is still disturbing to fully count on them in realworld scenarios. This is probably because there is no rational explanation for whether the comprehension process of LLMs is aligned with that of experts. In this paper, we propose SCOP to carefully examine how LLMs perform during the comprehension process from a cognitive view. Specifically, it is equipped with a systematical definition of five requisite skills during the comprehension process, a strict framework to construct testing data for these skills, and a detailed analysis of advanced open-sourced and closed-sourced LLMs using the testing data. With SCOP, we find that it is still challenging for LLMs to perform an expert-level comprehension process. Even so, we notice that LLMs share some similarities with experts, e.g., performing better at comprehending local information than global information. Further analysis reveals that LLMs can be somewhat unreliable -they might reach correct answers through flawed comprehension processes. Based on SCOP, we suggest that one direction for improving LLMs is to focus more on the comprehension process, ensuring all comprehension skills are thoroughly developed during training 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Large Language Models as Commonsense Knowledge for Large-Scale Task PlanningZirui Zhao, Wee Sun Lee, David HsuNeurIPS 2023 · 423 citations
- Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4Kent K. Chang, Mackenzie Cramer, Sandeep Soni, David BammanEMNLP 2023 · 70 citations
- The ACL OCL Corpus: Advancing Open Science in Computational LinguisticsShaurya Rohatgi, Yanxia Qin, Benjamin Aw, Niranjana Unnithan et al.EMNLP 2023 · 9 citations
- To Test Machine Comprehension, Start by Defining ComprehensionJesse Dunietz, Gregory Burnham, Akash Bharadwaj, Owen Rambow et al.ACL 2020 · 7 citations
- Feeding What You Need by Understanding What You LearnedXiaoqiang Wang, Bang Liu, Fangli Xu, Bo Long et al.ACL 2022 · 6 citations
Related papers
- CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and GenerationWeixiang Yan, Haitian Liu, Yunkun Wang, Yunzhe Li et al.ACL 2024
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon et al.ICML 2024 · 197 citations
- Evaluating LLM Reasoning in the Operations Research Domain with ORQAMahdi Mostajabdaveh, Timothy Tin Long Yu, Samarendra Chandan Bindu Dash, Rindra Ramamonjison et al.AAAI 2025 · 1 citation
- Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length ContextsYuho Lee, Jiaqi Deng, Nicole Hee-Yeon Kim, Hyangsuk Min et al.EMNLP 2025
- Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?Neeladri Bhuiya, Viktor Schlegel, Stefan WinklerEMNLP 2024 · 2 citations
