KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning
Peiqi Sui, Juan Diego Rodriguez, Philippe Laban, Dean Murphy, Joseph P. Dexter, Richard Jean So, Samuel Baker, Pramit Chaudhuri
摘要
Each year, tens of millions of essays are written and graded in college-level English courses. Students are asked to analyze literary and cultural texts through a process known as close reading, in which they gather textual details to formulate evidence-based arguments. Despite being viewed as a basis for critical thinking and widely adopted as a required element of university coursework, close reading has never been evaluated on large language models (LLMs), and multi-discipline benchmarks like MMLU do not include literature as a subject. To fill this gap, we present KRISTEVA, the first close reading benchmark 1 for evaluating interpretive reasoning, consisting of 1331 multiple-choice questions adapted from classroom data. With KRISTEVA, we propose three progressively more difficult sets of tasks to approximate different elements of the close reading process, which we use to test how well LLMs may seem to understand and reason about literary works: 1) extracting stylistic features, 2) retrieving relevant contextual information from parametric knowledge, and 3) multi-hop reasoning between style and external contexts. Our baseline results find that, while state-of-the-art LLMs possess some college-level close reading competency (accuracy 49.7% -69.7%), their performances still trail those of experienced human evaluators on 10 out of our 11 tasks. "It is not surprising that the detailed analysis of metaphors. . . sometimes feels like extracting cube-roots in the head." -I. A. Richards, (1936)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Critical Confabulation: Can LLMs Hallucinate for Social Good?Peiqi Sui, Eamon Duede, Hoyt Long, Richard Jean SoICLR 2026 · 被引用 2 次
- What Does AI Do for Cultural Interpretation? A Randomized Experiment on Close Reading Poems with Exposure to AI InterpretationJiayin Zhi, Hoyt Long, Richard Jean So, Mina LeeCHI 2026 · 被引用 1 次
它引用的顶会 Paper13
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 被引用 325 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 被引用 146 次
- Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4Kent K. Chang, Mackenzie Cramer, Sandeep Soni, David BammanEMNLP 2023 · 被引用 70 次
- IMPLI: Investigating NLI Models' Performance on Figurative LanguageKevin Stowe, Prasetya Ajie Utama, Iryna GurevychACL 2022 · 被引用 52 次
相关 Paper
- StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich TextZhouhong Gu, Haoning Ye, Xingzhou Chen, Zeyang Zhou 等ACL 2025
- Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?Neeladri Bhuiya, Viktor Schlegel, Stefan WinklerEMNLP 2024 · 被引用 2 次
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image ReasoningMingxin Huang, Yongxin Shi, Dezhi Peng, Songxuan Lai 等ICLR 2026 · 被引用 28 次
- Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language ModelsSophia Simeng Han, Howard Dai, Stephen Xia, Grant Zhang 等NeurIPS 2025 · 被引用 2 次
- MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular GraphsChristoph Bartmann, Johannes Schimunek, Mykyta Ielanskyi, Philipp Seidl 等ICLR 2026 · 被引用 5 次
