A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based Solution
Dongning Rao, Rongchu Zhou, Peng Chen, Zhihua Jiang
摘要
Low-resource language understanding is a challenging task, even for large language models (LLMs). An epitome of this problem is the CompRehensive lIterary chineSe readIng comprehenSion (CRISIS), whose difficulties include limited linguistic data, long input, and insight-required questions. Besides the compelling need to provide a larger dataset for CRISIS, excessive information, order bias, and entangled conundrums still plague the CRISIS solutions. Thus, we present the eVIdence cuRation with opTion shUffling and Abstract meaning representation-based cLauses segmenting (VIRTUAL) procedure for CRISIS, with the most extensive dataset. While the dataset is also named CRISIS, it results from a three-phase construction process, including question selection, data cleaning, and a silver-standard data augmentation step, which augments translations, celebrity profiles, government jobs, reign mottos, and dynasty to CRISIS. The six steps of VIRTUAL include embedding, shuffling, abstract meaning representation-based option segmenting, evidence extraction, solving, and voting. Notably, the evidence extraction algorithm facilitates the extraction of literary Chinese evidence sentences, translated evidence sentences, and annotations of keywords using a similaritybased ranking strategy. While CRISIS compiles understanding-required questions from seven sources, the experiments on CRISIS substantiate the effectiveness of VIRTUAL, with a 7 percent increase in accuracy compared to the baseline. Interestingly, both non-LLMs and LLMs exhibit order bias, and abstract meaning representation-based option segmenting is beneficial for CRISIS. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan 等NeurIPS 2023 · 被引用 420 次
- English Machine Reading Comprehension Datasets: A SurveyDaria Dzendzik, Jennifer Foster, Carl VogelEMNLP 2021 · 被引用 8 次
- A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasksElie Antoine, Frédéric Béchet, Géraldine Damnati, Philippe LanglaisEMNLP 2024 · 被引用 2 次
- Evaluating the Rationale Understanding of Critical Reasoning in Logical Reading ComprehensionAkira Kawabata, Saku SugawaraEMNLP 2023
相关 Paper
- CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language ModelsLing Shi, Deyi XiongACL 2025 · 被引用 5 次
- TongGu-VL: Advancing Visual-Language Understanding in Chinese Classical Studies through Parameter Sensitivity-Guided Instruction TuningJiahuan Cao, Yang Liu, Peirong Zhang, Yongxin Shi 等ACM MM 2025
- Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and ReasoningChongyuan Dai, Jinpeng Hu, Hongchang Shi, Zhuo Li 等ACL 2026 · 被引用 18 次
- ReCO: A Large Scale Chinese Reading Comprehension Dataset on OpinionBingning Wang, Ting Yao, Qi Zhang, Jingfang Xu 等AAAI 2020 · 被引用 26 次
- AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language ModelsLian Yan, Haotian Wang, Chen Tang, Haifeng Liu 等AAAI 2026 · 被引用 3 次
