A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based Solution
Dongning Rao, Rongchu Zhou, Peng Chen, Zhihua Jiang
Abstract
Low-resource language understanding is a challenging task, even for large language models (LLMs). An epitome of this problem is the CompRehensive lIterary chineSe readIng comprehenSion (CRISIS), whose difficulties include limited linguistic data, long input, and insight-required questions. Besides the compelling need to provide a larger dataset for CRISIS, excessive information, order bias, and entangled conundrums still plague the CRISIS solutions. Thus, we present the eVIdence cuRation with opTion shUffling and Abstract meaning representation-based cLauses segmenting (VIRTUAL) procedure for CRISIS, with the most extensive dataset. While the dataset is also named CRISIS, it results from a three-phase construction process, including question selection, data cleaning, and a silver-standard data augmentation step, which augments translations, celebrity profiles, government jobs, reign mottos, and dynasty to CRISIS. The six steps of VIRTUAL include embedding, shuffling, abstract meaning representation-based option segmenting, evidence extraction, solving, and voting. Notably, the evidence extraction algorithm facilitates the extraction of literary Chinese evidence sentences, translated evidence sentences, and annotations of keywords using a similaritybased ranking strategy. While CRISIS compiles understanding-required questions from seven sources, the experiments on CRISIS substantiate the effectiveness of VIRTUAL, with a 7 percent increase in accuracy compared to the baseline. Interestingly, both non-LLMs and LLMs exhibit order bias, and abstract meaning representation-based option segmenting is beneficial for CRISIS. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78bf3f38-d962-4d28-a74d-3e2f1d699781Builds on4
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan et al.NeurIPS 2023 · 420 citations
- English Machine Reading Comprehension Datasets: A SurveyDaria Dzendzik, Jennifer Foster, Carl VogelEMNLP 2021 · 8 citations
- A linguistically-motivated evaluation methodology for unraveling model's abilities in reading comprehension tasksElie Antoine, Frédéric Béchet, Géraldine Damnati, Philippe LanglaisEMNLP 2024 · 2 citations
- Evaluating the Rationale Understanding of Critical Reasoning in Logical Reading ComprehensionAkira Kawabata, Saku SugawaraEMNLP 2023
Related papers
- CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language ModelsLing Shi, Deyi XiongACL 2025 · 5 citations
- TongGu-VL: Advancing Visual-Language Understanding in Chinese Classical Studies through Parameter Sensitivity-Guided Instruction TuningJiahuan Cao, Yang Liu, Peirong Zhang, Yongxin Shi et al.ACM MM 2025
- Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and ReasoningChongyuan Dai, Jinpeng Hu, Hongchang Shi, Zhuo Li et al.ACL 2026 · 18 citations
- ReCO: A Large Scale Chinese Reading Comprehension Dataset on OpinionBingning Wang, Ting Yao, Qi Zhang, Jingfang Xu et al.AAAI 2020 · 26 citations
- AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language ModelsLian Yan, Haotian Wang, Chen Tang, Haifeng Liu et al.AAAI 2026 · 3 citations
