ReCO: A Large Scale Chinese Reading Comprehension Dataset on Opinion
Bingning Wang, Ting Yao, Qi Zhang, Jingfang Xu, Xiaochuan Wang
Abstract
This paper presents the ReCO, a human-curated Chinese Reading Comprehension dataset on Opinion. The questions in ReCO are opinion based queries issued to commercial search engine. The passages are provided by the crowdworkers who extract the support snippet from the retrieved documents. Finally, an abstractive yes/no/uncertain answer was given by the crowdworkers. The release of ReCO consists of 300k questions that to our knowledge is the largest in Chinese reading comprehension. A prominent characteristic of ReCO is that in addition to the original context paragraph, we also provided the support evidence that could be directly used to answer the question. Quality analysis demonstrates the challenge of ReCO that it requires various types of reasoning skills such as causal inference, logical reasoning, etc. Current QA models that perform very well on many question answering problems, such as BERT (Devlin et al. 2018), only achieves 77% accuracy on this dataset, a large margin behind humans nearly 92% performance, indicating ReCO present a good challenge for machine reading comprehension. The codes, dataset and leaderboard will be freely available at https://github.com/benywon/ReCO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 287cc3e8-18cf-45d1-abb9-ec61e9d339fcCited by top-tier papers4
- SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence UnderstandingTianyu Yu, Chengyue Jiang, Chao Lou, Shen Huang et al.AAAI 2024 · 30 citations
- English Machine Reading Comprehension Datasets: A SurveyDaria Dzendzik, Jennifer Foster, Carl VogelEMNLP 2021 · 8 citations
- RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-AgentsJize Wang, Han Wu, Zhiyuan You, Yiming Song et al.ACL 2026 · 2 citations
- Enhancing Document Understanding with Group Position Embedding: A Novel Approach to Incorporate Layout InformationYuke Zhu, Yue Zhang, Dongdong Liu, Chi Xie et al.ICLR 2025
Builds on1
Related papers
- JEC-QA: A Legal-Domain Question Answering DatasetHaoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang et al.AAAI 2020 · 212 citations
- Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionHongyu Gong, Yelong Shen, Dian Yu, Jianshu Chen et al.ACL 2020 · 39 citations
- CRT-QA: A Dataset of Complex Reasoning Question Answering over Tabular DataZhehao Zhang, Xitao Li, Yan Gao, Jian-Guang LouEMNLP 2023 · 3 citations
- ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic RelationsRujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon et al.EMNLP 2021 · 30 citations
- MLEC-QA: A Chinese Multi-Choice Biomedical Question Answering DatasetJing Li, Shangping Zhong, Kaizhi ChenEMNLP 2021 · 24 citations
