CFEVER: A Chinese Fact Extraction and VERification Dataset
Ying-Jia Lin, Chun-Yi Lin, Chia-Jen Yeh, Yi-Ting Li, Yun-Yu Hu, Chih-Hao Hsu, Mei-Feng Lee, Hung-Yu Kao
摘要
We present CFEVER, a Chinese dataset designed for Fact Extraction and VERification. CFEVER comprises 30,012 manually created claims based on content in Chinese Wikipedia. Each claim in CFEVER is labeled as "Supports", "Refutes", or "Not Enough Info" to depict its degree of factualness. Similar to the FEVER dataset, claims in the "Supports" and "Refutes" categories are also annotated with corresponding evidence sentences sourced from single or multiple pages in Chinese Wikipedia. Our labeled dataset holds a Fleiss' kappa value of 0.7934 for five-way inter-annotator agreement. In addition, through the experiments with the state-of-the-art approaches developed on the FEVER dataset and a simple baseline for CFEVER, we demonstrate that our dataset is a new rigorous benchmark for factual extraction and verification, which can be further used for developing automated systems to alleviate human fact-checking efforts. CFEVER is available at https://ikmlab.github.io/CFEVER .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation DetectionHerun Wan, Jiaying Wu, Minnan Luo, Zhi Zeng 等NeurIPS 2025 · 被引用 14 次
- RaCMC: Residual-Aware Compensation Network with Multi-Granularity Constraints for Fake News DetectionXinquan Yu, Ziqi Sheng, Wei Lu, Xiangyang Luo 等AAAI 2025 · 被引用 9 次
- How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and AnalysisHerun Wan, Minnan Luo, Zihan Ma, Guang Dai 等EMNLP 2025 · 被引用 3 次
- TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-CheckingXiaocheng Zhang, Xi Wang, Yifei Lu, Jianing Wang 等ACL 2026
它引用的顶会 Paper5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Mining Dual Emotion for Fake News DetectionXueyao Zhang, Juan Cao, Xirong Li, Qiang Sheng 等WWW 2021 · 被引用 332 次
- Explainable Automated Fact-Checking for Public Health ClaimsNeema Kotonya, Francesca ToniEMNLP 2020 · 被引用 10 次
相关 Paper
- DeSePtion: Dual Sequence Prediction and Adversarial Examples for Improved Fact-CheckingChristopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia 等ACL 2020 · 被引用 7 次
- DialFact: A Benchmark for Fact-Checking in DialoguePrakhar Gupta, Chien-Sheng Wu, Wenhao Liu, Caiming XiongACL 2022
- Reasoning Over Semantic-Level Graph for Fact CheckingWanjun Zhong, Jingjing Xu, Duyu Tang, Zenan Xu 等ACL 2020 · 被引用 154 次
- Fine-grained Fact Verification with Kernel Graph Attention NetworkZhenghao Liu, Chenyan Xiong, Maosong Sun, Zhiyuan LiuACL 2020 · 被引用 9 次
- Unsupervised Pretraining for Fact Verification by Language Model DistillationAdrián Bazaga, Pietro Lio, Gos MicklemICLR 2024 · 被引用 5 次
