TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-Checking
Xiaocheng Zhang, Xi Wang, Yifei Lu, Jianing Wang, Zhuangzhuang Ye, Mengjiao Bao, Peng Yan, Xiaohong Su
摘要
With the surge of online misinformation, Large Language Models (LLMs) and Reasoning Large Language Models (RLMs) serving as Automatic Fact-Checking (AFC) systems have emerged as a prominent paradigm for reliable, explainable verification. However, our empirical study reveals that this paradigm faces a critical risk asymmetry challenge when deployed in the real world under resource-constrained environments. While Hotspot Perception Ability (HPA), the capacity to dynamically allocate reasoning resources based on social impact, is essential to mitigate this risk, existing benchmarks lack the social metadata and evaluation framework to meet this urgent evaluation needs, thereby hindering the advancement of these AFC systems. To bridge this gap, we introduce TrendFact, the first benchmark capable of evaluating HPA and three fact-checking tasks. It consists of 7,643 curated samples sourced from trending platforms and professional datasets, with an evidence library containing 366,634 entries. To enable HPA assessment, we propose two novel metrics: the Explanation Consistency Score (ECS) to evaluate the reliability of verification reasoning, and the Hotspot Claim Perception Index (HCPI) to quantify the overall HPA of AFC systems. Extensive experiments demonstrate that existing AFC systems exhibit limited performance on TrendFact. Furthermore, our proposed FactISR framework effectively enhances HPA and computational efficiency for RLMs-served AFC systems. * Equal contribution † Corresponding Author Evidence: The Forbidden City served as the imperial palace for 24 emperors during the Ming and Qing dynasties. Construction of the Forbidden City began in 1406, during the fourth year of Emperor Yongle's reign in the Ming Dynasty, and was completed in 1420, the 18th year of his reign. Situated at the heart of Beijing's central axis, the Forbidden City covers a land area of 720,000 square meters and has a built-up area of approximately 150,000 square meters...... Label: REFUTE Explanation: Evidence indicates that the construction of the Forbidden City began in the fourth year of Emperor Yongle's reign during the Ming Dynasty ( 1406 ) and was completed in the 18th year of his reign (1420). As of 2020, the Forbidden City has been built for 600 years. Therefore, this claim is incorrect.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Generating Fact Checking ExplanationsPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinACL 2020 · 被引用 130 次
- Reinforcement Learning-based Counter-Misinformation Response Generation: A Case Study of COVID-19 Vaccine MisinformationBing He, Mustaque Ahamad, Srijan KumarWWW 2023 · 被引用 62 次
- Fact-Checking Complex Claims with Program-Guided ReasoningLiangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu 等ACL 2023 · 被引用 45 次
- Generating Literal and Implied Subquestions to Fact-check Complex ClaimsJifan Chen, Aniruddh Sriram, Eunsol Choi, Greg DurrettEMNLP 2022 · 被引用 30 次
- Misinformation as a Harm: Structured Approaches for Fact-Checking PrioritizationConnie Moon Sehat, Ryan Li, Peipei Nie, Tarunima Prabhakar 等CSCW 2024 · 被引用 22 次
相关 Paper
- Beyond Detection: Evaluating Fallacy Awareness of LLMs in Interactive ScenariosConghui Niu, Ningxin Wu, Ziran Zhao, Dong Yu 等ACL 2026
- TripleFact: Defending Data Contamination in the Evaluation of LLM-driven Fake News DetectionCheng Xu, Nan YanACL 2025
- LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News DetectionCheng Xu, Changhong Jin, Yingjie Niu, Nan Yan 等ACL 2026
- Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language ModelsYingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao 等ACL 2025 · 被引用 7 次
- Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language ModelsYancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan 等ACL 2025
