VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts
Xin Liu, Lechen Zhang, Sheza Munir, Yiyang Gu, Lu Wang
摘要
Large language models (LLMs) excel at generating long-form responses, but evaluating their factuality remains challenging due to complex inter-sentence dependencies within the generated facts. Prior solutions predominantly follow a decompose-decontextualize-verify pipeline but often fail to capture essential context and miss key relational facts. In this paper, we introduce VERIFACT, a factuality evaluation framework designed to enhance fact extraction by identifying and resolving incomplete and missing facts to support more accurate verification results. Moreover, we introduce FACTR-BENCH 1 , a benchmark that evaluates both precision and recall in long-form model responses, whereas prior work primarily focuses on precision. FACTRBENCH provides reference fact sets from advanced LLMs and human-written answers, enabling recall assessment. Empirical evaluations show that VERIFACT significantly enhances fact completeness and preserves complex facts with critical relational information, resulting in more accurate factuality evaluation. Benchmarking various open-and close-weight LLMs on FACTRBENCH indicate that larger models within same model family improve precision and recall, but high precision does not always correlate with high recall, underscoring the importance of comprehensive factuality assessment. If there wasn't a demand for gold as jewelry, would its price drop enough to make it usable in consumer electronics? If the demand for gold as jewelry were to disappear, the price of gold could drop by 20-50% or more, making it more competitive with other materials like copper and silver. 1. There is a demand for gold as jewelry Anthropic. 2024. Claude 3.5 sonnet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DeepFact: Co-Evolving Benchmarks and Agents for Deep Research FactualityYukun Huang, Leonardo F. R. Ribeiro, Momchil Hardalov, Bhuwan Dhingra 等ACL 2026 · 被引用 2 次
- IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural ThinkingZechen Sun, Yuyang Sun, Zecheng Tang, Juntao Li 等ACL 2026
它引用的顶会 Paper4
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie 等EMNLP 2023 · 被引用 224 次
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu 等NeurIPS 2024 · 被引用 182 次
相关 Paper
- FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality EvaluationFarima Fatahi Bayat, Lechen Zhang, Sheza Munir, Lu WangACL 2025
- Towards Effective Extraction and Evaluation of Factual ClaimsDasha Metropolitansky, Jonathan LarsonACL 2025 · 被引用 17 次
- Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual SettingsAustin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz 等ACL 2025 · 被引用 17 次
- VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across DomainsXuzhao Li, Xuchen Li, Shiyu Hu, Yongzhen Guo 等AAAI 2026 · 被引用 16 次
- FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language ModelsJavier Carnerero-Cano, Massimiliano Pronesti, Radu Marinescu, Tigran T. Tchrakian 等ACL 2026
