MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
Liyan Tang, Philippe Laban, Greg Durrett
摘要
Recognizing if LLM output can be grounded in evidence is central to many tasks in NLP: retrieval-augmented generation, summarization, document-grounded dialogue, and more. Current approaches to this kind of factchecking are based on verifying each piece of a model generation against potential evidence using an LLM. However, this process can be very computationally expensive, requiring many calls to a model to check a single response. In this work, we show how to build small fact-checking models that have GPT-4level performance but for 400x lower cost. We do this by constructing synthetic training data with GPT-4, which involves creating realistic yet challenging instances of factual errors via a structured generation procedure. Training on this data teaches models to check each fact in the claim and recognize synthesis of information across sentences. For evaluation, we unify datasets from recent work on factchecking and grounding LLM generations into a new benchmark, LLM-AGGREFACT. Our best system MiniCheck-FT5 (770M parameters) outperforms all systems of comparable size and reaches GPT-4 accuracy. We release LLM-AGGREFACT, code for data synthesis, and models. 1 Chunk Chunk 2 Abla#on 1 Chunk 2 Abla#on 2 Chunk 1 Chunk 2 Chunk 3 Chunk 2 Abl. 1 Chunk 2 Abl. 2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- QFFT, Question-Free Fine-Tuning for Adaptive ReasoningWanlong Liu, Junxiao Xu, Fei Yu, Yukang Lin 等NeurIPS 2025 · 被引用 26 次
- AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon ConversationsCheng Jiayang, Dongyu Ru, Lin Qiu, Yiyang Li 等ICLR 2026 · 被引用 21 次
- Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsPhilippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng WuEMNLP 2024 · 被引用 19 次
- DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text GenerationMiriam Wanner, Benjamin Van Durme, Mark DredzeEMNLP 2025 · 被引用 17 次
- Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual SettingsAustin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz 等ACL 2025 · 被引用 17 次
它引用的顶会 Paper23
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 被引用 531 次
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu 等ICML 2024 · 被引用 406 次
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 被引用 394 次
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 被引用 329 次
相关 Paper
- Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking StrategiesYuxuan Ye, Raúl Santos-Rodríguez, Edwin SimpsonACL 2026
- Improving Model Factuality with Fine-grained Critique-based EvaluatorYiqing Xie, Wenxuan Zhou, Pradyot Prakash, Di Jin 等ACL 2025 · 被引用 16 次
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu 等NeurIPS 2024 · 被引用 182 次
- SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri 等EMNLP 2023 · 被引用 29 次
- Evaluating Language Models as Synthetic Data GeneratorsSeungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan 等ACL 2025
