MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
Liyan Tang, Philippe Laban, Greg Durrett
Abstract
Recognizing if LLM output can be grounded in evidence is central to many tasks in NLP: retrieval-augmented generation, summarization, document-grounded dialogue, and more. Current approaches to this kind of factchecking are based on verifying each piece of a model generation against potential evidence using an LLM. However, this process can be very computationally expensive, requiring many calls to a model to check a single response. In this work, we show how to build small fact-checking models that have GPT-4level performance but for 400x lower cost. We do this by constructing synthetic training data with GPT-4, which involves creating realistic yet challenging instances of factual errors via a structured generation procedure. Training on this data teaches models to check each fact in the claim and recognize synthesis of information across sentences. For evaluation, we unify datasets from recent work on factchecking and grounding LLM generations into a new benchmark, LLM-AGGREFACT. Our best system MiniCheck-FT5 (770M parameters) outperforms all systems of comparable size and reaches GPT-4 accuracy. We release LLM-AGGREFACT, code for data synthesis, and models. 1 Chunk Chunk 2 Abla#on 1 Chunk 2 Abla#on 2 Chunk 1 Chunk 2 Chunk 3 Chunk 2 Abl. 1 Chunk 2 Abl. 2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5f2f464-9a6f-44b2-9652-fc30ddb2a92dCited by top-tier papers40
- QFFT, Question-Free Fine-Tuning for Adaptive ReasoningWanlong Liu, Junxiao Xu, Fei Yu, Yukang Lin et al.NeurIPS 2025 · 26 citations
- AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon ConversationsCheng Jiayang, Dongyu Ru, Lin Qiu, Yiyang Li et al.ICLR 2026 · 21 citations
- Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsPhilippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng WuEMNLP 2024 · 19 citations
- DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text GenerationMiriam Wanner, Benjamin Van Durme, Mark DredzeEMNLP 2025 · 17 citations
- Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual SettingsAustin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz et al.ACL 2025 · 17 citations
Builds on23
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 531 citations
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu et al.ICML 2024 · 406 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 329 citations
Related papers
- Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking StrategiesYuxuan Ye, Raúl Santos-Rodríguez, Edwin SimpsonACL 2026
- Improving Model Factuality with Fine-grained Critique-based EvaluatorYiqing Xie, Wenxuan Zhou, Pradyot Prakash, Di Jin et al.ACL 2025 · 16 citations
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu et al.NeurIPS 2024 · 182 citations
- SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of SummarizationPhilippe Laban, Wojciech Kryscinski, Divyansh Agarwal, Alexander R. Fabbri et al.EMNLP 2023 · 29 citations
- Evaluating Language Models as Synthetic Data GeneratorsSeungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan et al.ACL 2025
