Long-Form Information Alignment Evaluation Beyond Atomic Facts
Danna Zheng, Mirella Lapata, Jeff Z. Pan
Abstract
Information alignment evaluators are vital for various NLG evaluation tasks and trustworthy LLM deployment, reducing hallucinations and enhancing user trust. Current fine-grained methods, like FactScore, verify facts individually but neglect inter-fact dependencies, enabling subtle vulnerabilities. In this work, we introduce MONTAGELIE, a challenging benchmark that constructs deceptive narratives by "montaging" truthful statements without introducing explicit hallucinations. We demonstrate that both coarse-grained LLM-based evaluators and current fine-grained frameworks are susceptible to this attack, with AUC-ROC scores falling below 65%. To enable more robust fine-grained evaluation, we propose DOVESCORE, a novel framework that jointly verifies factual accuracy and event-order consistency. By modeling inter-fact relationships, DOVESCORE outperforms existing finegrained methods by over 8%, providing a more robust solution for long-form text alignment evaluation. Our code and datasets are available at https://github.com/dannalily/DoveScore . -Mike and Amy broke up. -Amy went to movies with John. -Mike hit Amy. Mike hit Amy. Mike and Amy broke up. Amy went to movies with John. Truth Montage Lie Amy went to movies with John. Mike hit Amy. Mike and Amy broke up.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d445a01a-da6f-4f2f-a90e-4f4ee723eb91Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis et al.EMNLP 2023 · 225 citations
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu et al.NeurIPS 2024 · 182 citations
Related papers
- AlignScore: Evaluating Factual Consistency with A Unified Alignment FunctionYuheng Zha, Yichi Yang, Ruichen Li, Zhiting HuACL 2023 · 44 citations
- FactVerse: A Benchmark for Factual Consistency in Interleaved Image-Text GenerationYubo Shan, Kun Zhang, Qiming Xu, Liping Cao et al.ACL 2026
- Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMsYuzhe Gu, Wenwei Zhang, Chengqi Lyu, Dahua Lin et al.ICLR 2025
- DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text GenerationMiriam Wanner, Benjamin Van Durme, Mark DredzeEMNLP 2025 · 17 citations
- VISTA: Verification In Sequential Turn-based AssessmentAshley Lewis, Andrew Perrault, Eric Fosler-Lussier, Michael WhiteACL 2026
