Agent-as-Judge for Factual Summarization of Long Narratives
Yeonseok Jeong, Minsoo Kim, Seung-won Hwang, Byung-Hak Kim
摘要
Large Language Models (LLMs) have demonstrated near-human performance in summarization tasks based on traditional metrics such as ROUGE and BERTScore. However, these metrics do not adequately capture critical aspects of summarization quality, such as factual accuracy, particularly for long narratives (>100K tokens). Recent advances, such as LLM-as-a-Judge, address the limitations of metrics based on lexical similarity but still exhibit factual inconsistencies, especially in understanding character relationships and states. In this work, we introduce NARRATIVEFACTSCORE (NFS), the first "Agent-as-a-Judge" framework that evaluates and refines factuality in narrative summarization. By leveraging a Character Knowledge Graph (CKG) extracted from input narrative, NARRATIVEFACTSCORE evaluates the factuality and provides actionable guidance for refinement, such as identifying missing or erroneous facts. Our experimental results demonstrate that constructing the CKG enables reasoning with 1/3 of the factuality computation used in the prior approach, and achieve three times higher correlation with human judgments. Furthermore, refinement with actionable guidance improves the quality of the summary. 1 * Corresponding Author 1 https://github.com/YeonseokJeong/NarrativeFactScore # 14. BAG END LIVING ROOM … Bilbo: It's mine, my own. my precious (Frodo rushes into Bag End. He stops and picks up the ring at his feet.) … # 25. BAG END KITCHEN … Gandalf: Sauron needs only this ring to cover all the lands in the second darkness. He is seeking it, seeking it, all his thought is bent on it. … Frodo: Alright! ... input story Gandalf warned Frodo, who carries the Ring, that its master is Sauron. Sauron is searching for the Ring and is pursuing Gandalf.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- NexusSum: Hierarchical LLM Agents for Long-Form Narrative SummarizationHyuntak Kim, Byung-Hak KimACL 2025
- Agent Newsroom: Efficient Chronological Report Generation via Dynamic Multi-Agent CollaborationZhenhua Wang, Chunlei Wang, Yue Geng, Bang WangACL 2026
它引用的顶会 Paper13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 被引用 254 次
相关 Paper
- An Analysis of Multilingual FActScoreVu Trong Kim, Michael Krumdick, Varshini Reddy, Franck Dernoncourt 等EMNLP 2024 · 被引用 2 次
- How Far are We from Robust Long Abstractive Summarization?Huan Yee Koh, Jiaxin Ju, He Zhang, Ming Liu 等EMNLP 2022 · 被引用 16 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code SummarizationSuyoung Bae, CheolWon Na, Jaehoon Lee, Yumin Lee 等ACL 2026 · 被引用 1 次
- SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QASher Badshah, Ali Emami, Hassan SajjadACL 2026 · 被引用 1 次
