COSMIC: Mutual Information for Task-Agnostic Summarization Evaluation
Maxime Darrin, Philippe Formont, Jackie Chi Kit Cheung, Pablo Piantanida
摘要
Assessing the quality of summarizers poses significant challenges-gold summaries are hard to obtain and their suitability depends on the use context of the summarization system. Who is the user of the system, and what do they intend to do with the summary? In response, we propose a novel task-oriented evaluation approach that assesses summarizers based on their capacity to produce summaries while preserving task outcomes. We theoretically establish both a lower and upper bound on the expected error rate of these tasks, which depends on the mutual information between source texts and generated summaries. We introduce COSMIC, a practical implementation of this metric, and demonstrate its strong correlation with human judgment-based metrics, as well as its effectiveness in predicting downstream task performance. Comparative analyses against established metrics like BERTScore and ROUGE highlight the competitive performance of COSMIC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- GraphNarrator: Generating Textual Explanations for Graph Neural NetworksBo Pan, Zhen Xiong, Guanchen Wu, Zheng Zhang 等ACL 2025 · 被引用 7 次
- Learning Task-Agnostic Representations through Multi-Teacher DistillationPhilippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger 等NeurIPS 2025 · 被引用 6 次
- An Information Theoretic Perspective on Agentic System DesignShizhe He, Avanika Narayan, Ishan S. Khare, Scott W. Linderman 等ICLR 2026 · 被引用 6 次
- Information Estimation with Discrete DiffusionAlberto Foresti, Giulio Franzese, Pietro MichiardiICLR 2026
它引用的顶会 Paper11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 被引用 337 次
相关 Paper
- MTAS: A Reference-Free Approach for Evaluating Abstractive Summarization SystemsXiaoyan Zhu, Mingyue Jiang, Xiao-Yi Zhang, Liming Nie 等FSE 2024 · 被引用 2 次
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu 等EMNLP 2020 · 被引用 3 次
- Unsupervised Reference-Free Summary Quality Evaluation via Contrastive LearningHanlu Wu, Tengfei Ma, Lingfei Wu, Tariro Manyumwa 等EMNLP 2020 · 被引用 47 次
- Extrinsic Evaluation of Machine Translation MetricsNikita Moghe, Tom Sherborne, Mark Steedman, Alexandra BirchACL 2023 · 被引用 12 次
- SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at ScaleNaman Bansal, Mousumi Akter, Shubhra Kanti Karmaker SantuEMNLP 2022 · 被引用 2 次
