InfoLM: A New Metric to Evaluate Summarization & Data2Text Generation
Pierre Jean A. Colombo, Chloé Clavel, Pablo Piantanida
Abstract
Assessing the quality of natural language generation (NLG) systems through human annotation is very expensive. Additionally, human annotation campaigns are time-consuming and include non-reusable human labour. In practice, researchers rely on automatic metrics as a proxy of quality. In the last decade, many string-based metrics (e.g., BLEU or ROUGE) have been introduced. However, such metrics usually rely on exact matches and thus, do not robustly handle synonyms. In this paper, we introduce InfoLM a family of untrained metrics that can be viewed as a string-based metric that addresses the aforementioned flaws thanks to a pre-trained masked language model. This family of metrics also makes use of information measures allowing the possibility to adapt InfoLM to different evaluation criteria. Using direct assessment, we demonstrate that InfoLM achieves statistically significant improvement and two figure correlation gains in many configurations compared to existing metrics on both summarization and data2text generation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6e9cd9f-10b1-44fc-b1e9-46399e6f72dbCited by top-tier papers12
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry et al.NeurIPS 2022 · 24 citations
- Inherent Trade-Offs between Diversity and Stability in Multi-Task BenchmarksGuanhua Zhang, Moritz HardtICML 2024 · 22 citations
- What are the best Systems? New Perspectives on NLP BenchmarkingPierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan ClémençonNeurIPS 2022 · 20 citations
- CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model GenerationPei Ke, Bosi Wen, Andrew Feng, Xiao Liu et al.ACL 2024 · 9 citations
- GraphNarrator: Generating Textual Explanations for Graph Neural NetworksBo Pan, Zhen Xiong, Guanchen Wu, Zheng Zhang et al.ACL 2025 · 7 citations
Builds on4
- Heavy-tailed Representations, Text Polarity Classification & Data AugmentationHamid Jalalzai, Pierre Colombo, Chloé Clavel, Éric Gaussier et al.NeurIPS 2020 · 33 citations
- Automatic Text Evaluation through the Lens of Wasserstein BarycentersPierre Colombo, Guillaume Staerman, Chloé Clavel, Pablo PiantanidaEMNLP 2021 · 21 citations
- What are the best Systems? New Perspectives on NLP BenchmarkingPierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan ClémençonNeurIPS 2022 · 20 citations
- Re-evaluating Evaluation in Text SummarizationManik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu et al.EMNLP 2020 · 3 citations
Related papers
- Spurious Correlations in Reference-Free Evaluation of Text GenerationEsin Durmus, Faisal Ladhak, Tatsunori HashimotoACL 2022
- On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar et al.ACL 2023 · 10 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringPei Ke, Fei Huang, Fei Mi, Yasheng Wang et al.ACL 2023 · 2 citations
- Reassessing automatic evaluation metrics for code summarization tasksDevjeet Roy, Sarah Fakhoury, Venera ArnaoudovaFSE 2021 · 103 citations
