On the Usefulness of Embeddings, Clusters and Strings for Text Generation Evaluation
Tiago Pimentel, Clara Meister, Ryan Cotterell
摘要
A good automatic evaluation metric for language generation ideally correlates highly with human judgements of text quality. Yet, there is a dearth of such metrics, which inhibits the rapid and efficient progress of language generators. One exception is the recently proposed MAUVE. In theory, MAUVE measures an informationtheoretic divergence between two probability distributions over strings: one representing the language generator under evaluation and the other representing the true natural language distribution. MAUVE's authors argue that its success comes from the qualitative properties of their proposed divergence. Yet in practice, as this divergence is uncomputable, MAUVE approximates it by measuring the divergence between multinomial distributions over clusters instead, where cluster assignments are attained by grouping strings based on a pre-trained language model's embeddings. As we show, however, this is not a tight approximation-in either theory or practice. This begs the question: why does MAUVE work so well? In this work, we show that MAUVE was right for the wrong reasons, and that its newly proposed divergence is not necessary for its high performance. In fact, classical divergences paired with its proposed cluster-based approximation may actually serve as better evaluation metrics. We finish the paper with a probing analysis; this analysis leads us to conclude that-by encoding syntactic-and coherence-level features of text, while ignoring surface-level features-such cluster-based substitutes to string distributions may simply be better for evaluating state-of-the-art language generators. 1 * Equal contribution. 1 Code available at https://github.com/rycolab/clusters-in-language-evaluation . 2 We define a language generator as a probability distribution qw over strings w. Specifically, we consider this distribution as used during generation. E.g., if decoding is performed with nucleus sampling, we consider the final distribution where every sentence with tokens not in the nucleus is assigned a probability of 0. 3 Most measures we consider are not metrics in a strict sense; we use the term "metric" out of convention.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- kNN-LM Does Not Improve Open-ended Text GenerationShufan Wang, Yixiao Song, Andrew Drozdov, Aparna Garimella 等EMNLP 2023 · 被引用 3 次
- On the Efficacy of Sampling AdaptersClara Meister, Tiago Pimentel, Luca Malagutti, Ethan Wilcox 等ACL 2023 · 被引用 3 次
- Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMsYanzhu Guo, Simone Conia, Zelin Zhou, Min Li 等ACL 2025
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
相关 Paper
- Information-Theoretic Generative Clustering of DocumentsXin Du, Kumiko Tanaka-IshiiAAAI 2025 · 被引用 1 次
- On the Relation between Quality-Diversity Evaluation and Distribution-Fitting Goal in Text GenerationJianing Li, Yanyan Lan, Jiafeng Guo, Xueqi ChengICML 2020 · 被引用 7 次
- On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar 等ACL 2023 · 被引用 10 次
- A Theoretical Framework for Statistical Evaluability of Generative ModelsShashaank Aiyer, Yishay Mansour, Shay Moran, Han ShaoICML 2026 · 被引用 1 次
- Mutual Information Divergence: A Unified Metric for Multimodal Generative ModelsJin-Hwa Kim, Yunji Kim, Jiyoung Lee, Kang Min Yoo 等NeurIPS 2022 · 被引用 49 次
