Exploring Precision and Recall to assess the quality and diversity of LLMs
Florian Le Bronnec, Alexandre Verine, Benjamin Négrevergne, Yann Chevaleyre, Alexandre Allauzen
摘要
We introduce a novel evaluation framework for Large Language Models (LLMs) such as LLAMA-2 and MISTRAL, focusing on importing Precision and Recall metrics from image generation to text generation. This approach allows for a nuanced assessment of the quality and diversity of generated text without the need for aligned corpora. By conducting a comprehensive evaluation of state-of-the-art language models, the study reveals new insights into their performance on open-ended generation tasks, which are not adequately captured by traditional benchmarks. The findings highlight a trade-off between the quality and diversity of generated samples, particularly when models are fine-tuned on instruction dataset or with human feedback. This work extends the toolkit for distribution-based NLP evaluation, offering insights into the practical capabilities and challenges that current LLMs face in generating diverse and high-quality text. We release our code and data 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Program Synthesis via Test-Time TransductionKang-il Lee, Jahyun Koo, Seunghyun Yoon, Minbeom Kim 等NeurIPS 2025 · 被引用 4 次
- Equalized Generative Treatment: Matching f-divergences for Fairness in Generative ModelsAlexandre Verine, Rafael Pinot, Florian Le BronnecICML 2026 · 被引用 1 次
- Measuring Diversity in Synthetic DatasetsYuchang Zhu, Huizhe Zhang, Bingzhe Wu, Jintang Li 等ICML 2025
- Improving Diversity in Language Models: When Temperature Fails, Change the LossAlexandre Verine, Florian Le Bronnec, Kunhao Zheng, Alexandre Allauzen 等ICML 2025
它引用的顶会 Paper13
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text GenerationZdenek Kasner, Ondrej DusekACL 2024 · 被引用 10 次
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 被引用 254 次
- Open-Domain Text Evaluation via Contrastive Distribution MethodsSidi Lu, Hongyi Liu, Asli Celikyilmaz, Tianlu Wang 等ICML 2024 · 被引用 2 次
- Themis: A Reference-free NLG Evaluation Language Model with Flexibility and InterpretabilityXinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin 等EMNLP 2024 · 被引用 2 次
- RepEval: Effective Text Evaluation with LLM RepresentationShuqian Sheng, Yi Xu, Tianhang Zhang, Zanwei Shen 等EMNLP 2024 · 被引用 5 次
