Exploring Precision and Recall to assess the quality and diversity of LLMs
Florian Le Bronnec, Alexandre Verine, Benjamin Négrevergne, Yann Chevaleyre, Alexandre Allauzen
Abstract
We introduce a novel evaluation framework for Large Language Models (LLMs) such as LLAMA-2 and MISTRAL, focusing on importing Precision and Recall metrics from image generation to text generation. This approach allows for a nuanced assessment of the quality and diversity of generated text without the need for aligned corpora. By conducting a comprehensive evaluation of state-of-the-art language models, the study reveals new insights into their performance on open-ended generation tasks, which are not adequately captured by traditional benchmarks. The findings highlight a trade-off between the quality and diversity of generated samples, particularly when models are fine-tuned on instruction dataset or with human feedback. This work extends the toolkit for distribution-based NLP evaluation, offering insights into the practical capabilities and challenges that current LLMs face in generating diverse and high-quality text. We release our code and data 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25ab44f4-7656-46da-a8d9-ce011292b268Cited by top-tier papers4
- Program Synthesis via Test-Time TransductionKang-il Lee, Jahyun Koo, Seunghyun Yoon, Minbeom Kim et al.NeurIPS 2025 · 4 citations
- Equalized Generative Treatment: Matching f-divergences for Fairness in Generative ModelsAlexandre Verine, Rafael Pinot, Florian Le BronnecICML 2026 · 1 citation
- Measuring Diversity in Synthetic DatasetsYuchang Zhu, Huizhe Zhang, Bingzhe Wu, Jintang Li et al.ICML 2025
- Improving Diversity in Language Models: When Temperature Fails, Change the LossAlexandre Verine, Florian Le Bronnec, Kunhao Zheng, Alexandre Allauzen et al.ICML 2025
Builds on13
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text GenerationZdenek Kasner, Ondrej DusekACL 2024 · 10 citations
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 254 citations
- Open-Domain Text Evaluation via Contrastive Distribution MethodsSidi Lu, Hongyi Liu, Asli Celikyilmaz, Tianlu Wang et al.ICML 2024 · 2 citations
- Themis: A Reference-free NLG Evaluation Language Model with Flexibility and InterpretabilityXinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin et al.EMNLP 2024 · 2 citations
- RepEval: Effective Text Evaluation with LLM RepresentationShuqian Sheng, Yi Xu, Tianhang Zhang, Zanwei Shen et al.EMNLP 2024 · 5 citations
