On the Limitations of Reference-Free Evaluations of Generated Text
Daniel Deutsch, Rotem Dror, Dan Roth
Abstract
There is significant interest in developing evaluation metrics which accurately estimate the quality of generated text without the aid of a human-written reference text, which can be time consuming and expensive to collect or entirely unavailable in online applications. However, in this work, we demonstrate that these reference-free metrics are inherently biased and limited in their ability to evaluate generated text, and we argue that they should not be used to measure progress on tasks like machine translation or summarization. We show how reference-free metrics are equivalent to using one generation model to evaluate another, which has several limitations: (1) the metrics can be optimized at test time to find the approximate best-possible output, (2) they are inherently biased toward models which are more similar to their own, and (3) they can be biased against higher-quality outputs, including those written by humans. Therefore, we recommend that reference-free metrics should be used as diagnostic tools for analyzing and understanding model behavior instead of measures of how well models perform a task, in which the goal is to achieve as high of a score as possible. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4dba8c82-7b28-481e-a5ec-8afee540514fCited by top-tier papers19
- CARE: Confounder-Aware Aggregation for Reliable LLM EvaluationJitian Zhao, Changho Shin, Tzu-Heng Huang, Satya Sai Srinath Namburi GNVV et al.ICML 2026 · 7 citations
- Document Summarization with Conformal Importance GuaranteesBruce Kuwahara, Chen-Yuan Lin, Xiao Shi Huang, Kin Kwan Leung et al.NeurIPS 2025 · 7 citations
- Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language ModelsJerry Huang, Prasanna Parthasarathi, Mehdi Rezagholizadeh, Sarath ChandarEMNLP 2024 · 6 citations
- RADE: Reference-Assisted Dialogue Evaluation for Open-Domain DialogueZhengliang Shi, Weiwei Sun, Shuo Zhang, Zhen Zhang et al.ACL 2023 · 5 citations
- What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production VariabilityMario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández et al.EMNLP 2023 · 5 citations
Builds on10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Finding a Balanced Degree of Automation for Summary EvaluationShiyue Zhang, Mohit BansalEMNLP 2021 · 16 citations
Related papers
- Spurious Correlations in Reference-Free Evaluation of Text GenerationEsin Durmus, Faisal Ladhak, Tatsunori HashimotoACL 2022
- CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text GenerationPei Ke, Hao Zhou, Yankai Lin, Peng Li et al.ACL 2022
- MT-Ranker: Reference-free machine translation evaluation by inter-system rankingIbraheem Muhammad Moosa, Rui Zhang, Wenpeng YinICLR 2024 · 13 citations
- On the Limitations of Cross-lingual Encoders as Exposed by Reference-Free Machine Translation EvaluationWei Zhao, Goran Glavas, Maxime Peyrard, Yang Gao et al.ACL 2020 · 53 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
