On the Evaluation Metrics for Paraphrase Generation
Lingfeng Shen, Lemao Liu, Haiyun Jiang, Shuming Shi
Abstract
In this paper we revisit automatic metrics for paraphrase evaluation and obtain two findings that disobey conventional wisdom: (1) Reference-free metrics achieve better performance than their reference-based counterparts. (2) Most commonly used metrics do not align well with human annotation. Underlying reasons behind the above findings are explored through additional experiments and in-depth analyses. Based on the experiments and analyses, we propose ParaScore, a new evaluation metric for paraphrase generation. It possesses the merits of referencebased and reference-free metrics and explicitly models lexical divergence. Based on our analysis and improvements, our proposed reference-based outperforms than referencefree metrics. Experimental results demonstrate that ParaScore significantly outperforms existing metrics. Our codes and toolkit are released in https://github.com/ shadowkiller33/ParaScore .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b110c188-b6d4-4568-b34e-ce9591120c67Cited by top-tier papers8
- Fusing Models with Complementary ExpertiseHongyi Wang, Felipe Maia Polo, Yuekai Sun, Souvik Kundu et al.ICLR 2024 · 44 citations
- LAMPAT: Low-Rank Adaption for Multilingual Paraphrasing Using Adversarial TrainingKhoi M. Le, Trinh Pham, Tho Quan, Anh Tuan LuuAAAI 2024 · 12 citations
- The Trickle-down Impact of Reward Inconsistency on RLHFLingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin et al.ICLR 2024 · 10 citations
- On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar et al.ACL 2023 · 10 citations
- AutoMetrics: Approximate Human Judgments with Automatically Generated EvaluatorsMichael J Ryan, Yanzhe Zhang, Amol Salunkhe, Yi Chu et al.ICLR 2026 · 3 citations
Builds on11
- Unsupervised Paraphrasing by Simulated AnnealingXianggen Liu, Lili Mou, Fandong Meng, Hao Zhou et al.ACL 2020 · 74 citations
- AESOP: Paraphrase Generation with Adaptive Syntactic ControlJiao Sun, Xuezhe Ma, Nanyun PengEMNLP 2021 · 47 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Unsupervised Dual Paraphrasing for Two-stage Semantic ParsingRuisheng Cao, Su Zhu, Chenyu Yang, Chen Liu et al.ACL 2020 · 37 citations
- Unsupervised Paraphrasing via Deep Reinforcement LearningA. B. Siddique, Samet Oymak, Vagelis HristidisKDD 2020 · 27 citations
Related papers
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Towards Better Characterization of ParaphrasesTimothy Liu, De Wen SohACL 2022 · 9 citations
- Spurious Correlations in Reference-Free Evaluation of Text GenerationEsin Durmus, Faisal Ladhak, Tatsunori HashimotoACL 2022
- RevisEval: Improving LLM-as-a-Judge via Response-Adapted ReferencesQiyuan Zhang, Yufei Wang, Tiezheng Yu, Yuxin Jiang et al.ICLR 2025
- LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsShinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki et al.AAAI 2026 · 2 citations
