On the Evaluation Metrics for Paraphrase Generation
Lingfeng Shen, Lemao Liu, Haiyun Jiang, Shuming Shi
摘要
In this paper we revisit automatic metrics for paraphrase evaluation and obtain two findings that disobey conventional wisdom: (1) Reference-free metrics achieve better performance than their reference-based counterparts. (2) Most commonly used metrics do not align well with human annotation. Underlying reasons behind the above findings are explored through additional experiments and in-depth analyses. Based on the experiments and analyses, we propose ParaScore, a new evaluation metric for paraphrase generation. It possesses the merits of referencebased and reference-free metrics and explicitly models lexical divergence. Based on our analysis and improvements, our proposed reference-based outperforms than referencefree metrics. Experimental results demonstrate that ParaScore significantly outperforms existing metrics. Our codes and toolkit are released in https://github.com/ shadowkiller33/ParaScore .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Fusing Models with Complementary ExpertiseHongyi Wang, Felipe Maia Polo, Yuekai Sun, Souvik Kundu 等ICLR 2024 · 被引用 44 次
- LAMPAT: Low-Rank Adaption for Multilingual Paraphrasing Using Adversarial TrainingKhoi M. Le, Trinh Pham, Tho Quan, Anh Tuan LuuAAAI 2024 · 被引用 12 次
- The Trickle-down Impact of Reward Inconsistency on RLHFLingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin 等ICLR 2024 · 被引用 10 次
- On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar 等ACL 2023 · 被引用 10 次
- AutoMetrics: Approximate Human Judgments with Automatically Generated EvaluatorsMichael J Ryan, Yanzhe Zhang, Amol Salunkhe, Yi Chu 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper11
- Unsupervised Paraphrasing by Simulated AnnealingXianggen Liu, Lili Mou, Fandong Meng, Hao Zhou 等ACL 2020 · 被引用 74 次
- AESOP: Paraphrase Generation with Adaptive Syntactic ControlJiao Sun, Xuezhe Ma, Nanyun PengEMNLP 2021 · 被引用 47 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- Unsupervised Dual Paraphrasing for Two-stage Semantic ParsingRuisheng Cao, Su Zhu, Chenyu Yang, Chen Liu 等ACL 2020 · 被引用 37 次
- Unsupervised Paraphrasing via Deep Reinforcement LearningA. B. Siddique, Samet Oymak, Vagelis HristidisKDD 2020 · 被引用 27 次
相关 Paper
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Towards Better Characterization of ParaphrasesTimothy Liu, De Wen SohACL 2022 · 被引用 9 次
- Spurious Correlations in Reference-Free Evaluation of Text GenerationEsin Durmus, Faisal Ladhak, Tatsunori HashimotoACL 2022
- RevisEval: Improving LLM-as-a-Judge via Response-Adapted ReferencesQiyuan Zhang, Yufei Wang, Tiezheng Yu, Yuxin Jiang 等ICLR 2025
- LLM-Free Image Captioning Evaluation in Reference-Flexible SettingsShinnosuke Hirano, Yuiga Wada, Kazuki Matsuda, Seitaro Otsuki 等AAAI 2026 · 被引用 2 次
