Identifying Reliable Evaluation Metrics for Scientific Text Revision
Léane Jourdan, Nicolas Hernandez, Florian Boudin, Richard Dufour
Abstract
Evaluating text revision in scientific writing remains a challenge, as traditional metrics such as ROUGE and BERTScore primarily focus on similarity rather than capturing meaningful improvements. In this work, we analyse and identify the limitations of these metrics and explore alternative evaluation methods that better align with human judgments. We first conduct a manual annotation study to assess the quality of different revisions. Then, we investigate reference-free evaluation metrics from related NLP domains. Additionally, we examine LLMas-a-judge approaches, analysing their ability to assess revisions with and without a gold reference. Our results show that LLMs effectively assess instruction-following but struggle with correctness, while domain-specific metrics provide complementary insights. We find that a hybrid approach combining LLM-as-a-judge evaluation and task-specific metrics offers the most reliable assessment of revision quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88ce9d89-1b3f-4ce2-bb5f-51ecf922da3dCited by top-tier papers1
Ask how each one uses itBuilds on4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Text Revision By On-the-Fly Representation OptimizationJingjing Li, Zichao Li, Tao Ge, Irwin King et al.AAAI 2022 · 20 citations
- Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie CalibrationDaniel Deutsch, George F. Foster, Markus FreitagEMNLP 2023 · 14 citations
- Improving Iterative Text Revision by Learning Where to Edit from Other Revision TasksZae Myung Kim, Wanyu Du, Vipul Raheja, Dhruv Kumar et al.EMNLP 2022 · 8 citations
Related papers
- RevisEval: Improving LLM-as-a-Judge via Response-Adapted ReferencesQiyuan Zhang, Yufei Wang, Tiezheng Yu, Yuxin Jiang et al.ICLR 2025
- Language Model Augmented Relevance ScoreRuibo Liu, Jason Wei, Soroush VosoughiACL 2021
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMsDenis Janiak, Jakub Binkowski, Albert Sawczyn, Bogdan Gabrys et al.EMNLP 2025
- A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better InterpretabilityXinyu Hu, Mingqi Gao, Li Lin, Zhenghan Yu et al.ACL 2025
- FineSurE: Fine-grained Summarization Evaluation using LLMsHwanjun Song, Hang Su, Igor Shalyminov, Jason Cai et al.ACL 2024
