INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, Lei Li
Abstract
Automatically evaluating the quality of language generation is critical. Although recent learned metrics show high correlation with human judgement, these metrics do not provide explicit explanation of their verdict, nor associate the scores with defects in the generated text. To address this limitation, we present IN-STRUCTSCORE, a fine-grained explainable evaluation metric for text generation. By harnessing both explicit human instruction and the implicit knowledge of GPT-4, we fine-tune a text evaluation metric based on LLaMA, producing both a score for generated text and a human readable diagnostic report. We evaluate INSTRUCTSCORE on a variety of generation tasks, including translation, captioning, data-to-text, and commonsense generation. Experiments show that our 7B model surpasses all other unsupervised metrics, including those based on 175B GPT-3 and GPT-4. Surprisingly, our INSTRUCTSCORE, even without direct supervision from human-rated data, achieves performance levels on par with state-of-the-art metrics like COMET22, which were fine-tuned on human ratings. Prompt: You are evaluating a model output based on a reference. Reference: Normally the administration office downstairs would call me when there's a delivery. Output: Usually when there is takeaway, the management office downstairs will call.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10bae14a-bcf3-440a-8e82-0652e68bdd40Cited by top-tier papers31
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context LearningBill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri et al.ICLR 2024 · 299 citations
- A Closer Look at the Limitations of Instruction TuningSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S. et al.ICML 2024 · 90 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software EngineeringRuiqi Wang, Jiyu Guo, Cuiyun Gao, Guodong Fan et al.ISSTA 2025 · 27 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
Related papers
- VIEScore: Towards Explainable Metrics for Conditional Image Synthesis EvaluationMax Ku, Dongfu Jiang, Cong Wei, Xiang Yue et al.ACL 2024 · 24 citations
- From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test SetMara Finkelstein, Daniel Deutsch, Parker Riley, Juraj Juraska et al.ICML 2025
- CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model GenerationPei Ke, Bosi Wen, Andrew Feng, Xiao Liu et al.ACL 2024 · 9 citations
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildBill Yuchen Lin, Yuntian Deng, Khyathi Raghavi Chandu, Abhilasha Ravichander et al.ICLR 2025
- DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question AnsweringPei Ke, Fei Huang, Fei Mi, Yasheng Wang et al.ACL 2023 · 2 citations
