Perception Score: A Learned Metric for Open-ended Text Generation Evaluation
Jing Gu, Qingyang Wu, Zhou Yu
2021Year
3Citations
4Top-tier citations
Abstract
Automatic evaluation for open-ended natural language generation tasks remains a challenge. We propose a learned evaluation metric: Perception Score. It utilizes a pre-trained model and considers context information for conditional generation. Perception Score assigns a holistic score along with the uncertainty measurement. We conduct experiments on three open-ended conditional generation tasks and two open-ended unconditional generation tasks. Perception Score achieves state-of-the-art results on all the tasks consistently in terms of correlation with human evaluation scores.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Think Together and Work Better: Combining Humans' and LLMs' Think-Aloud Outcomes for Effective Text EvaluationSeongYeub Chu, Jong Woo Kim, Mun Yong YiCHI 2025 · 7 citations
- CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklistsYukyung Lee, JoongHoon Kim, Jaehee Kim, Hyowon Cho et al.EMNLP 2025 · 2 citations
- SMURF: SeMantic and linguistic UndeRstanding Fusion for Caption Evaluation via Typicality AnalysisJoshua Feinglass, Yezhou YangACL 2021
- Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine TextYao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith et al.ACL 2022
Builds on3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Learning to Compare for Better Training and Evaluation of Open Domain Natural Language Generation ModelsWangchunshu Zhou, Ke XuAAAI 2020 · 49 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text GenerationPei Ke, Hao Zhou, Yankai Lin, Peng Li et al.ACL 2022
- OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation MetricsJian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu et al.ACL 2021
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou et al.ACL 2020 · 69 citations
- Conformal Reliability: A New Evaluation Metric for Conditional GenerationYachen Gao, Xinwei Sun, Yikai Wang, Ye Shi et al.ICML 2026
- QRelScore: Better Evaluating Generated Questions with Deeper Understanding of Context-aware RelevanceXiaoqiang Wang, Bang Liu, Siliang Tang, Lingfei WuEMNLP 2022 · 6 citations
