Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probability
Tu Anh Dinh, Jan Niehues
摘要
Quality Estimation (QE) is estimating the quality of the model output during inference when the ground truth is not available. Deriving output quality from the models' output probability is the most trivial and low-effort way. However, we show that the output probability of text-generation models can appear underconfident. At each output step, there can be multiple correct options, making the probability distribution spread out more. Thus, lower probability does not necessarily mean lower output quality. Due to this observation, we propose a QE approach called BOOSTEDPROB 1 , which boosts the model's confidence in cases where there are multiple viable output options. With no increase in complexity, BOOSTEDPROB is notably better than raw model probability in different settings, achieving on average +0.194 improvement in Pearson correlation to groundtruth quality. It also comes close to or outperforms more costly approaches like supervised or ensemble-based QE in certain settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
相关 Paper
- Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam SearchXiang Geng, Yu Zhang, Zhejian Lai, Shuaijie She 等EMNLP 2023 · 被引用 2 次
- Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation OutputRaksha Shenoy, Nico Herbig, Antonio Krüger, Josef van GenabithEMNLP 2021 · 被引用 3 次
- Classification-based Quality Estimation: Small and Efficient Models for Real-world ApplicationsShuo Sun, Ahmed El-Kishky, Vishrav Chaudhary, James Cross 等EMNLP 2021 · 被引用 1 次
- Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object DetectionXiang Li, Wenhai Wang, Xiaolin Hu, Jun Li 等CVPR 2021
- DirectQE: Direct Pretraining for Machine Translation Quality EstimationQu Cui, Shujian Huang, Jiahuan Li, Xiang Geng 等AAAI 2021 · 被引用 24 次
