Are Generative Models Underconfident? Better Quality Estimation with Boosted Model Probability
Tu Anh Dinh, Jan Niehues
Abstract
Quality Estimation (QE) is estimating the quality of the model output during inference when the ground truth is not available. Deriving output quality from the models' output probability is the most trivial and low-effort way. However, we show that the output probability of text-generation models can appear underconfident. At each output step, there can be multiple correct options, making the probability distribution spread out more. Thus, lower probability does not necessarily mean lower output quality. Due to this observation, we propose a QE approach called BOOSTEDPROB 1 , which boosts the model's confidence in cases where there are multiple viable output options. With no increase in complexity, BOOSTEDPROB is notably better than raw model probability in different settings, achieving on average +0.194 improvement in Pearson correlation to groundtruth quality. It also comes close to or outperforms more costly approaches like supervised or ensemble-based QE in certain settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88bc5f6c-76e3-4572-a6bc-85590cbaab41Cited by top-tier papers1
Ask how each one uses itBuilds on13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
Related papers
- Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam SearchXiang Geng, Yu Zhang, Zhejian Lai, Shuaijie She et al.EMNLP 2023 · 2 citations
- Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation OutputRaksha Shenoy, Nico Herbig, Antonio Krüger, Josef van GenabithEMNLP 2021 · 3 citations
- Classification-based Quality Estimation: Small and Efficient Models for Real-world ApplicationsShuo Sun, Ahmed El-Kishky, Vishrav Chaudhary, James Cross et al.EMNLP 2021 · 1 citation
- Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object DetectionXiang Li, Wenhai Wang, Xiaolin Hu, Jun Li et al.CVPR 2021
- DirectQE: Direct Pretraining for Machine Translation Quality EstimationQu Cui, Shujian Huang, Jiahuan Li, Xiang Geng et al.AAAI 2021 · 24 citations
