Sigmoid Head for Quality Estimation under Language Ambiguity
Tu Anh Dinh, Jan Niehues
Abstract
Language model (LM) probability is not a reliable quality estimator, as natural language is ambiguous. When multiple output options are valid, the model's probability distribution is spread across them, which can misleadingly indicate low output quality. This issue is caused by two reasons: (1) LMs' final output activation is softmax, which does not allow multiple correct options to receive high probabilities simultaneuously and (2) LMs' training data is single, one-hot encoded references, indicating that there is only one correct option at each output step. We propose training a module for Quality Estimation on top of pre-trained LMs to address these limitations. The module, called Sigmoid Head, is an extra unembedding head with sigmoid activation to tackle the first limitation. To tackle the second limitation, during the negative sampling process to train the Sigmoid Head, we use a heuristic to avoid selecting potentially alternative correct tokens. Our Sigmoid Head is computationally efficient during training and inference. The probability from Sigmoid Head is notably better quality signal compared to the original softmax head. As the Sigmoid Head does not rely on humanannotated quality data, it is more robust to outof-domain settings compared to supervised QE. 1 Implementation available at https://github.com/ TuAnh23/sigmoid-head-qe .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2365e90-00de-4925-a480-756f5dff8eedBuilds on6
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield et al.ACL 2020 · 132 citations
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language GenerationLorenz Kuhn, Yarin Gal, Sebastian FarquharICLR 2023 · 49 citations
- Out-of-Distribution Detection and Selective Generation for Conditional Language ModelsJie Ren, Jiaming Luo, Yao Zhao, Kundan Krishna et al.ICLR 2023 · 12 citations
Related papers
- Are Generative Models Underconfident? Better Quality Estimation with Boosted Model ProbabilityTu Anh Dinh, Jan NiehuesEMNLP 2025
- A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM OutputsArtem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun et al.EMNLP 2025 · 3 citations
- Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score DistributionZhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue et al.CVPR 2025
- Calibrating Sequence likelihood Improves Conditional Language GenerationYao Zhao, Misha Khalman, Rishabh Joshi, Shashi Narayan et al.ICLR 2023 · 38 citations
- Revisiting MLLM Based Image Quality Assessment: Errors and RemedyZhenchen Tang, Songlin Yang, Bo Peng, Zichuan Wang et al.AAAI 2026 · 2 citations
