What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production Variability
Mario Giulianelli, Joris Baan, Wilker Aziz, Raquel Fernández, Barbara Plank
Abstract
In Natural Language Generation (NLG) tasks, for any input, multiple communicative goals are plausible, and any goal can be put into words, or produced, in multiple ways. We characterise the extent to which human production varies lexically, syntactically, and semantically across four NLG tasks, connecting human production variability to aleatoric or data uncertainty. We then inspect the space of output strings shaped by a generation system's predicted probability distribution and decoding algorithm to probe its uncertainty. For each test input, we measure the generator's calibration to human production variability. Following this instance-level approach, we analyse NLG models and decoding strategies, demonstrating that probing a generator with multiple samples and, when possible, multiple references, provides the level of detail necessary to gain understanding of a model's representation of uncertainty. 1 * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Enhancing Uncertainty Modeling with Semantic Graph for Hallucination DetectionKedi Chen, Qin Chen, Jie Zhou, Xinqi Tao et al.AAAI 2025 · 13 citations
- How Far Can We Extract Diverse Perspectives from Large Language Models?Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, Dongyeop KangEMNLP 2024 · 6 citations
- Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesMario Giulianelli, Sarenne Wallbridge, Raquel FernándezEMNLP 2023 · 4 citations
- FastLog: An End-to-End Method to Efficiently Generate and Insert Logging StatementsXiaoyuan Xie, Zhipeng Cai, Songqiang Chen, Jifeng XuanISSTA 2024 · 3 citations
- Lexical Diversity-aware Relevance Assessment for Retrieval-Augmented GenerationZhange Zhang, Yuqing Ma, Yulong Wang, Shan He et al.ACL 2025 · 2 citations
Builds on11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun et al.NeurIPS 2021 · 606 citations
- Towards a Unified Multi-Dimensional Evaluator for Text GenerationMing Zhong, Yang Liu, Da Yin, Yuning Mao et al.EMNLP 2022 · 103 citations
- Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language GenerationMingkai Deng, Bowen Tan, Zhengzhong Liu, Eric P. Xing et al.EMNLP 2021 · 49 citations
Related papers
- Improving Uncertainty Estimation through Semantically Diverse Language GenerationLukas Aichberger, Kajetan Schweighofer, Mykyta Ielanskyi, Sepp HochreiterICLR 2025
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language ModelsKyle Cox, Jiawei Xu, Yikun Han, Rong Xu et al.AAAI 2025 · 6 citations
- Distinguishing the Knowable from the Unknowable with Language ModelsGustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak et al.ICML 2024 · 44 citations
- Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language GenerationMykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp HochreiterICLR 2026 · 10 citations
- Uncertainty in Language Models: Assessment through Rank-CalibrationXinmeng Huang, Shuo Li, Mengxin Yu, Matteo Sesia et al.EMNLP 2024 · 9 citations
