MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Salman Avestimehr
Abstract
Generative Large Language Models (LLMs) are widely utilized for their excellence in various tasks. However, their tendency to produce inaccurate or misleading outputs poses a potential risk, particularly in high-stakes environments. Therefore, estimating the correctness of generative LLM outputs is an important task for enhanced reliability. Uncertainty Estimation (UE) in generative LLMs is an evolving domain, where SOTA probability-based methods commonly employ length-normalized scoring. In this work, we propose Meaning-Aware Response Scoring (MARS) as an alternative to length-normalized scoring for UE methods. MARS is a novel scoring function that considers the semantic contribution of each token in the generated sequence in the context of the question. We demonstrate that integrating MARS into UE methods results in a universal and significant improvement in UE performance. We conduct experiments using three distinct closed-book question-answering datasets across five popular pre-trained LLMs. Lastly, we validate the efficacy of MARS on a Medical QA dataset. Code can be found here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Robust Hallucination Detection in LLMs via Adaptive Token SelectionMengjia Niu, Hamed Haddadi, Guansong PangNeurIPS 2025 · 24 citations
- Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence MeasureLukas Aichberger, Kajetan Schweighofer, Sepp HochreiterICLR 2026 · 12 citations
- Addressing Pitfalls in the Evaluation of Uncertainty Estimation Methods for Natural Language GenerationMykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger, Sepp HochreiterICLR 2026 · 10 citations
- Data Imputation with Limited Data Redundancy Using Data LakesChenyu Yang, Yuyu Luo, Chuanxuan Cui, Ju Fan et al.VLDB 2025 · 9 citations
- Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-AnsweringYavuz Faruk Bakman, Sungmin Kang, Zhiqi Huang, Duygu Nur Yaldiz et al.ICLR 2026 · 6 citations
Builds on10
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- Adapting Large Language Models via Reading ComprehensionDaixuan Cheng, Shaohan Huang, Furu WeiICLR 2024 · 146 citations
- Uncertainty Estimation of Transformer Predictions for Misclassification DetectionArtem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun et al.ACL 2022 · 59 citations
- WebQA: Multihop and Multimodal QAYingshan Chang, Guihong Cao, Mridu Narang, Jianfeng Gao et al.CVPR 2022 · 58 citations
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language GenerationLorenz Kuhn, Yarin Gal, Sebastian FarquharICLR 2023 · 49 citations
Related papers
- UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language ModelsRoman Vashurin, Maiya Goloburda, Preslav Nakov, Maxim PanovEMNLP 2025 · 1 citation
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language ModelsJinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny et al.ACL 2024 · 28 citations
- Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language GenerationZhen Lin, Shubhendu Trivedi, Jimeng SunEMNLP 2024 · 1 citation
- Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic SpaceXin Qiu, Risto MiikkulainenNeurIPS 2024
- TokUR: Token-Level Uncertainty Estimation for Large Language Model ReasoningTunyu Zhang, Haizhou Shi, Yibin Wang, Hengyi Wang et al.ICLR 2026 · 19 citations
