Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic Space
Xin Qiu, Risto Miikkulainen
Abstract
With the widespread application of Large Language Models (LLMs) to various domains, concerns regarding the trustworthiness of LLMs in safety-critical scenarios have been raised, due to their unpredictable tendency to hallucinate and generate misinformation. Existing LLMs do not have an inherent functionality to provide the users with an uncertainty/confidence metric for each response it generates, making it difficult to evaluate trustworthiness. Although several studies aim to develop uncertainty quantification methods for LLMs, they have fundamental limitations, such as being restricted to classification tasks, requiring additional training and data, considering only lexical instead of semantic information, and being prompt-wise but not response-wise. A new framework is proposed in this paper to address these issues. Semantic density extracts uncertainty/confidence information for each response from a probability distribution perspective in semantic space. It has no restriction on task types and is "off-the-shelf" for new models and tasks. Experiments on seven state-of-the-art LLMs, including the latest Llama 3 and Mixtral-8x22B models, on four free-form question-answering benchmarks demonstrate the superior performance and robustness of semantic density compared to prior approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d54aa1dc-68c5-4544-8de3-822eda537058Cited by top-tier papers23
- CoCoA: A Minimum Bayes Risk Framework Bridging Confidence and Consistency for Uncertainty Quantification in LLMsRoman Vashurin, Maiya Goloburda, Albina Ilina, Aleksandr Rubashevskii et al.NeurIPS 2025 · 33 citations
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement LearningXin Qiu, Yulu Gan, Conor Hayes, Qiyao Liang et al.ICML 2026 · 24 citations
- Hallucination Detection in LLMs with Topological Divergence on Attention GraphsAlexandra Bazarova, Andrei Volodichev, Aleksandr Yugay, Andrey Shulga et al.ACL 2026 · 14 citations
- Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuningHeming Zou, Yixiu Mao, Yun Qu, Qi Wang et al.ICML 2026 · 13 citations
- Thought calibration: Efficient and confident test-time scalingMenghua Wu, Cai Zhou, Stephen Bates, Tommi S. JaakkolaEMNLP 2025 · 12 citations
Builds on9
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Uncertainty Estimation in Autoregressive Structured PredictionAndrey Malinin, Mark J. F. GalesICLR 2021 · 439 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 197 citations
- Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingBairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas et al.ICML 2024 · 113 citations
Related papers
- Semantic Volume: Quantifying and Detecting Both External and Internal Uncertainty in LLMsXiaomin Li, Zhou Yu, Ziji Zhang, Yingying Zhuang et al.AAAI 2026 · 11 citations
- Probabilities Are All You Need: A Probability-Only Approach to Uncertainty Estimation in Large Language ModelsManh Nguyen, Sunil Gupta, Hung LeAAAI 2026 · 4 citations
- Improving Uncertainty Estimation through Semantically Diverse Language GenerationLukas Aichberger, Kajetan Schweighofer, Mykyta Ielanskyi, Sepp HochreiterICLR 2025
- IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model GenerationHaozhi Fan, Jinhao Duan, Kaidi XuACL 2026 · 1 citation
- MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMsYavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao et al.ACL 2024 · 6 citations
