Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models
Kyle Cox, Jiawei Xu, Yikun Han, Rong Xu, Tianhao Li, Chi-Yang Hsu, Tianlong Chen, Walter Gerych, Ying Ding
Abstract
An interesting behavior in large language models (LLMs) is prompt sensitivity. When provided with different but semantically equivalent versions of the same prompt, models may produce very different distributions of answers. This suggests that the uncertainty reflected in a model's output distribution for one prompt may not reflect the model's uncertainty about the meaning of the prompt. We model prompt sensitivity as a type of generalization error, and show that sampling across the semantic "concept space" with paraphrasing perturbations improves uncertainty calibration without compromising accuracy. Additionally, we introduce a new metric for uncertainty decomposition in black-box LLMs that improves upon entropy-based decomposition by modeling semantic continuities in natural language generation. We show that this decomposition metric can be used to quantify how much LLM uncertainty is attributed to prompt sensitivity. Our work introduces a new way to improve uncertainty calibration in prompt-sensitive language models, and provides evidence that some LLMs fail to exhibit consistent general reasoning about the meanings of their inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3aa67887-615c-4106-8d0d-cc0ea43d43dfCited by top-tier papers4
- SAFER: Risk-Constrained Sample-then-Filter in Large Language ModelsQingni Wang, Yue Fan, Xin WangICLR 2026 · 8 citations
- Representation Consistency for Accurate and Coherent LLM Answer AggregationJunqi Jiang, Tom Bewley, Salim I. Amoukou, Francesco Leofante et al.NeurIPS 2025 · 6 citations
- Semantic Uncertainty Quantification of Hallucinations in LLMs: A Quantum Tensor Network Based MethodPragatheeswaran Vipulanandan, Kamal Premaratne, Dilip SarkarICLR 2026 · 4 citations
- Diagnosing the Reliability of LLM-as-a-Judge via Item Response TheoryJunhyuk Choi, Sohhyung Park, chanhee cho, Hyeonchu Park et al.ICML 2026
Builds on7
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
- Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingBairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas et al.ICML 2024 · 113 citations
- Premise Order Matters in Reasoning with Large Language ModelsXinyun Chen, Ryan A. Chi, Xuezhi Wang, Denny ZhouICML 2024 · 59 citations
Related papers
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language GenerationLorenz Kuhn, Yarin Gal, Sebastian FarquharICLR 2023 · 49 citations
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 197 citations
- Understanding the Prompt SensitivityYang Liu, Chenhui ChuACL 2026 · 192 citations
- Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language ModelsHaoyi Song, Ruihan Ji, Naichen Shi, Fan Lai et al.NeurIPS 2025 · 6 citations
- Estimating Semantic Alphabet Size for LLM Uncertainty QuantificationLucas H. McCabe, Rimon Melamed, Tom Hartvigsen, H. Howie HuangICLR 2026 · 7 citations
