Code-MUE: Measuring Code LLMs’ Uncertainty through Execution-Based Semantic Interaction Graphs
Xiaoning Ren, Yinxing Xue, Lei Ma, Yuheng Huang
Abstract
As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety consequences. Reliable automation, therefore, demands the ability to distinguish between confident, well-supported predictions and stochastic guessing. However, existing uncertainty estimation methods face a critical gap: white-box and grey-box techniques are often inapplicable to closed-source models, while standard black-box text metrics fail to capture the unique fragility of code, where syntactic variation does not necessarily imply semantic divergence. To bridge this syntax-semantics gap, we introduce Code-MUE, a purely black-box framework that measures uncertainty through execution-based Semantic Interaction Graphs. Unlike prior approaches that rely on superficial textual similarity, Code-MUE grounds uncertainty in observable runtime behavior and calculates the von Neumann entropy of the solution space to quantify global semantic diversity. A large-scale empirical study across eight state-of-the-art LLMs demonstrates that Code-MUE achieves a strong negative correlation with functional correctness, with Spearman’s correlation reaching up to −0.98. It significantly outperforms lexical and embedding-based baselines while enabling robust risk detection and selective prediction in practical workflows.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 896923ac-d74f-4c84-9295-1a8ca4c5d5bcBuilds on27
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine et al.ICLR 2026 · 218 citations
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 197 citations
- Deep Think with ConfidenceYichao Fu, Xuewei Wang, Hao Zhang, Yuandong Tian et al.ICLR 2026 · 171 citations
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar et al.ICSE 2024 · 96 citations
Related papers
- Rethinking Code Complexity Through the Lens of Large Language ModelsChen Xie, Xiaodong Gu, Yuling Shi, Beijun ShenICML 2026
- Fine-grained Uncertainty Decomposition in Large Language Models: A Spectral ApproachNassim Walha, Sebastian G. Gruber, Thomas Decker, Yinchong Yang et al.AAAI 2026 · 2 citations
- Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language ModelsHaoyi Song, Ruihan Ji, Naichen Shi, Fan Lai et al.NeurIPS 2025 · 6 citations
- LLM-based Vulnerability Discovery through the Lens of Code MetricsFelix Weissberg, Lukas Pirch, Erik Imgrund, Jonas Möller et al.ICSE 2026
- Enabling Global, Human-Centered Explanations for LLMs: From Tokens to Interpretable Code and Test GenerationDipin Khati, Daniel Rodriguez-Cardenas, David N. Palacio, Alejandro Velasco et al.ICSE 2026
