Code-MUE: Measuring Code LLMs’ Uncertainty through Execution-Based Semantic Interaction Graphs
Xiaoning Ren, Yinxing Xue, Lei Ma, Yuheng Huang
摘要
As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety consequences. Reliable automation, therefore, demands the ability to distinguish between confident, well-supported predictions and stochastic guessing. However, existing uncertainty estimation methods face a critical gap: white-box and grey-box techniques are often inapplicable to closed-source models, while standard black-box text metrics fail to capture the unique fragility of code, where syntactic variation does not necessarily imply semantic divergence. To bridge this syntax-semantics gap, we introduce Code-MUE, a purely black-box framework that measures uncertainty through execution-based Semantic Interaction Graphs. Unlike prior approaches that rely on superficial textual similarity, Code-MUE grounds uncertainty in observable runtime behavior and calculates the von Neumann entropy of the solution space to quantify global semantic diversity. A large-scale empirical study across eight state-of-the-art LLMs demonstrates that Code-MUE achieves a strong negative correlation with functional correctness, with Spearman’s correlation reaching up to −0.98. It significantly outperforms lexical and embedding-based baselines while enabling robust risk detection and selective prediction in practical workflows.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine 等ICLR 2026 · 被引用 218 次
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 被引用 197 次
- Deep Think with ConfidenceYichao Fu, Xuewei Wang, Hao Zhang, Yuandong Tian 等ICLR 2026 · 被引用 171 次
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar 等ICSE 2024 · 被引用 96 次
相关 Paper
- Rethinking Code Complexity Through the Lens of Large Language ModelsChen Xie, Xiaodong Gu, Yuling Shi, Beijun ShenICML 2026
- Fine-grained Uncertainty Decomposition in Large Language Models: A Spectral ApproachNassim Walha, Sebastian G. Gruber, Thomas Decker, Yinchong Yang 等AAAI 2026 · 被引用 2 次
- Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language ModelsHaoyi Song, Ruihan Ji, Naichen Shi, Fan Lai 等NeurIPS 2025 · 被引用 6 次
- LLM-based Vulnerability Discovery through the Lens of Code MetricsFelix Weissberg, Lukas Pirch, Erik Imgrund, Jonas Möller 等ICSE 2026
- Enabling Global, Human-Centered Explanations for LLMs: From Tokens to Interpretable Code and Test GenerationDipin Khati, Daniel Rodriguez-Cardenas, David N. Palacio, Alejandro Velasco 等ICSE 2026
