R-U-SURE? Uncertainty-Aware Code Suggestions By Maximizing Utility Across Random User Intents
Daniel D. Johnson, Daniel Tarlow, Christian Walder
摘要
Large language models show impressive results at predicting structured text such as code, but also commonly introduce errors and hallucinations in their output. When used to assist software developers, these models may make mistakes that users must go back and fix, or worse, introduce subtle bugs that users may miss entirely. We propose Randomized Utility-driven Synthesis of Uncertain REgions (R-U-SURE), an approach for building uncertainty-aware suggestions based on a decision-theoretic model of goal-conditioned utility, using random samples from a generative model as a proxy for the unobserved possible intents of the end user. Our technique combines minimum-Bayes-risk decoding, dual decomposition, and decision diagrams in order to efficiently produce structured uncertainty summaries, given only sample access to an arbitrary generative model of code and an optional AST parser. We demonstrate R-U-SURE on three developerassistance tasks, and show that it can be applied different user interaction patterns without retraining the model and leads to more accurate uncertainty estimates than token-probability baselines. We also release our implementation as an open-source library at https://github. com/google-research/r_u_sure .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Calibration and Correctness of Language Models for CodeClaudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel 等ICSE 2025 · 被引用 21 次
- Guess & Sketch: Language Model Guided TranspilationCeline Lee, Abdulrahman Mahmoud, Michal Kurek, Simone Campanoni 等ICLR 2024 · 被引用 8 次
- Calibration of Large Language Models on Code SummarizationYuvraj Virk, Premkumar T. Devanbu, Toufique AhmedFSE 2025 · 被引用 5 次
- VERSE: Verification-based Self-Play for Code InstructionsHao Jiang, Qi Liu, Rui Li, Yuze Zhao 等AAAI 2025 · 被引用 3 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 被引用 408 次
相关 Paper
- CoCoA: A Minimum Bayes Risk Framework Bridging Confidence and Consistency for Uncertainty Quantification in LLMsRoman Vashurin, Maiya Goloburda, Albina Ilina, Aleksandr Rubashevskii 等NeurIPS 2025 · 被引用 33 次
- Task-Awareness Improves LLM Generations and UncertaintyTim Tomov, Dominik Fuchsgruber, Stephan GünnemannICML 2026 · 被引用 2 次
- DeLLMa: Decision Making Under Uncertainty with Large Language ModelsOllie Liu, Deqing Fu, Dani Yogatama, Willie NeiswangerICLR 2025
- Uncertainty-Aware Decoding with Minimum Bayes RiskNico Daheim, Clara Meister, Thomas Möllenhoff, Iryna GurevychICLR 2025
- Improving Uncertainty Estimation through Semantically Diverse Language GenerationLukas Aichberger, Kajetan Schweighofer, Mykyta Ielanskyi, Sepp HochreiterICLR 2025
