Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code Generation
Haoxiang Jia, Robbie Morris, He Ye, Federica Sarro, Sergey Mechtaev
摘要
The growing use of large language models (LLMs) has increased the importance of natural language (NL) in software engineering. However, ambiguity of NL can harm software quality, as unclear problem descriptions may lead to incorrect program generation. Detecting and resolving such ambiguity is challenging, motivating our introduction of the automated repair of ambiguous NL descriptions, which we approach by reducing code generation uncertainty and better aligning NL with input–output examples. Ambiguity repair is difficult for LLMs because they must understand how their interpretation of a description changes when the text is altered. We find that directly prompting LLMs to clarify ambiguity often produces irrelevant or inconsistent edits. To address this, we decompose this task into two simpler steps: (1) analyzing and repairing the LLM’s interpretation of the description — captured by the distribution of programs it induces — using traditional testing and program repair, and (2) refining the description based on distribution changes via a method we call contrastive specification inference. We implement this approach in a tool called SPEC-FIX and evaluate it using four state-of-the-art LLMs (GPT-4o, GPT-4o-mini, DeepSeek-V3, and Qwen2.5-Coder-32B-Instruct) on three popular code generation benchmarks (HumanEval+, MBPP+ and LiveCodeBench). Without human intervention or external information, SPECFIX modified 43.58% of descriptions, improving Pass@1 on the modified set by 30.9%. This yields a 4.09% absolute improvement across the entire benchmark. Repairs also transfer across models: descriptions repaired for one model improve other models’ performance by 10.48%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury 等ICSE 2023 · 被引用 213 次
- NEZHA: Efficient Domain-Independent Differential TestingTheofilos Petsios, Adrian Tang, Salvatore J. Stolfo, Angelos D. Keromytis 等S&P 2017 · 被引用 132 次
- Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingBairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas 等ICML 2024 · 被引用 113 次
- CodeT: Code Generation with Generated TestsBei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan 等ICLR 2023 · 被引用 64 次
相关 Paper
- Debugging Engine Enhanced by Prior Knowledge: Can We Teach LLM How to Debug?Kunyi Li, Sai Wu, Xiu Tang, Chang Yao 等FSE 2026 · 被引用 1 次
- Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Bonan Kou, Shengmai Chen, Zhijie Wang, Lei Ma 等FSE 2024 · 被引用 8 次
- Fixing Large Language Models' Specification Misunderstanding for Better Code GenerationZhao Tian, Junjie Chen, Xiangyu ZhangICSE 2025 · 被引用 6 次
- ConTested: Consistency-Aided Tested Code Generation with LLMJinhao Dong, Jun Sun, Wenjie Zhang, Jin Song Dong 等ISSTA 2025 · 被引用 6 次
- Aligning the Objective of LLM-Based Program RepairJunjielong Xu, Ying Fu, Shin Hwei Tan, Pinjia HeICSE 2025 · 被引用 5 次
