Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code Generation
Haoxiang Jia, Robbie Morris, He Ye, Federica Sarro, Sergey Mechtaev
Abstract
The growing use of large language models (LLMs) has increased the importance of natural language (NL) in software engineering. However, ambiguity of NL can harm software quality, as unclear problem descriptions may lead to incorrect program generation. Detecting and resolving such ambiguity is challenging, motivating our introduction of the automated repair of ambiguous NL descriptions, which we approach by reducing code generation uncertainty and better aligning NL with input–output examples. Ambiguity repair is difficult for LLMs because they must understand how their interpretation of a description changes when the text is altered. We find that directly prompting LLMs to clarify ambiguity often produces irrelevant or inconsistent edits. To address this, we decompose this task into two simpler steps: (1) analyzing and repairing the LLM’s interpretation of the description — captured by the distribution of programs it induces — using traditional testing and program repair, and (2) refining the description based on distribution changes via a method we call contrastive specification inference. We implement this approach in a tool called SPEC-FIX and evaluate it using four state-of-the-art LLMs (GPT-4o, GPT-4o-mini, DeepSeek-V3, and Qwen2.5-Coder-32B-Instruct) on three popular code generation benchmarks (HumanEval+, MBPP+ and LiveCodeBench). Without human intervention or external information, SPECFIX modified 43.58% of descriptions, improving Pass@1 on the modified set by 30.9%. This yields a 4.09% absolute improvement across the entire benchmark. Repairs also transfer across models: descriptions repaired for one model improve other models’ performance by 10.48%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86aa94c8-c8dc-4c74-b508-c00bbf3df5d4Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury et al.ICSE 2023 · 213 citations
- NEZHA: Efficient Domain-Independent Differential TestingTheofilos Petsios, Adrian Tang, Salvatore J. Stolfo, Angelos D. Keromytis et al.S&P 2017 · 132 citations
- Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingBairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas et al.ICML 2024 · 113 citations
- CodeT: Code Generation with Generated TestsBei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan et al.ICLR 2023 · 64 citations
Related papers
- Debugging Engine Enhanced by Prior Knowledge: Can We Teach LLM How to Debug?Kunyi Li, Sai Wu, Xiu Tang, Chang Yao et al.FSE 2026 · 1 citation
- Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Bonan Kou, Shengmai Chen, Zhijie Wang, Lei Ma et al.FSE 2024 · 8 citations
- Fixing Large Language Models' Specification Misunderstanding for Better Code GenerationZhao Tian, Junjie Chen, Xiangyu ZhangICSE 2025 · 6 citations
- ConTested: Consistency-Aided Tested Code Generation with LLMJinhao Dong, Jun Sun, Wenjie Zhang, Jin Song Dong et al.ISSTA 2025 · 6 citations
- Aligning the Objective of LLM-Based Program RepairJunjielong Xu, Ying Fu, Shin Hwei Tan, Pinjia HeICSE 2025 · 5 citations
