Fixing Large Language Models' Specification Misunderstanding for Better Code Generation
Zhao Tian, Junjie Chen, Xiangyu Zhang
Abstract
Code generation is to automatically generate source code conforming to a given programming specification, which has received extensive attention especially with the development of large language models (LLMs). Due to the inherent difficulty of code generation, the code generated by LLMs may not be aligned with the specification. Although thought-eliciting prompting techniques have been proposed to enhance the code generation performance of LLMs, producing correct understanding for complicated programming problems remains challenging, resulting in unsatisfactory performance. Also, some feedbackbased prompting techniques have been proposed to fix incorrect code using error messages produced by test execution. However, when the generated code deviates significantly from the ground truth, they encounter difficulties in improving performance based on such coarse-grained information. In this work, we propose a novel prompting technique, called <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex>, to improve the code generation performance of LLMs by devising both sophisticated thought-eliciting prompting and feedback-based prompting and making the first exploration on their synergy. It first exploits test case analysis to obtain specification understanding and enables a self-improvement process to identify and refine the misunderstanding in the thoughteliciting prompting phase. <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> further fixes the specification understanding towards the direction reducing the gap between the provided understanding (from the first phase) and the actual understanding implicity utilized by LLMs for code generation in the feedback-based prompting phase. By improving the understanding with <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex>, the code generation performance of LLMs can be largely improved. Our evaluation on two advanced LLMs (ChatGPT and DeepSeek-Coder) with six widely-used benchmarks by comparing with 15 baselines, demonstrates the effectiveness of <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex>. For example, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> outperforms the most effective baseline with an average improvement of 35.62 % in terms of Pass@1 across all subjects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb0b48cc-8b7d-467b-ac64-fa227e925e47Cited by top-tier papers5
- Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code GenerationHaoxiang Jia, Robbie Morris, He Ye, Federica Sarro et al.ASE 2025 · 6 citations
- RefAgent: A Multi-agent LLM-based Framework for Automatic Software RefactoringKhouloud Oueslati, Maxime Lamothe, Foutse KhomhICSE 2026 · 1 citation
- AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code GenerationKaifeng He, Mingwei Liu, Chong Wang, Zike Li et al.FSE 2026
- Aligning Requirement for Large Language Model's Code GenerationZhao Tian, Junjie ChenICSE 2026
- Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case StudyMingwei Liu, Zheng Pei, Yanlin Wang, Zihao Wang et al.FSE 2026
Builds on18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- MoT: Modularization-of-Thought Prompting for Effective Code GenerationRuwei Pan, Hongyu ZhangISSTA 2026
- Enhanced Prompting Framework for Code Summarization with Large Language ModelsMinying Fang, Xing Yuan, Yuying Li, Haojie Li et al.ISSTA 2025 · 3 citations
- Intention is All you Need: Refining your Code from your IntentionQi Guo, Xiaofei Xie, Shangqing Liu, Ming Hu et al.ICSE 2025 · 7 citations
- Live in the Loop: Rapid Run-time Feedback for PromptsToni Mattis, Abdullatif Ghajar, Tom Beckmann, Robert HirschfeldCHI 2026 · 2 citations
- CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksHongchao Jiang, Yiming Chen, Yushi Cao, Hung-Yi Lee et al.ACL 2026 · 33 citations
