ClarifyGPT: A Framework for Enhancing LLM-Based Code Generation via Requirements Clarification
Fangwen Mu, Lin Shi, Song Wang, Zhuohao Yu, Binquan Zhang, Chenxue Wang, Shichao Liu, Qing Wang
摘要
Large Language Models (LLMs), such as ChatGPT, have demonstrated impressive capabilities in automatically generating code from provided natural language requirements. However, in real-world practice, it is inevitable that the requirements written by users might be ambiguous or insufficient. Current LLMs will directly generate programs according to those unclear requirements, regardless of interactive clarification, which will likely deviate from the original user intents. To bridge that gap, we introduce a novel framework named C larify GPT, which aims to enhance code generation by empowering LLMs with the ability to identify ambiguous requirements and ask targeted clarifying questions. Specifically, C larify GPT first detects whether a given requirement is ambiguous by performing a code consistency check. If it is ambiguous, C larify GPT prompts an LLM to generate targeted clarifying questions. After receiving question responses, C larify GPT refines the ambiguous requirement and inputs it into the same LLM to generate a final code solution. To evaluate our C larify GPT, we invite ten participants to use C larify GPT for code generation on two benchmarks: MBPP-sanitized and MBPP-ET. The results show that C larify GPT elevates the performance (Pass@1) of GPT-4 from 70.96% to 80.80% on MBPP-sanitized. Furthermore, to conduct large-scale automated evaluations of C larify GPT across different LLMs and benchmarks without requiring user participation, we introduce a high-fidelity simulation method to simulate user responses. The results demonstrate that C larify GPT can significantly enhance code generation performance compared to the baselines. In particular, C larify GPT improves the average performance of GPT-4 and ChatGPT across five benchmarks from 62.43% to 69.60% and from 54.32% to 62.37%, respectively. A human evaluation also confirms the effectiveness of C larify GPT in detecting ambiguous requirements and generating high-quality clarifying questions. We believe that C larify GPT can effectively facilitate the practical application of LLMs in real-world development environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- What Makes a Good Natural Language Prompt?Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi 等ACL 2025 · 被引用 13 次
- SpecRover: Code Intent Extraction via LLMsHaifeng Ruan, Yuntong Zhang, Abhik RoychoudhuryICSE 2025 · 被引用 12 次
- LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented GenerationDominik Fuchß, Tobias Hey, Jan Keim, Haoyu Liu 等ICSE 2025 · 被引用 8 次
- RustAssure: Differential Symbolic Testing for LLM-Transpiled C-to-Rust CodeYubo Bai, Tapti PalitASE 2025 · 被引用 8 次
- An LLM-Based Agent-Oriented Approach for Automated Code Design Issue LocalizationFraol Batole, David O'Brien, Tien N. Nguyen, Robert Dyer 等ICSE 2025 · 被引用 7 次
它引用的顶会 Paper16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language ModelsTong Zhang, Peixin Qin, Yang Deng, Chen Huang 等ACL 2024
- Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code GenerationHaoxiang Jia, Robbie Morris, He Ye, Federica Sarro 等ASE 2025 · 被引用 6 次
- ConTested: Consistency-Aided Tested Code Generation with LLMJinhao Dong, Jun Sun, Wenjie Zhang, Jin Song Dong 等ISSTA 2025 · 被引用 6 次
- Python Code Generation by Asking Clarification QuestionsHaau-Sing Li, Mohsen Mesgar, André F. T. Martins, Iryna GurevychACL 2023 · 被引用 3 次
- Do Large Language Models Pay Similar Attention Like Human Programmers When Generating Code?Bonan Kou, Shengmai Chen, Zhijie Wang, Lei Ma 等FSE 2024 · 被引用 8 次
