Python Code Generation by Asking Clarification Questions
Haau-Sing Li, Mohsen Mesgar, André F. T. Martins, Iryna Gurevych
摘要
Code generation from text requires understanding the user's intent from a natural language description and generating an executable code snippet that satisfies this intent. While recent pretrained language models demonstrate remarkable performance for this task, these models fail when the given natural language description is under-specified. In this work, we introduce a novel and more realistic setup for this task. We hypothesize that the underspecification of a natural language description can be resolved by asking clarification questions. Therefore, we collect and introduce a new dataset named CodeClarQA containing pairs of natural language descriptions and code with created synthetic clarification questions and answers. The empirical results of our evaluation of pretrained language model performance on code generation show that clarifications result in more precisely generated code, as shown by the substantial improvement of model performance in all evaluation metrics. Alongside this, our task and dataset introduce new challenges to the community, including when and what clarification questions should be asked. Our code and dataset are available on GitHub. 1 * Work done while being a postdoc at UKP Lab.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ClarifyGPT: A Framework for Enhancing LLM-Based Code Generation via Requirements ClarificationFangwen Mu, Lin Shi, Song Wang, Zhuohao Yu 等FSE 2024 · 被引用 49 次
- Topological Active Inference for Task DisambiguationYangbo Wei, Zhen Huang, Shaoqiang Lu, Junhong Qian 等ICML 2026
- Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMsClara Lachenmaier, Hannah Bultmann, Sina ZarrießACL 2026
- Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code GenerationMert Inan, Anthony Sicilia, Alex Xie, Saujas Vaduguru 等EMNLP 2025
它引用的顶会 Paper6
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
- Building and Evaluating Open-Domain Dialogue Corpora with Clarifying QuestionsMohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton 等EMNLP 2021 · 被引用 61 次
- Natural Language to Code Generation in Interactive Data Science NotebooksPengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao 等ACL 2023 · 被引用 17 次
相关 Paper
- Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD GenerationBo Yuan, Zelin Zhao, Petr Molodyk, Bin Hu 等ICML 2026 · 被引用 6 次
- Prompting with Pseudo-Code InstructionsMayank Mishra, Prince Kumar, Riyaz A. Bhat, Rudra Murthy V 等EMNLP 2023 · 被引用 4 次
- Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual QuestionsPu Jian, Donglei Yu, Wen Yang, Shuo Ren 等ACL 2025
- CoSQA: 20, 000+ Web Queries for Code Search and Question AnsweringJunjie Huang, Duyu Tang, Linjun Shou, Ming Gong 等ACL 2021
- Large Language Models are Few-Shot Summarizers: Multi-Intent Comment Generation via In-Context LearningMingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang 等ICSE 2024 · 被引用 124 次
