Rocks Coding, Not Development: A Human-Centric, Experimental Evaluation of LLM-Supported SE Tasks
Wei Wang, Huilong Ning, Gaowei Zhang, Libo Liu, Yi Wang
摘要
and Telecommunications, China Recently, large language models (LLM) based generative AI has been gaining momentum for their impressive high-quality performances in multiple domains, particularly after the release of the ChatGPT. Many believe that they have the potential to perform general-purpose problem-solving in software development and replace human software developers. Nevertheless, there are in a lack of serious investigation into the capability of these LLM techniques in fulfilling software development tasks. In a controlled 2 × 2 between-subject experiment with 109 participants, we examined whether and to what degree working with ChatGPT was helpful in the coding task and typical software development task and how people work with ChatGPT. We found that while ChatGPT performed well in solving simple coding problems, its performance in supporting typical software development tasks was not that good. We also observed the interactions between participants and ChatGPT and found the relations between the interactions and the outcomes. Our study thus provides first-hand insights into using ChatGPT to fulfill software engineering tasks with real-world developers and motivates the need for novel interaction mechanisms that help developers effectively work with large language models to achieve desired outcomes.
CCS Concepts: • Human-centered computing → Laboratory experiments; • Software and its engineering → Software development techniques; Software development methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- OBsmith: LLM-Powered JavaScript Obfuscator TestingShan Jiang, Chenguang Zhu, Sarfraz KhurshidOOPSLA 2026 · 被引用 2 次
- A Comparison of Conversational Models and Humans in Answering Technical Questions: the Firefox CaseJoão Correia, Daniel Coutinho, Marco Castelluccio, Caio Barbosa 等ICSE 2026
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 被引用 465 次
相关 Paper
- Beyond Code Generation: An Observational Study of ChatGPT Usage in Software Engineering PracticeRanim Khojah, Mazen Mohamad, Philipp Leitner, Francisco Gomes de Oliveira NetoFSE 2024 · 被引用 56 次
- From Code Generation to Conceptual Learning: Student Use of LLMs in a Web Programming CourseHajara-Yasmin Isa, Matthew Weston, Muhammad Rizky Wellyanto, Ishita Karna 等CHI 2026 · 被引用 1 次
- Current and Future Use of Large Language Models for Knowledge WorkMichelle Brachman, Amina H. El-Ashry, Casey Dugan, Werner GeyerCSCW 2025 · 被引用 10 次
- ChatDev: Communicative Agents for Software DevelopmentChen Qian, Wei Liu, Hongzhang Liu, Nuo Chen 等ACL 2024
- SWE-GPT: A Process-Centric Language Model for Automated Software ImprovementYingwei Ma, Rongyu Cao, Yongchang Cao, Yue Zhang 等ISSTA 2025 · 被引用 1 次
