Rocks Coding, Not Development: A Human-Centric, Experimental Evaluation of LLM-Supported SE Tasks
Wei Wang, Huilong Ning, Gaowei Zhang, Libo Liu, Yi Wang
Abstract
and Telecommunications, China Recently, large language models (LLM) based generative AI has been gaining momentum for their impressive high-quality performances in multiple domains, particularly after the release of the ChatGPT. Many believe that they have the potential to perform general-purpose problem-solving in software development and replace human software developers. Nevertheless, there are in a lack of serious investigation into the capability of these LLM techniques in fulfilling software development tasks. In a controlled 2 × 2 between-subject experiment with 109 participants, we examined whether and to what degree working with ChatGPT was helpful in the coding task and typical software development task and how people work with ChatGPT. We found that while ChatGPT performed well in solving simple coding problems, its performance in supporting typical software development tasks was not that good. We also observed the interactions between participants and ChatGPT and found the relations between the interactions and the outcomes. Our study thus provides first-hand insights into using ChatGPT to fulfill software engineering tasks with real-world developers and motivates the need for novel interaction mechanisms that help developers effectively work with large language models to achieve desired outcomes.
CCS Concepts: • Human-centered computing → Laboratory experiments; • Software and its engineering → Software development techniques; Software development methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- OBsmith: LLM-Powered JavaScript Obfuscator TestingShan Jiang, Chenguang Zhu, Sarfraz KhurshidOOPSLA 2026 · 2 citations
- A Comparison of Conversational Models and Humans in Answering Technical Questions: the Firefox CaseJoão Correia, Daniel Coutinho, Marco Castelluccio, Caio Barbosa et al.ICSE 2026
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
Related papers
- Beyond Code Generation: An Observational Study of ChatGPT Usage in Software Engineering PracticeRanim Khojah, Mazen Mohamad, Philipp Leitner, Francisco Gomes de Oliveira NetoFSE 2024 · 56 citations
- From Code Generation to Conceptual Learning: Student Use of LLMs in a Web Programming CourseHajara-Yasmin Isa, Matthew Weston, Muhammad Rizky Wellyanto, Ishita Karna et al.CHI 2026 · 1 citation
- Current and Future Use of Large Language Models for Knowledge WorkMichelle Brachman, Amina H. El-Ashry, Casey Dugan, Werner GeyerCSCW 2025 · 10 citations
- ChatDev: Communicative Agents for Software DevelopmentChen Qian, Wei Liu, Hongzhang Liu, Nuo Chen et al.ACL 2024
- SWE-GPT: A Process-Centric Language Model for Automated Software ImprovementYingwei Ma, Rongyu Cao, Yongchang Cao, Yue Zhang et al.ISSTA 2025 · 1 citation
