Beyond Code Generation: An Observational Study of ChatGPT Usage in Software Engineering Practice
Ranim Khojah, Mazen Mohamad, Philipp Leitner, Francisco Gomes de Oliveira Neto
Abstract
Large Language Models (LLMs) are frequently discussed in academia and the general public as support tools for virtually any use case that relies on the production of text, including software engineering. Currently, there is much debate, but little empirical evidence, regarding the practical usefulness of LLM-based tools such as ChatGPT for engineers in industry. We conduct an observational study of 24 professional software engineers who have been using ChatGPT over a period of one week in their jobs, and qualitatively analyse their dialogues with the chatbot as well as their overall experience (as captured by an exit survey). We nd that rather than expecting ChatGPT to generate ready-to-use software artifacts (e.g., code), practitioners more often use ChatGPT to receive guidance on how to solve their tasks or learn about a topic in more abstract terms. We also propose a theoretical framework for how the (i) purpose of the interaction, (ii) internal factors (e.g., the user's personality), and (iii) external factors (e.g., company policy) together shape the experience (in terms of perceived usefulness and trust). We envision that our framework can be used by future research to further the academic discussion on LLM usage by software engineering practitioners, and to serve as a reference point for the design of future empirical LLM research in this domain. CCS Concepts: • Software and its engineering; • Human-centered computing → Natural language interfaces;
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f46ca837-d03a-40b6-ad7f-60ddc78265acCited by top-tier papers8
- Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming SupportKevin Pu, Daniel Lazaro, Ian Arawjo, Haijun Xia et al.CHI 2025 · 31 citations
- "Create a Fear of Missing Out" - ChatGPT Implements Unsolicited Deceptive Designs in Generated Websites Without WarningVeronika Krauß, Mark McGill, Thomas Kosch, Yolanda Maira Thiel et al.CHI 2025 · 17 citations
- "Maybe We Need Some More Examples:" Individual and Team Drivers of Developer GenAI Tool UseCourtney Miller, Rudrajit Choudhuri, Mara Ulloa, Sankeerti Haniyur et al.ICSE 2026 · 1 citation
- Beyond the Desk: Barriers and Future Opportunities for AI to Assist Scientists in Embodied Physical TasksIrene Hou, Alexander Qin, Lauren Cheng, Philip J. GuoCHI 2026 · 1 citation
- Toward Systematic Counterfactual Fairness Evaluation of Large Language Models: The CAFFE FrameworkAlessandra Parziale, Gianmario Voria, Valeria Pontillo, Gemma Catolino et al.ICSE 2026
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 408 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- On the Robustness of Code Generation Techniques: An Empirical Study on GitHub CopilotAntonio Mastropaolo, Luca Pascarella, Emanuela Guglielmi, Matteo Ciniselli et al.ICSE 2023 · 124 citations
- An empirical study of bots in software development: characteristics and challenges from a practitioner's perspectiveLinda Erlenhov, Francisco Gomes de Oliveira Neto, Philipp LeitnerFSE 2020 · 47 citations
Related papers
- Rocks Coding, Not Development: A Human-Centric, Experimental Evaluation of LLM-Supported SE TasksWei Wang, Huilong Ning, Gaowei Zhang, Libo Liu et al.FSE 2024 · 17 citations
- "If the Machine Is As Good As Me, Then What Use Am I?" - How the Use of ChatGPT Changes Young Professionals' Perception of Productivity and AccomplishmentCharlotte Kobiella, Yarhy Said Flores López, Franz Waltenberger, Fiona Draxler et al.CHI 2024 · 54 citations
- Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage ScholarsJiyao Wang, Haolong Hu, Zuyuan Wang, Song Yan et al.CHI 2024 · 21 citations
- Code Red! On the Harmfulness of Applying Off-the-Shelf Large Language Models to Programming TasksAli Al-Kaswan, Sebastian Deatc, Begüm Koç, Arie van Deursen et al.FSE 2025 · 1 citation
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
