Pair programming conversations with agents vs. developers: challenges and opportunities for SE community
Peter Robe, Sandeep Kaur Kuttal, Jake AuBuchon, Jacob C. Hart
Abstract
Recent research has shown feasibility of an interactive pair-programming conversational agent, but implementing such an agent poses three challenges: a lack of benchmark datasets, absence of software engineering specific labels, and the need to understand developer conversations. To address these challenges, we conducted a Wizard of Oz study with 14 participants pair programming with a simulated agent and collected 4,443 developer-agent utterances. Based on this dataset, we created 26 software engineering labels using an open coding process to develop a hierarchical classification scheme. To understand labeled developer-agent conversations, we compared the accuracy of three state-of-the-art transformer-based language models, BERT, GPT-2, and XLNet, which performed interchangeably. In order to begin creating a developer-agent dataset, researchers and practitioners need to conduct resource intensive Wizard of Oz studies. Presently, there exists vast amounts of developer-developer conversations on video hosting websites. To investigate the feasibility of using developer-developer conversations, we labeled a publicly available developer-developer dataset (3,436 utterances) with our hierarchical classification scheme and found that a BERT model trained on developer-developer data performed 10% worse than the BERT trained on developer-agent data, but when using transfer-learning, accuracy improved. Finally, our qualitative analysis revealed that developer-developer conversations are more implicit, neutral, and opinionated than developer-agent conversations. Our results have implications for software engineering researchers and practitioners developing conversational agents.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6afa80c1-69c4-4ff1-9d9c-47a264cc84f3Cited by top-tier papers2
- A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and ChallengesJenny T. Liang, Chenyang Yang, Brad A. MyersICSE 2024 · 126 citations
- Toward Systematic Counterfactual Fairness Evaluation of Large Language Models: The CAFFE FrameworkAlessandra Parziale, Gianmario Voria, Valeria Pontillo, Gemma Catolino et al.ICSE 2026
Related papers
- Trade-offs for Substituting a Human with an Agent in a Pair Programming Context: The Good, the Bad, and the UglySandeep Kaur Kuttal, Bali Ong, Kate Kwasny, Peter RobeCHI 2021 · 44 citations
- Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT ModelsJinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang et al.ICSE 2021 · 124 citations
- The Personality Dimensions GPT-3 Expresses During Human-Chatbot InteractionsNikola Kovacevic, Christian Holz, Markus Gross, Rafael WampflerUbiComp 2024 · 15 citations
- How can we assess human-agent interactions? Case studies in software agent designValerie Chen, Rohit Malhotra, Xingyao Wang, Juan Michelini et al.ICML 2026
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 4 citations
