CREF: An LLM-Based Conversational Software Repair Framework for Programming Tutors
Boyang Yang, Haoye Tian, Weiguo Pian, Haoran Yu, Haitao Wang, Jacques Klein, Tegawendé F. Bissyandé, Shunfu Jin
Abstract
With the proven effectiveness of Large Language Models (LLMs) in code-related tasks, researchers have explored their potential for program repair. However, existing repair benchmarks might have influenced LLM training data, potentially causing data leakage. To evaluate LLMs’ realistic repair capabilities, (i) we introduce an extensive, non-crawled benchmark TutorCode, comprising 1,239 C++ defect codes and associated information such as tutor guidance, solution description, failing test cases, and the corrected code. Our work assesses LLM’s repair performance on TutorCode, measuring repair correctness (TOP-5 and AVG-5) and patch precision (RPSR). (ii) We then provide a comprehensive investigation into which types of extra information can help LLMs improve their repair performance. Among these types, tutor guidance was the most effective information. To fully harness LLMs’ conversational capabilities and the benefits of augmented information, (iii) we introduce a novel conversational semi-automatic repair framework CREF assisting human programming tutors. It demonstrates a remarkable AVG-5 improvement of 17.2%-24.6% compared to the baseline, achieving an impressive AVG-5 of 76.6% when utilizing GPT-4. These results highlight the potential for enhancing LLMs’ repair capabilities through tutor interactions and historical conversations. The successful application of CREF in a real-world educational setting demonstrates its effectiveness in reducing tutors’ workload and improving students’ learning experience, showing promise for code review and other software engineering tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98db17a8-5c97-4c0c-b77a-12ceff1cf630Cited by top-tier papers9
- GPT-4 as a Homework Tutor Can Improve Student Engagement and Learning OutcomesAlessandro Vanzo, Sankalan Pal Chowdhury, Mrinmaya SachanACL 2025 · 16 citations
- DeclarUI: Bridging Design and Development with Automated Declarative UI Code GenerationTing Zhou, Yanjie Zhao, Xinyi Hou, Xiaoyu Sun et al.FSE 2025 · 13 citations
- Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue RepairKai Huang, Jian Zhang, Xiaofei Xie, Chunyang ChenASE 2025 · 5 citations
- Diagnosing Performance Issues in Application-Defined ResourcesYigong Hu, You-Liang Huang, Haodong Zheng, Yicheng Liu et al.OSDI 2026 · 1 citation
- Input Reduction Enhanced LLM-based Program RepairBoyang Yang, Luyao Ren, Xin Yin, Jiadong Ren et al.ICSE 2026 · 1 citation
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ BugsJian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu et al.ASE 2025
- Automated Program Repair via Conversation: Fixing 162 out of 337 Bugs for $0.42 Each using ChatGPTChunqiu Steven Xia, Lingming ZhangISSTA 2024 · 105 citations
- Exploring Parameter-Efficient Fine-Tuning of Large Language Model on Automated Program RepairGuochang Li, Chen Zhi, Jialiang Chen, Junxiao Han et al.ASE 2024 · 8 citations
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury et al.ICSE 2023 · 213 citations
- HELO-APR: Enhancing Low-Resource Program Repair through Cross-Lingual Knowledge TransferZhipeng Wang, Boyang Yang, Yidong Wan, Liuye Guo et al.ISSTA 2026
