CORE: Resolving Code Quality Issues using LLMs
Nalin Wadhwa, Jui Pradhan, Atharv Sonwane, Surya Prakash Sahu, Nagarajan Natarajan, Aditya Kanade, Suresh Parthasarathy, Sriram K. Rajamani
Abstract
As software projects progress, quality of code assumes paramount importance as it affects reliability, maintainability and security of software. For this reason, static analysis tools are used in developer workflows to flag code quality issues. However, developers need to spend extra efforts to revise their code to improve code quality based on the tool findings. In this work, we investigate the use of (instruction-following) large language models (LLMs) to assist developers in revising code to resolve code quality issues.
We present a tool, CORE (short for COde REvisions), architected using a pair of LLMs organized as a duo comprised of a proposer and a ranker. Providers of static analysis tools recommend ways to mitigate the tool warnings and developers follow them to revise their code. The proposer LLM of CORE takes the same set of recommendations and applies them to generate candidate code revisions. The candidates which pass the static quality checks are retained. However, the LLM may introduce subtle, unintended functionality changes which may go un-detected by the static analysis. The ranker LLM evaluates the changes made by the proposer using a rubric that closely follows the acceptance criteria that a developer would enforce. CORE uses the scores assigned by the ranker LLM to rank the candidate revisions before presenting them to the developer.
We conduct a variety of experiments on two public benchmarks to show the ability of CORE: (1) to generate code revisions acceptable to both static analysis tools and human reviewers (the latter evaluated with user study on a subset of the Python benchmark), (2) to reduce human review efforts by detecting and eliminating revisions with unintended changes, (3) to readily work across multiple languages (Python and Java), static analysis tools (CodeQL and SonarQube) and quality checks (52 and 10 checks, respectively), and (4) to achieve fix rate comparable to a rule-based automated program repair tool but with much smaller engineering efforts (on the Java benchmark). CORE could revise 59.2% Python files (across 52 quality checks) so that they pass scrutiny by both a tool and a human reviewer. The ranker LLM reduced false positives by 25.8% in these cases. CORE produced revisions that passed the static analysis tool in 76.8% Java files (across 10 quality checks) comparable to 78.3% of a specialized program repair tool, with significantly much less engineering efforts. We release code, data, and supplementary material publicly at http://aka.ms/COREMSRI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- An LLM-Based Agent-Oriented Approach for Automated Code Design Issue LocalizationFraol Batole, David O'Brien, Tien N. Nguyen, Robert Dyer et al.ICSE 2025 · 7 citations
- PortGPT: Towards Automated Backporting Using Large Language ModelsZhaoyang Li, Zheng Yu, Jingyi Song, Meng Xu et al.S&P 2026 · 2 citations
- OBsmith: LLM-Powered JavaScript Obfuscator TestingShan Jiang, Chenguang Zhu, Sarfraz KhurshidOOPSLA 2026 · 2 citations
- Safe4U: Identifying Unsound Safe Encapsulations of Unsafe Calls in Rust using LLMsHuan Li, Bei Wang, Xing Hu, Xin XiaISSTA 2025 · 1 citation
- LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis PipelineZheng Zhang, Haonan Li, Xingyu Li, Hang Zhang et al.NDSS 2026 · 1 citation
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
Related papers
- Fact-Aligned and Template-Constrained Static Analyzer Rule Enhancement with LLMsZongze Jiang, Ming Wen, Ge Wen, Hai JinASE 2025
- LLM-Based Repair of Static Nullability ErrorsNima Karimipour, Pascal Joos, Michael Pradel, Martin Kellogg et al.ISSTA 2026
- Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code GenerationHaoxiang Jia, Robbie Morris, He Ye, Federica Sarro et al.ASE 2025 · 6 citations
- Rethinking the Evaluation of Secure Code GenerationShih-Chieh Dai, Jun Xu, Guanhong TaoICSE 2026 · 1 citation
- Training Language Models to Generate Quality Code with Program Analysis FeedbackFeng Yao, Zilong Wang, Liyuan Liu, Junxia Cui et al.NeurIPS 2025 · 11 citations
