Code Comment Inconsistency Detection and Rectification Using a Large Language Model
Guoping Rong, Yongda Yu, Song Liu, Xin Tan, Tianyi Zhang, Haifeng Shen, Jidong Hu
Abstract
Comments are widely used in source code. If a comment is consistent with the code snippet it intends to annotate, it would aid code comprehension. Otherwise, Code Comment Inconsistency (CCI) is not only detrimental to the understanding of code, but more importantly, it would negatively impact the development, testing, and maintenance of software. To tackle this issue, existing research has been primarily focused on detecting inconsistencies with varied performance. It is evident that detection alone does not solve the problem; it merely paves the way for solving it. A complete solution requires detecting inconsistencies and, more importantly, rectifying them by amending comments. However, this type of work is scarce. In this paper, we contribute C4RLLaMA, a fine-tuned large language model based on the open-source CodeLLaMA. It not only has the ability to rectify inconsistencies by correcting relevant comment content but also outperforms state-of-the-art approaches in detecting inconsistencies. Experiments with various datasets confirm that C4RLLaMA consistently surpasses both post hoc and just-in-time CCI detection approaches. More importantly, C4RLLaMA outperforms substantially the only known CCI rectification approach in terms of multiple performance metrics. To further examine C4RLLaMA's efficacy in rectifying inconsistencies, we conducted a manual evaluation, and the results showed that the percentage of correct comment updates by C4RLLaMA was 65.0% and 55.9% in just-in-time and post hoc, respectively, implying C4RLLaMA's real potential in practical use.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6d351208-b2b6-41f9-a83a-5937b7367d04Cited by top-tier papers6
- CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code ReasoningMan Ho Lam, Chaozheng Wang, Jen-Tse Huang, Michael R. LyuNeurIPS 2025 · 16 citations
- SciCoQA: Quality Assurance for Scientific Paper-Code AlignmentTim Baumgärtner, Iryna GurevychACL 2026 · 5 citations
- Testora: Using Natural Language Intent to Detect Behavioral RegressionsMichael PradelICSE 2026 · 1 citation
- CASCADE: Detecting Inconsistencies between Code and Documentation with Automatic Test GenerationTobias Kiecker, Jan Arne Sparka, Martin Reuter, Albert Ziegler et al.FSE 2026 · 1 citation
- Identifying Multi-parameter Constraint Errors in Python Data Science Library API DocumentationXiufeng Xu, Fuman Xie, Chenguang Zhu, Guangdong Bai et al.ISSTA 2025
Related papers
- Deep Just-In-Time Inconsistency Detection Between Comments and Source CodeSheena Panthaplackel, Junyi Jessy Li, Milos Gligoric, Raymond J. MooneyAAAI 2021 · 62 citations
- Towards Better Answers: Automated Stack Overflow Post UpdatingYubo Mai, Zhipeng Gao, Haoye Wang, Tingting Bi et al.ICSE 2025 · 4 citations
- DocPrism: Multi-lingual Detection of Incorrectness Inconsistencies between Code and DocumentationXiaomeng Xu, Zahin Wahab, Reid Holmes, Caroline LemieuxISSTA 2026
- Detecting Code-Comment Inconsistencies in Smart Contracts by Combining LLM and Program AnalysisJiashuo Zhang, Jiachi Chen, Ting Zhang, Yue Li et al.FSE 2026
- Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language ModelsYanlin Wang, Tianyue Jiang, Mingwei Liu, Jiachi Chen et al.FSE 2025 · 6 citations
