CPC: automatically classifying and propagating natural language comments via program analysis
Juan Zhai, Xiangzhe Xu, Yu Shi, Guanhong Tao, Minxue Pan, Shiqing Ma, Lei Xu, Weifeng Zhang, Lin Tan, Xiangyu Zhang
摘要
Code comments provide abundant information that have been leveraged to help perform various software engineering tasks, such as bug detection, specification inference, and code synthesis. However, developers are less motivated to write and update comments, making it infeasible and error-prone to leverage comments to facilitate software engineering tasks. In this paper, we propose to leverage program analysis to systematically derive, refine, and propagate comments. For example, by propagation via program analysis, comments can be passed on to code entities that are not commented such that code bugs can be detected leveraging the propagated comments. Developers usually comment on different aspects of code elements like methods, and use comments to describe various contents, such as functionalities and properties. To more effectively utilize comments, a fine-grained and elaborated taxonomy of comments and a reliable classifier to automatically categorize a comment are needed. In this paper, we build a comprehensive taxonomy and propose using program analysis to propagate comments. We develop a prototype CPC, and evaluate it on 5 projects. The evaluation results demonstrate 41573 new comments can be derived by propagation from other code locations with 88% accuracy. Among them, we can derive precise functional comments for 87 native methods that have neither existing comments nor source code. Leveraging the propagated comments, we detect 37 new bugs in open source large projects, 30 of which have been confirmed and fixed by developers, and 304 defects in existing comments (by looking at inconsistencies between existing and propagated comments), including 12 incomplete comments and 292 wrong comments. This demonstrates the effectiveness of our approach. Our user study confirms propagated comments align well with existing comments in terms of quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Large Language Models are Few-Shot Summarizers: Multi-Intent Comment Generation via In-Context LearningMingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang 等ICSE 2024 · 被引用 124 次
- C2S: translating natural language comments to formal program specificationsJuan Zhai, Yu Shi, Minxue Pan, Guian Zhou 等FSE 2020 · 被引用 44 次
- Source Code Summarization in the Era of Large Language ModelsWeisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang 等ICSE 2025 · 被引用 37 次
- Developer-Intent Driven Code Comment GenerationFangwen Mu, Xiao Chen, Lin Shi, Song Wang 等ICSE 2023 · 被引用 25 次
- PyMT5: multi-mode translation of natural language and Python code with transformersColin B. Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy 等EMNLP 2020 · 被引用 24 次
相关 Paper
- Uncovering the iceberg from the tip: Generating API Specifications for Bug Detection via Specification Propagation AnalysisMiaoqian Lin, Kai Chen, Yi Yang, Jinghua LiuNDSS 2025
- Inference of Error Specifications and Bug Detection Using Structural SimilaritiesNora Dossche, Bart CoppensUSENIX Security 2024 · 被引用 2 次
- Deep Just-In-Time Inconsistency Detection Between Comments and Source CodeSheena Panthaplackel, Junyi Jessy Li, Milos Gligoric, Raymond J. MooneyAAAI 2021 · 被引用 62 次
- Doc2OracLL: Investigating the Impact of Documentation on LLM-Based Test Oracle GenerationSoneya Binta Hossain, Raygan Taylor, Matthew B. DwyerFSE 2025 · 被引用 3 次
- Towards Balanced Defect Prediction with Better Information PropagationXianda Zheng, Yuan-Fang Li, Huan Gao, Yuncheng Hua 等AAAI 2021 · 被引用 2 次
