DocPrism: Multi-lingual Detection of Incorrectness Inconsistencies between Code and Documentation
Xiaomeng Xu, Zahin Wahab, Reid Holmes, Caroline Lemieux
摘要
Code-documentation inconsistencies are common and undesirable: they can lead to developer misunderstandings and software defects. This paper introduces DocPrism, a lightweight multi-language, code-documentation inconsistency detection tool. DocPrism uses a standard large language model (LLM) to analyze and explain inconsistencies, and focuses on outputting incorrectness inconsistencies. Plain use of LLMs for this task yields unacceptably high inconsistency flag rates—i.e., over 90% of functions are flagged as inconsistent with their documentation. One substantial reason is that LLMs identify natural gaps between high-level documentation and code as incompleteness inconsistencies. We introduce and apply the Local Categorization, External Filtering (LCEF) methodology: LCEF uses an LLM’s local completion skills, rather than its long-term reasoning skills, to focus on reporting incorrectness inconsistencies. In our ablation study, LCEF reduces DocPrism’s inconsistency flag rate from 98% to 14%, and increases F1 score from 0.22 to 0.77, compared to standard prompting techniques. On a broad evaluation across Python, TypeScript, C++, and Java, DocPrism maintains a low flag rate of 17%, and achieves a precision of 0.63 without performing any fine-tuning. We also establish a conservative lower bound across four programming languages, showing that inconsistency errors are present in 11% of code-documentation pairs. In addition, DocPrism achieves precision comparable to the state-of-the-art on an established synthetic dataset, but substantially outperforms it on our real-world Java dataset in precision (DocPrism: 0.47–0.67 vs. SOTA: 0.05–0.14).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Deep Just-In-Time Inconsistency Detection Between Comments and Source CodeSheena Panthaplackel, Junyi Jessy Li, Milos Gligoric, Raymond J. MooneyAAAI 2021 · 被引用 62 次
- Leveraging Large Language Model to Assist Detecting Rust Code Comment InconsistencyYichi Zhang, Zixi Liu, Yang Feng, Baowen XuASE 2024 · 被引用 4 次
- Code Comment Inconsistency Detection and Rectification Using a Large Language ModelGuoping Rong, Yongda Yu, Song Liu, Xin Tan 等ICSE 2025 · 被引用 4 次
- Identifying Multi-parameter Constraint Errors in Python Data Science Library API DocumentationXiufeng Xu, Fuman Xie, Chenguang Zhu, Guangdong Bai 等ISSTA 2025
相关 Paper
- CASCADE: Detecting Inconsistencies between Code and Documentation with Automatic Test GenerationTobias Kiecker, Jan Arne Sparka, Martin Reuter, Albert Ziegler 等FSE 2026 · 被引用 1 次
- Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language ModelsYanlin Wang, Tianyue Jiang, Mingwei Liu, Jiachi Chen 等FSE 2025 · 被引用 6 次
- PRISM: A Benchmark for Unveiling Cross-modal Knowledge Inconsistency in Large Vision-Language ModelsMingjie Wei, Wei-Nan Zhang, Chen Zhang, Yifeng Ding 等ACM MM 2025 · 被引用 1 次
- Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChainMarcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar 等ICLR 2024 · 被引用 39 次
- Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?Madeline Endres, Sarah Fakhoury, Saikat Chakraborty, Shuvendu K. LahiriFSE 2024 · 被引用 27 次
