SCPatcher: Mining Crowd Security Discussions to Enrich Secure Coding Practices
Ziyou Jiang, Lin Shi, Guowei Yang, Qing Wang
摘要
Secure coding practices (SCPs) have been proposed to guide software developers to write code securely to prevent potential security vulnerabilities. Yet, they are typically one-sentence principles without detailed specifications, e.g., “Properly free allocated memory upon the completion of functions and at all exit points.”, which makes them difficult to follow in practice, especially for software developers who are not yet experienced in secure programming. To address this problem, this paper proposes SCPatcher, an automated approach to enrich secure coding practices by mining crowd security discussions on online knowledge-sharing platforms, such as Stack Overflow. In particular, for each security post, SCPatcher first extracts the area of coding examples and coding explanations with a fix-prompt tuned Large Language Model (LLM) via Prompt Learning. Then, it hierarchically slices the lengthy code into coding examples and summarizes the coding explanations with the areas. Finally, SCPatcher matches the CWE and Public SCP, integrating them with extracted coding examples and explanations to form the SCP specifications, which are the wild SCPs with details, proposed by the developers. To evaluate the performance of SCPatcher, we conduct experiments on 3,907 security posts from Stack Overflow. The experimental results show that SCPatcher outperforms all baselines in extracting the coding examples with 2.73 % MLine on average, as well as coding explanations with 3.97 % F1 on average. Moreover, we apply SCPatcher on 447 new security posts to further evaluate its practicality, and the extracted SCP specifications enrich the public SCPs with 3,074 lines of code and 1,967 sentences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Few-Shot Text Generation with Natural Language InstructionsTimo Schick, Hinrich SchützeEMNLP 2021 · 被引用 101 次
- CAST: Enhancing Code Summarization with Hierarchical Splitting and Reconstruction of Abstract Syntax TreesEnsheng Shi, Yanlin Wang, Lun Du, Hongyu Zhang 等EMNLP 2021 · 被引用 42 次
相关 Paper
- PatUntrack: Automated Generating Patch Examples for Issue Reports without Tracked Insecure CodeZiyou Jiang, Lin Shi, Guowei Yang, Qing WangASE 2024
- ProSec: Fortifying Code LLMs with Proactive Security AlignmentXiangzhe Xu, Zian Su, Jinyao Guo, Kaiyuan Zhang 等ICML 2025
- APPATCH: Automated Adaptive Prompting Large Language Models for Real-World Software Vulnerability PatchingYu Nong, Haoran Yang, Long Cheng, Hongxin Hu 等USENIX Security 2025
- CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decodingDong Li, Meng Yan, Yaosheng Zhang, Zhongxin Liu 等ISSTA 2024 · 被引用 10 次
- Instruction Tuning for Secure Code GenerationJingxuan He, Mark Vero, Gabriela Krasnopolska, Martin T. VechevICML 2024 · 被引用 69 次
