SeRe: A Security-Related Code Review Dataset Aligned with Real-World Review Activities
Zixiao Zhao, Yanjie Jiang, Hui Liu, Kui Liu, Lu Zhang
摘要
Software security vulnerabilities can lead to severe consequences, making early detection essential. Although code review serves as a critical defense mechanism against security flaws, relevant feedback remains scarce due to limited attention to security issues or a lack of expertise among reviewers. Existing datasets and studies primarily focus on general-purpose code review comments, either lacking security-specific annotations or being too limited in scale to support large-scale research. To bridge this gap, we introduce SeRe, a security-related code review dataset, constructed using an active learning-based ensemble classification approach. The proposed approach iteratively refines model predictions through human annotations, achieving high precision while maintaining reasonable recall. Using the fine-tuned ensemble classifier, we extracted 6,732 security-related reviews from 373,824 raw review instances, ensuring representativeness across multiple programming languages. Statistical analysis indicates that SeRe generally aligns with real-world security-related review distribution. To assess both the utility of SeRe and the effectiveness of existing code review comment generation approaches, we benchmark state-of-the-art approaches on security-related feedback generation. By releasing SeRe along with our benchmark results, we aim to advance research in automated security-focused code review and contribute to the development of more effective secure software engineering practices.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan 等FSE 2022 · 被引用 195 次
- Using Pre-Trained Models to Boost Code Review AutomationRosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella 等ICSE 2022 · 被引用 149 次
- CommentFinder: a simpler, faster, more accurate code review comments recommendationYang Hong, Chakkrit Tantithamthavorn, Patanamon Thongtanunam, Aldeida AletiFSE 2022 · 被引用 57 次
- AUGER: automatically generating review comments with pre-training modelsLingwei Li, Li Yang, Huaxi Jiang, Jun Yan 等FSE 2022 · 被引用 56 次
- Software security during modern code review: the developer's perspectiveLarissa Braz, Alberto BacchelliFSE 2022 · 被引用 28 次
相关 Paper
- SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-Aware Fine-TuningFang Liu, Simiao Liu, Yinghao Zhu, Xiaoli Lian 等ICSE 2026
- Empowering Lightweight Language Models for Security Code Review via Context-Aware DistillationZixiao Zhao, Yanjie Jiang, Hui Liu, Lu ZhangISSTA 2026
- An Empirical Study of Static Analysis Tools for Secure Code ReviewWachiraphan Charoenwet, Patanamon Thongtanunam, Van-Thuan Pham, Christoph TreudeISSTA 2024 · 被引用 19 次
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li 等FSE 2026
- ProSec: Fortifying Code LLMs with Proactive Security AlignmentXiangzhe Xu, Zian Su, Jinyao Guo, Kaiyuan Zhang 等ICML 2025
