Empowering Lightweight Language Models for Security Code Review via Context-Aware Distillation
Zixiao Zhao, Yanjie Jiang, Hui Liu, Lu Zhang
摘要
Security code review, which specifically examines software from a security perspective, is indispensable for preempting vulnerabilities and bolstering software reliability. However, existing automated approaches face a dilemma. They either rely on large language models with prohibitive deployment costs, or employ lightweight models that struggle with complex reasoning and global contextual understanding. In this paper, we propose LSCR, a context-aware distillation approach that empowers lightweight language models for security code review by distilling static-analysis–style security rationale from powerful teacher models. Rather than directly transferring final review outputs, LSCR guides student models to internalize how high-level security judgments are systematically derived from low-level code evidence, leveraging repository-level context during training. By embedding this evidence-driven analysis paradigm into lightweight models, LSCR enables more reliable security code reviews under constrained model capacity. When compared with competitive lightweight baselines, LSCR attains an average improvement of over 12% in security issue classification accuracy and more than 13.3% gains in BLEU score for review generation. Human evaluation reveals that LSCR increases the proportion of instrumental reviews by 52.6% on average over state-of-the-art baselines, while reducing misleading feedback by 22.6%. LSCR effectively narrows the performance gap between small-scale and large-scale models, making on-premise security code review more feasible in practice.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-Aware Fine-TuningFang Liu, Simiao Liu, Yinghao Zhu, Xiaoli Lian 等ICSE 2026
- LAURA: Enhancing Code Review Generation with Context-Enriched Retrieval-Augmented LLMYuxin Zhang, Yuxia Zhang, Zeyu Sun, Yanjie Jiang 等ASE 2025 · 被引用 8 次
- SeRe: A Security-Related Code Review Dataset Aligned with Real-World Review ActivitiesZixiao Zhao, Yanjie Jiang, Hui Liu, Kui Liu 等ICSE 2026
- Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity DetectionShize Zhou, Peiyu Liu, Lirong Fu, Tong Ye 等ACL 2026
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li 等FSE 2026
