Empowering Lightweight Language Models for Security Code Review via Context-Aware Distillation
Zixiao Zhao, Yanjie Jiang, Hui Liu, Lu Zhang
Abstract
Security code review, which specifically examines software from a security perspective, is indispensable for preempting vulnerabilities and bolstering software reliability. However, existing automated approaches face a dilemma. They either rely on large language models with prohibitive deployment costs, or employ lightweight models that struggle with complex reasoning and global contextual understanding. In this paper, we propose LSCR, a context-aware distillation approach that empowers lightweight language models for security code review by distilling static-analysis–style security rationale from powerful teacher models. Rather than directly transferring final review outputs, LSCR guides student models to internalize how high-level security judgments are systematically derived from low-level code evidence, leveraging repository-level context during training. By embedding this evidence-driven analysis paradigm into lightweight models, LSCR enables more reliable security code reviews under constrained model capacity. When compared with competitive lightweight baselines, LSCR attains an average improvement of over 12% in security issue classification accuracy and more than 13.3% gains in BLEU score for review generation. Human evaluation reveals that LSCR increases the proportion of instrumental reviews by 52.6% on average over state-of-the-art baselines, while reducing misleading feedback by 22.6%. LSCR effectively narrows the performance gap between small-scale and large-scale models, making on-premise security code review more feasible in practice.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get aeead6c9-55fa-4ffb-8b04-b731b79e8c69Related papers
- SecureReviewer: Enhancing Large Language Models for Secure Code Review through Secure-Aware Fine-TuningFang Liu, Simiao Liu, Yinghao Zhu, Xiaoli Lian et al.ICSE 2026
- LAURA: Enhancing Code Review Generation with Context-Enriched Retrieval-Augmented LLMYuxin Zhang, Yuxia Zhang, Zeyu Sun, Yanjie Jiang et al.ASE 2025 · 8 citations
- SeRe: A Security-Related Code Review Dataset Aligned with Real-World Review ActivitiesZixiao Zhao, Yanjie Jiang, Hui Liu, Kui Liu et al.ICSE 2026
- Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity DetectionShize Zhou, Peiyu Liu, Lirong Fu, Tong Ye et al.ACL 2026
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li et al.FSE 2026
