Securing Retrieval-Augmented Code Generation via Contextual Knowledge Injection: A Case for Embedded IoT Applications
Tong Sun, Jingyi Su, Yi Gao, Wei Dong
摘要
Repository-grounded retrieval-augmented code generation (RACG) is increasingly used in embedded IoT development by retrieving code and documentation from a pinned RTOS/SDK repository (e.g., Zephyr OS). In this setting, security risks are often version-inherited: even without retrieval poisoning, generated applications may invoke benign-looking public APIs that transitively reach vulnerable internal routines in the pinned snapshot, thereby inheriting known CVEs. Existing secure RACG pipelines largely focus on task-level intent and generic vulnerability patterns, which can miss repository- and version-specific exposure. Meanwhile, conventional CVE scanners can flag vulnerable locations but cannot determine whether those vulnerabilities are reachable through the public APIs that the generator commits to during repository-grounded generation. In this paper, we present IoTRAGuarder, a contextual knowledge injection framework that aligns security hardening with generation-time API selection under repository grounding. IoTRAGuarder (i) recovers auditable reverse call chains from CVE-localized internals to exposing public APIs via static analysis plus an evidence-gated LLM to bridge indirections and macro-driven "call-graph islands", (ii) constructs a version-aware security knowledge base that binds affected version intervals to exposed public APIs with prompt-ready constraints, safer alternatives, or avoidance/upgrade guidance, and (iii) performs dual-layer, API-aligned online retrieval to inject concise, version-matched constraints into the final prompt. We evaluate IoTRAGuarder on 44 real-world Zephyr tasks across four LLMs. Compared to the prior state-of-the-art secure RACG baseline, IoTRAGuarder improves the overall security success rate from 5.11% to 78.41%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt 等S&P 2022 · 被引用 725 次
- ReACC: A Retrieval-Augmented Code Completion FrameworkShuai Lu, Nan Duan, Hojae Han, Daya Guo 等ACL 2022 · 被引用 208 次
- Large Language Models for Code: Security Hardening and Adversarial TestingJingxuan He, Martin T. VechevCCS 2023 · 被引用 98 次
- RT-TEE: Real-time System Availability for Cyber-physical Systems using ARM TrustZoneJinwen Wang, Ao Li, Haoran Li, Chenyang Lu 等S&P 2022 · 被引用 69 次
- Instruction Tuning for Secure Code GenerationJingxuan He, Mark Vero, Gabriela Krasnopolska, Martin T. VechevICML 2024 · 被引用 69 次
相关 Paper
- Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge InjectionBo Lin, Shangwen Wang, Yihao Qin, Liqian Chen 等CCS 2025
- CPscan: Detecting Bugs Caused by Code Pruning in IoT KernelsLirong Fu, Shouling Ji, Kangjie Lu, Peiyu Liu 等CCS 2021 · 被引用 7 次
- RESCUE: Retrieval Augmented Secure Code GenerationJiahao Shi, Tianyi ZhangICLR 2026 · 被引用 15 次
- Safe RAG by RAG: Untying the Bell That RAG Rang with the RAG HandXun Liang, Mengwei Wang, Yuefeng Ma, Simin NiuAAAI 2026
- RTCON: Context-Adaptive Function-Level Fuzzing for RTOS KernelsEunkyu Lee, Junyoung Park, Insu YunNDSS 2026 · 被引用 1 次
