Automated Code Annotation with LLMs for Establishing TEE Boundaries
Varun Gadey, Melanie Gotz, Christoph Sendner, Sampo Sovio, Alexandra Dmitrienko
摘要
—Modern systems increasingly rely on Trusted Ex-ecution Environments (TEEs), such as Intel SGX and ARM TrustZone, to securely isolate sensitive code and reduce the Trusted Computing Base (TCB). However, identifying the precise regions of code especially those involving cryptographic logic that should reside within a TEE remains challenging, as it requires deep manual inspection and is not supported by automated tools yet. To solve this open problem, we propose LLM based Code Annotation Logic (LLM-CAL), a tool that automates the identification of security-sensitive code regions with a focus on cryptographic implementations by leveraging most recent and advanced Large Language Models (LLMs). Our approach leverages foundational LLMs (Gemma-2B, CodeGemma-2B, and LLaMA-7B), which we fine-tuned using a newly collected and manually labeled dataset of over 4,000 C source files. We encode local context features, global semantic information, and structural metadata into compact input sequences that guide the model in capturing subtle patterns of security sensitivity in code. The fine-tuning process is based on quantized LoRA—a parameter-efficient technique that introduces lightweight, trainable adapters into the LLM architecture. To support practical deployment, we developed a scalable pipeline for data preprocessing and inference. LLM-CAL achieves an F1 score of 98.40% and a recall of 97.50% in identifying sensitive and non-sensitive code. It represents the first effort to automate the annotation of cryptographic security-sensitive code for TEE-enabled platforms, aiming to minimize the Trusted Computing Base (TCB) and optimize TEE usage to enhance overall system security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov 等ICML 2024 · 被引用 820 次
- T-SGX: Eradicating Controlled-Channel Attacks Against Enclave ProgramsMing-Wei Shih, Sangho Lee, Taesoo Kim, Marcus PeinadoNDSS 2017 · 被引用 431 次
- Telling Your Secrets without Page Faults: Stealthy Page Table-Based Attacks on Enclaved ExecutionJo Van Bulck, Nico Weichbrodt, Rüdiger Kapitza, Frank Piessens 等USENIX Security 2017 · 被引用 316 次
相关 Paper
- FHE-Coder: Benchmarking Secure Agentic Code Generation for Fully Homomorphic EncryptionMayank Kumar, Jiaqi Xue, Mengxin Zheng, Qian LouICLR 2026
- Repairing LLM Executions for Secure Automatic ProgrammingAli El Husseini, Yacine Izza, Blaise Genest, Abhik RoychoudhuryICSE 2026
- CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decodingDong Li, Meng Yan, Yaosheng Zhang, Zhongxin Liu 等ISSTA 2024 · 被引用 10 次
- Beyond Static Pattern Matching? Rethinking Automatic Cryptographic API Misuse Detection in the Era of LLMsYifan Xia, Zichen Xie, Peiyu Liu, Kangjie Lu 等ISSTA 2025 · 被引用 2 次
- Efficient Code Analysis via Graph Representation Learning-Guided Large Language ModelsHang Gao, Tao Peng, Baoquan Cui, Hong Huang 等ICML 2026 · 被引用 1 次
