Automated Code Annotation with LLMs for Establishing TEE Boundaries
Varun Gadey, Melanie Gotz, Christoph Sendner, Sampo Sovio, Alexandra Dmitrienko
Abstract
—Modern systems increasingly rely on Trusted Ex-ecution Environments (TEEs), such as Intel SGX and ARM TrustZone, to securely isolate sensitive code and reduce the Trusted Computing Base (TCB). However, identifying the precise regions of code especially those involving cryptographic logic that should reside within a TEE remains challenging, as it requires deep manual inspection and is not supported by automated tools yet. To solve this open problem, we propose LLM based Code Annotation Logic (LLM-CAL), a tool that automates the identification of security-sensitive code regions with a focus on cryptographic implementations by leveraging most recent and advanced Large Language Models (LLMs). Our approach leverages foundational LLMs (Gemma-2B, CodeGemma-2B, and LLaMA-7B), which we fine-tuned using a newly collected and manually labeled dataset of over 4,000 C source files. We encode local context features, global semantic information, and structural metadata into compact input sequences that guide the model in capturing subtle patterns of security sensitivity in code. The fine-tuning process is based on quantized LoRA—a parameter-efficient technique that introduces lightweight, trainable adapters into the LLM architecture. To support practical deployment, we developed a scalable pipeline for data preprocessing and inference. LLM-CAL achieves an F1 score of 98.40% and a recall of 97.50% in identifying sensitive and non-sensitive code. It represents the first effort to automate the annotation of cryptographic security-sensitive code for TEE-enabled platforms, aiming to minimize the Trusted Computing Base (TCB) and optimize TEE usage to enhance overall system security.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa8ae2d8-4307-44d7-807b-d5589914e12cBuilds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov et al.ICML 2024 · 820 citations
- T-SGX: Eradicating Controlled-Channel Attacks Against Enclave ProgramsMing-Wei Shih, Sangho Lee, Taesoo Kim, Marcus PeinadoNDSS 2017 · 431 citations
- Telling Your Secrets without Page Faults: Stealthy Page Table-Based Attacks on Enclaved ExecutionJo Van Bulck, Nico Weichbrodt, Rüdiger Kapitza, Frank Piessens et al.USENIX Security 2017 · 316 citations
Related papers
- FHE-Coder: Benchmarking Secure Agentic Code Generation for Fully Homomorphic EncryptionMayank Kumar, Jiaqi Xue, Mengxin Zheng, Qian LouICLR 2026
- Repairing LLM Executions for Secure Automatic ProgrammingAli El Husseini, Yacine Izza, Blaise Genest, Abhik RoychoudhuryICSE 2026
- CoSec: On-the-Fly Security Hardening of Code LLMs via Supervised Co-decodingDong Li, Meng Yan, Yaosheng Zhang, Zhongxin Liu et al.ISSTA 2024 · 10 citations
- Beyond Static Pattern Matching? Rethinking Automatic Cryptographic API Misuse Detection in the Era of LLMsYifan Xia, Zichen Xie, Peiyu Liu, Kangjie Lu et al.ISSTA 2025 · 2 citations
- Efficient Code Analysis via Graph Representation Learning-Guided Large Language ModelsHang Gao, Tao Peng, Baoquan Cui, Hong Huang et al.ICML 2026 · 1 citation
