TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone
Xunjie Wang, Jiacheng Shi, Zihan Zhao, Yang Yu, Zhichao Hua, Jinyu Gu
Abstract
Large Language Models (LLMs) deployed on mobile devices offer benefits like user privacy and reduced network latency, but introduce a significant security risk: the leakage of proprietary models to end users.
To mitigate this risk, we propose a system design for protecting on-device LLMs using Arm Trusted Execution Environment (TEE), TrustZone. Our system addresses two primary challenges: (1) The dilemma between memory efficiency and fast inference (caching model parameters within TEE memory). (2) The lack of efficient and secure Neural Processing Unit (NPU) time-sharing between Rich Execution Environment (REE) and TEE.
Our approach incorporates two key innovations. First, we employ pipelined restoration, leveraging the deterministic memory access patterns of LLM inference to prefetch parameters on demand, hiding memory allocation, I/O and decryption latency under computation time. Second, we introduce a co-driver design, creating a minimal data plane NPU driver in the TEE that collaborates with the full-fledged REE driver. This reduces the TEE TCB size and eliminates control plane reinitialization overhead during NPU world switches.
We implemented our system on the emerging OpenHarmony OS and the llama.cpp inference framework, and evaluated it with various LLMs on an Arm Rockchip device. Compared to a strawman TEE baseline lacking our optimizations, our system reduces TTFT by up to 90.9% and increases decoding speed by up to 23.2%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be881c78-e429-40b8-86fd-0db954dbaf29Builds on35
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu et al.NeurIPS 2022 · 816 citations
- Sanctum: Minimal Hardware Extensions for Strong Software IsolationVictor Costan, Ilia A. Lebedev, Srinivas DevadasUSENIX Security 2016 · 649 citations
- ARMageddon: Cache Attacks on Mobile DevicesMoritz Lipp, Daniel Gruss, Raphael Spreitzer, Clémentine Maurice et al.USENIX Security 2016 · 451 citations
- Keystone: an open framework for architecting trusted execution environmentsDayeol Lee, David Kohlbrenner, Shweta Shinde, Krste Asanovic et al.EuroSys 2020 · 381 citations
Related papers
- ASGARD: Protecting On-Device Deep Neural Networks with Virtualization-Based Trusted Execution EnvironmentsMyungsuk Moon, Minhee Kim, Joonkyo Jung, Dokyung SongNDSS 2025
- ShadowNet: A Secure and Efficient On-device Model Inference System for Convolutional Neural NetworksZhichuang Sun, Ruimin Sun, Changming Liu, Amrita Roy Chowdhury et al.S&P 2023
- TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge DeploymentQinfeng Li, Zhiqiang Shen, Zhenghan Qin, Yangfan Xie et al.ACM MM 2024 · 9 citations
- Fast On-device LLM Inference with NPUsDaliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu et al.ASPLOS 2025 · 38 citations
- Memory-Efficient and Secure DNN Inference on TrustZone-enabled Consumer IoT DevicesXueshuo Xie, Haoxu Wang, Zhaolong Jian, Tao Li et al.INFOCOM 2024 · 11 citations
