SLIM: Secure and Efficient Inference for Large Language Models on Untrusted Devices via TEEs
Wei Wang, Zihao Guan, Xing Zhou, Yan Ding, Yusong Tan, Jie Yu, Bao Li
Abstract
Deploying large language models (LLMs) on untrusted hardware entails a risk of weight extraction, which can lead to unauthorized replication and misuse of the model. A practical approach is to leverage Trusted Execution Environments (TEEs) and protect model security by obfuscating model weights. However, existing obfuscation schemes struggle to simultaneously provide strong security guarantees and high performance: schemes with security guarantees incur substantial overhead due to frequent TEE interactions, whereas schemes that achieve efficient inference are insecure. We propose SLIM, a secure inference framework that exploits the iterative structure of LLMs to let transformed representations cascade through consecutive obfuscated layers, thereby minimizing interactions with the TEE. SLIM introduces a T-Way Mixing algorithm that performs consecutive inter-vector covering using carefully constructed block-diagonal Householder matrices and combines it with successive random permutations, providing thorough weight obfuscation while keeping TEE-side computation lightweight. Evaluations demonstrate that SLIM provides robust security guarantees and significantly outperforms prior state-of-the-art obfuscation schemes in terms of performance, delivering up to a speedup while preserving fidelity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang et al.NeurIPS 2025 · 336 citations
- DeepSteal: Advanced Model Extractions Leveraging Efficient Weight Stealing in MemoriesAdnan Siraj Rakin, Md Hafizul Islam Chowdhuryy, Fan Yao, Deliang FanS&P 2022 · 163 citations
- Mind Your Weight(s): A Large-scale Study on Insufficient Machine Learning Model Protection in Mobile AppsZhichuang Sun, Ruimin Sun, Long Lu, Alan MisloveUSENIX Security 2021 · 101 citations
- No Privacy Left Outside: On the (In-)Security of TEE-Shielded DNN Partition for On-Device MLZiqi Zhang, Chen Gong, Yifeng Cai, Yuanyuan Yuan et al.S&P 2024 · 53 citations
- NNSplitter: An Active Defense Solution for DNN Model via Automated Weight ObfuscationTong Zhou, Yukui Luo, Shaolei Ren, Xiaolin XuICML 2023 · 30 citations
Related papers
- LoRO: Real-Time on-Device Secure Inference for LLMs via TEE-Based Low Rank ObfuscationGaojian Xiong, Yu Sun, Jianhua Liu, Jian Cui et al.NeurIPS 2025 · 6 citations
- Game of Arrows: On the (In-)Security of Weight Obfuscation for On-Device TEE-Shielded LLM Partition AlgorithmsPengli Wang, Bingyou Dong, Yifeng Cai, Zheng Zhang et al.USENIX Security 2025
- Understanding the Security Boundary of Obfuscation-based On-Device LLM ProtectionHanyi Zhou, Chenyang Li, Yuanzhe Pang, Ke Xu et al.CCS 2026
- TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZoneXunjie Wang, Jiacheng Shi, Zihan Zhao, Yang Yu et al.EuroSys 2026 · 1 citation
- DeepIvy: Toward Accelerator-Speed Secure Model Inference via CPU-Side TEEsKunbei Cai, Md Hafizul Islam Chowdhuryy, Wujie Wen, Fan YaoCCS 2026
