LoRO: Real-Time on-Device Secure Inference for LLMs via TEE-Based Low Rank Obfuscation
Gaojian Xiong, Yu Sun, Jianhua Liu, Jian Cui, Jianwei Liu
Abstract
While Large Language Models (LLMs) have gained remarkable success, they are consistently at risk of being stolen when deployed on untrusted edge devices. As a solution, TEE-based secure inference has been proposed to protect valuable model property. However, we identify a statistical vulnerability in existing protection methods, and furtherly compromise their security guarantees by proposed Model Stealing Attack with Prior. To eliminate this vulnerability, LoRO is presented in this paper, which leverages dense mask to completely obfuscate parameters. LoRO includes two innovations: (1) Low Rank Mask, which uses low-rank factors to generate dense masks efficiently. The computing complexity in TEE is hence reduced by an exponential amount to achieve inference speed up, while providing robust model confidentiality. (2) Factors Multiplexing, which reuses several cornerstone factors to generate masks for all layers. Compared to one-mask-per-layer, the secure memory requirement is reduced from GB-level to tens of MB, hence avoiding the hundred-fold latency introduced by secure memory paging. Experimental results indicate that LoRO achieve a 0 . 94 × Model Stealing (MS) accuracy, while SOTA methods presents 3 . 37 × at least. The averaged inference latency of LoRO is only 1 . 49 × , compared to the 112 × of TEE-shielded inference. Moreover, LoRO results no accuracy loss, and requires no re-training and structure modification. LoRO can solve the concerns regarding model thefts on edge devices in an efficient and secure manner, facilitating the wide edge application of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8b72fc9-7fdd-4934-9d2b-366c5a27c261Cited by top-tier papers1
Ask how each one uses itBuilds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- CrypTFlow2: Practical 2-Party Secure InferenceDeevashwer Rathee, Mayank Rathee, Nishant Kumar, Nishanth Chandran et al.CCS 2020 · 294 citations
- ZeroTrace : Oblivious Memory Primitives from Intel SGXSajin Sasy, Sergey Gorbunov, Christopher W. FletcherNDSS 2018 · 244 citations
Related papers
- TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge DeploymentQinfeng Li, Zhiqiang Shen, Zhenghan Qin, Yangfan Xie et al.ACM MM 2024 · 9 citations
- SLIM: Secure and Efficient Inference for Large Language Models on Untrusted Devices via TEEsWei Wang, Zihao Guan, Xing Zhou, Yan Ding et al.ICML 2026
- TSQP: Safeguarding Real-Time Inference for Quantization Neural Networks on Edge DevicesYu Sun, Gaojian Xiong, Jianhua Liu, Zheng Liu et al.S&P 2025
- AegisGuard: RL-Guided Adapter Tuning for TEE-Based Efficient & Secure On-Device InferenceChe Wang, Ziqi Zhang, Yinggui Wang, Tiantong Wang et al.NeurIPS 2025
- CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge DeploymentQinfeng Li, Tianyue Luo, Xuhong Zhang, Yangfan Xie et al.NeurIPS 2025 · 9 citations
