SOTER: Guarding Black-box Inference for General Neural Networks at the Edge
Tianxiang Shen, Ji Qi, Jianyu Jiang, Xian Wang, Siyuan Wen, Xusheng Chen, Shixiong Zhao, Sen Wang, Li Chen, Xiapu Luo, Fengwei Zhang, Heming Cui
Abstract
The prosperity of AI and edge computing has pushed more and more well-trained DNN models to be deployed on thirdparty edge devices to compose mission-critical applications. This necessitates protecting model confidentiality at untrusted devices, and using a co-located accelerator (e.g., GPU) to speed up model inference locally. Recently, the community has sought to improve the security with CPU trusted execution environments (TEE). However, existing solutions either run an entire model in TEE, suffering from extremely high inference latency, or take a partition-based approach to handcraft partial model via parameter obfuscation techniques to run on an untrusted GPU, achieving lower inference latency at the expense of both the integrity of partitioned computations outside TEE and accuracy of obfuscated parameters.
We propose SOTER, the first system that can achieve model confidentiality, integrity, low inference latency and high accuracy in the partition-based approach. Our key observation is that there is often an associativity property among many inference operators in DNN models. Therefore, SOTER automatically transforms a major fraction of associative operators into parameter-morphed, thus confidentiality-preserved operators to execute on untrusted GPU, and fully restores the execution results to accurate results with associativity in TEE. Based on these steps, SOTER further designs an oblivious fingerprinting technique to safely detect integrity breaches of morphed operators outside TEE to ensure correct executions of inferences. Experimental results on six prevalent models in the three most popular categories show that, even with stronger model protection, SOTER achieves comparable performance with partition-based baselines while retaining the same high accuracy as insecure inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9aa98951-4e39-457c-b8af-b384a268498dCited by top-tier papers17
- No Privacy Left Outside: On the (In-)Security of TEE-Shielded DNN Partition for On-Device MLZiqi Zhang, Chen Gong, Yifeng Cai, Yuanyuan Yuan et al.S&P 2024 · 53 citations
- GroupCover: A Secure, Efficient and Scalable Inference Framework for On-device Model Protection based on TEEsZheng Zhang, Na Wang, Ziqi Zhang, Yao Zhang et al.ICML 2024 · 13 citations
- PipeLLM: Fast and Confidential Large Language Model Services with Speculative Pipelined EncryptionYifan Tan, Cheng Tan, Zeyu Mi, Haibo ChenASPLOS 2025 · 10 citations
- CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge DeploymentQinfeng Li, Tianyue Luo, Xuhong Zhang, Yangfan Xie et al.NeurIPS 2025 · 9 citations
- TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge DeploymentQinfeng Li, Zhiqiang Shen, Zhenghan Qin, Yangfan Xie et al.ACM MM 2024 · 9 citations
Builds on12
- GAZELLE: A Low Latency Framework for Secure Neural Network InferenceChiraag Juvekar, Vinod Vaikuntanathan, Anantha P. ChandrakasanUSENIX Security 2018 · 1,075 citations
- Leaky Cauldron on the Dark Land: Understanding Memory Side-Channel Hazards in SGXWenhao Wang, Guoxing Chen, Xiaorui Pan, Yinqian Zhang et al.CCS 2017 · 403 citations
- ZeroTrace : Oblivious Memory Primitives from Intel SGXSajin Sasy, Sergey Gorbunov, Christopher W. FletcherNDSS 2018 · 244 citations
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin et al.AAAI 2020 · 201 citations
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 194 citations
Related papers
- ShadowNet: A Secure and Efficient On-device Model Inference System for Convolutional Neural NetworksZhichuang Sun, Ruimin Sun, Changming Liu, Amrita Roy Chowdhury et al.S&P 2023
- GuardNN: secure accelerator architecture for privacy-preserving deep learningWeizhe Hua, Muhammad Umar, Zhiru Zhang, G. Edward SuhDAC 2022 · 28 citations
- ASGARD: Protecting On-Device Deep Neural Networks with Virtualization-Based Trusted Execution EnvironmentsMyungsuk Moon, Minhee Kim, Joonkyo Jung, Dokyung SongNDSS 2025
- TSQP: Safeguarding Real-Time Inference for Quantization Neural Networks on Edge DevicesYu Sun, Gaojian Xiong, Jianhua Liu, Zheng Liu et al.S&P 2025
- TBNet: A Neural Architectural Defense Framework Facilitating DNN Model Protection in Trusted Execution EnvironmentsZiyu Liu, Tong Zhou, Yukui Luo, Xiaolin XuDAC 2024 · 4 citations
