TSQP: Safeguarding Real-Time Inference for Quantization Neural Networks on Edge Devices
Yu Sun, Gaojian Xiong, Jianhua Liu, Zheng Liu, Jian Cui
Abstract
Quantization Neural Networks (QNNs) has been widely adopted in resource-constrained edge devices due to their real-time capabilities and low resource requirement. However, concerns have arisen regarding that deployed models are white-box available to model thefts. To address this issue, TEE-shielded secure inference has been introduced as a secure and efficient solution. Nevertheless, existing methods neglect the compatibility with 8-bit quantized computation, which leads to severe integer overflow issue during inference. This issue could result a disastrous degradation in QNNs (to random guessing level), completely destroying model utility. Moreover, the model confidentiality and inference integrity also face a substantial threat due to the limited data representation space. To safeguard accurate and efficient inference for QNNs, TEE-Shielded QNN Partition (TSQP) are proposed, which presents three key insights: Firstly, Quantization Manager is designed to convert white-box inference to black-box by shielding critical scales in TEE. Additionally, overflow concerns are effectively addressed using reduced-range approaches. Secondly, by leveraging the Information Bottleneck theory to enhance model training, we introduce Parameter De-Similarity to defend against powerful Model Stealing attacks that existing methods are vulnerable to. Thirdly, the Integrity Monitor is suggested to detect inference integrity breaches in an oblivious manner. In contrast, existing method can be bypassed due to the lack of obliviousness. Experimental results demonstrate that proposed TSQP maintains high accuracy and achieves accurate integrity breaches detection. Our method achieves more than speedup compared to full TEE inference, while reducing Model Stealing attacks accuracy from to . To our best knowledge, proposed method is the first TEE-shielded secure inference solution that achieves model confidentiality, inference integrity and model utility on QNNs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 43b0e998-0751-4a9a-8b53-b554a7a9eb20Cited by top-tier papers8
- LoRO: Real-Time on-Device Secure Inference for LLMs via TEE-Based Low Rank ObfuscationGaojian Xiong, Yu Sun, Jianhua Liu, Jian Cui et al.NeurIPS 2025 · 6 citations
- TensorShield: Safeguarding On-Device Inference by Shielding Critical DNN Tensors with TEETong Sun, Bowen Jiang, Hailong Lin, Borui Li et al.CCS 2025 · 3 citations
- TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZoneXunjie Wang, Jiacheng Shi, Zihan Zhao, Yang Yu et al.EuroSys 2026 · 1 citation
- MEPS: Privacy-preserving Edge-cloud Video Foundation Model Inference with Privacy ProtectabilitySiping Shi, Rui Lu, Dan Wang, Bihai ZhangUSENIX Security 2026
- AegisGuard: RL-Guided Adapter Tuning for TEE-Based Efficient & Secure On-Device InferenceChe Wang, Ziqi Zhang, Yinggui Wang, Tiantong Wang et al.NeurIPS 2025
Related papers
- No Privacy Left Outside: On the (In-)Security of TEE-Shielded DNN Partition for On-Device MLZiqi Zhang, Chen Gong, Yifeng Cai, Yuanyuan Yuan et al.S&P 2024 · 53 citations
- SOTER: Guarding Black-box Inference for General Neural Networks at the EdgeTianxiang Shen, Ji Qi, Jianyu Jiang, Xian Wang et al.USENIX ATC 2022 · 67 citations
- Graph in the Vault: Protecting Edge GNN Inference with Trusted Execution EnvironmentRuyi Ding, Tianhong Xu, Aidong Adam Ding, Yunsi FeiDAC 2025 · 4 citations
- Memory-Efficient and Secure DNN Inference on TrustZone-enabled Consumer IoT DevicesXueshuo Xie, Haoxu Wang, Zhaolong Jian, Tao Li et al.INFOCOM 2024 · 11 citations
- ShadowNet: A Secure and Efficient On-device Model Inference System for Convolutional Neural NetworksZhichuang Sun, Ruimin Sun, Changming Liu, Amrita Roy Chowdhury et al.S&P 2023
