TSQP: Safeguarding Real-Time Inference for Quantization Neural Networks on Edge Devices
Yu Sun, Gaojian Xiong, Jianhua Liu, Zheng Liu, Jian Cui
摘要
Quantization Neural Networks (QNNs) has been widely adopted in resource-constrained edge devices due to their real-time capabilities and low resource requirement. However, concerns have arisen regarding that deployed models are white-box available to model thefts. To address this issue, TEE-shielded secure inference has been introduced as a secure and efficient solution. Nevertheless, existing methods neglect the compatibility with 8-bit quantized computation, which leads to severe integer overflow issue during inference. This issue could result a disastrous degradation in QNNs (to random guessing level), completely destroying model utility. Moreover, the model confidentiality and inference integrity also face a substantial threat due to the limited data representation space. To safeguard accurate and efficient inference for QNNs, TEE-Shielded QNN Partition (TSQP) are proposed, which presents three key insights: Firstly, Quantization Manager is designed to convert white-box inference to black-box by shielding critical scales in TEE. Additionally, overflow concerns are effectively addressed using reduced-range approaches. Secondly, by leveraging the Information Bottleneck theory to enhance model training, we introduce Parameter De-Similarity to defend against powerful Model Stealing attacks that existing methods are vulnerable to. Thirdly, the Integrity Monitor is suggested to detect inference integrity breaches in an oblivious manner. In contrast, existing method can be bypassed due to the lack of obliviousness. Experimental results demonstrate that proposed TSQP maintains high accuracy and achieves accurate integrity breaches detection. Our method achieves more than speedup compared to full TEE inference, while reducing Model Stealing attacks accuracy from to . To our best knowledge, proposed method is the first TEE-shielded secure inference solution that achieves model confidentiality, inference integrity and model utility on QNNs.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- LoRO: Real-Time on-Device Secure Inference for LLMs via TEE-Based Low Rank ObfuscationGaojian Xiong, Yu Sun, Jianhua Liu, Jian Cui 等NeurIPS 2025 · 被引用 6 次
- TensorShield: Safeguarding On-Device Inference by Shielding Critical DNN Tensors with TEETong Sun, Bowen Jiang, Hailong Lin, Borui Li 等CCS 2025 · 被引用 3 次
- TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZoneXunjie Wang, Jiacheng Shi, Zihan Zhao, Yang Yu 等EuroSys 2026 · 被引用 1 次
- MEPS: Privacy-preserving Edge-cloud Video Foundation Model Inference with Privacy ProtectabilitySiping Shi, Rui Lu, Dan Wang, Bihai ZhangUSENIX Security 2026
- AegisGuard: RL-Guided Adapter Tuning for TEE-Based Efficient & Secure On-Device InferenceChe Wang, Ziqi Zhang, Yinggui Wang, Tiantong Wang 等NeurIPS 2025
相关 Paper
- No Privacy Left Outside: On the (In-)Security of TEE-Shielded DNN Partition for On-Device MLZiqi Zhang, Chen Gong, Yifeng Cai, Yuanyuan Yuan 等S&P 2024 · 被引用 53 次
- SOTER: Guarding Black-box Inference for General Neural Networks at the EdgeTianxiang Shen, Ji Qi, Jianyu Jiang, Xian Wang 等USENIX ATC 2022 · 被引用 67 次
- Graph in the Vault: Protecting Edge GNN Inference with Trusted Execution EnvironmentRuyi Ding, Tianhong Xu, Aidong Adam Ding, Yunsi FeiDAC 2025 · 被引用 4 次
- Memory-Efficient and Secure DNN Inference on TrustZone-enabled Consumer IoT DevicesXueshuo Xie, Haoxu Wang, Zhaolong Jian, Tao Li 等INFOCOM 2024 · 被引用 11 次
- ShadowNet: A Secure and Efficient On-device Model Inference System for Convolutional Neural NetworksZhichuang Sun, Ruimin Sun, Changming Liu, Amrita Roy Chowdhury 等S&P 2023
