TensorTEE: Unifying Heterogeneous TEE Granularity for Efficient Secure Collaborative Tensor Computing
Husheng Han, Xinyao Zheng, Yuanbo Wen, Yifan Hao, Erhu Feng, Ling Liang, Jianan Mu, Xiaqing Li, Tianyun Ma, Pengwei Jin, Xinkai Song, Zidong Du
摘要
Heterogeneous collaborative computing with NPU and CPU has received widespread attention due to its substantial performance benefits. To ensure data confidentiality and integrity during computing, Trusted Execution Environments (TEE) is considered a promising solution because of its comparatively lower overhead. However, existing heterogeneous TEE designs are inefficient for collaborative computing due to fine and different memory granularities between CPU and NPU. 1) The cacheline granularity of CPU TEE intensifies memory pressure due to its extra memory access, and 2) the cacheline granularity MAC of NPU escalates the pressure on the limited memory storage. 3) Data transfer across heterogeneous enclaves relies on the transit of non-secure regions, resulting in cumbersome re-encryption and scheduling.
To address these issues, we propose TensorTEE , a unified tensor-granularity heterogeneous TEE for efficient secure collaborative tensor computing. First, we virtually support tensor granularity in CPU TEE to eliminate the off-chip metadata access by detecting and maintaining tensor structures on-chip. Second, we propose tensor-granularity MAC management with predictive execution to avoid computational stalls while eliminating off-chip MAC storage and access. Moreover, based on the unified granularity, we enable direct data transfer without re-encryption and scheduling dilemmas. Our evaluation is built on enhanced Gem5 and a cycle-accurate NPU simulator. The results show that Ten-sorTEE improves the performance of Large Language Model
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Unified Memory Protection with Multi-granular MAC and Integrity Tree for Heterogeneous ProcessorsSunho Lee, Seonjin Na, Jeongwon Choi, Jinwon Pyo 等ISCA 2025 · 被引用 2 次
- SoK: Analysis of Accelerator TEE DesignsChenxu Wang, Junjie Huang, Yujun Liang, Xuanyao Peng 等NDSS 2026 · 被引用 2 次
它引用的顶会 Paper34
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 被引用 502 次
- Keystone: an open framework for architecting trusted execution environmentsDayeol Lee, David Kohlbrenner, Shweta Shinde, Krste Asanovic 等EuroSys 2020 · 被引用 381 次
- F1: A Fast and Programmable Accelerator for Fully Homomorphic EncryptionNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Srinivas Devadas 等MICRO 2021 · 被引用 294 次
相关 Paper
- Enabling Rack-scale Confidential Computing using Heterogeneous Trusted Execution EnvironmentJianping Zhu, Rui Hou, XiaoFeng Wang, Wenhao Wang 等S&P 2020 · 被引用 95 次
- TEEM³: Core-Independent and Cooperating Trusted Execution EnvironmentsNils Asmussen, Sebastian Haas, Carsten Weinhold, Nicholas Gordon 等ASPLOS 2026
- HyperTEE: A Decoupled TEE Architecture with Secure Enclave ManagementYunkai Bai, Peinan Li, Yubiao Huang, Michael C. Huang 等MICRO 2024 · 被引用 5 次
- CRONUS: Fault-isolated, Secure and High-performance Heterogeneous Computing for Trusted Execution EnvironmentJianyu Jiang, Ji Qi, Tianxiang Shen, Xusheng Chen 等MICRO 2022 · 被引用 27 次
- VirTEE: a full backward-compatible TEE with native live migration and secure I/OJianqiang Wang, Pouya Mahmoody, Ferdinand Brasser, Patrick Jauernig 等DAC 2022 · 被引用 11 次
