Enabling Rack-scale Confidential Computing using Heterogeneous Trusted Execution Environment
Jianping Zhu, Rui Hou, XiaoFeng Wang, Wenhao Wang, Jiangfeng Cao, Boyan Zhao, Zhongpu Wang, Yuhui Zhang, Jiameng Ying, Lixin Zhang, Dan Meng
摘要
With its huge real-world demands, large-scale confidential computing still cannot be supported by today’s Trusted Execution Environment (TEE), due to the lack of scalable and effective protection of high-throughput accelerators like GPUs, FPGAs, and TPUs etc. Although attempts have been made recently to extend the CPU-like enclave to GPUs, these solutions require change to the CPU or GPU chips, may introduce new security risks due to the side-channel leaks in CPU-GPU communication and are still under the resource constraint of today’s CPU TEE.To address these problems, we present the first Heterogeneous TEE design that can truly support large-scale compute or data intensive (CDI) computing, without any chip-level change. Our approach, called HETEE, is a device for centralized management of all computing units (e.g., GPUs and other accelerators) of a server rack. It is uniquely designed to work with today’s data centres and clouds, leveraging modern resource pooling technologies to dynamically compartmentalize computing tasks, and enforce strong isolation and reduce TCB through hardware support. More specifically, HETEE utilizes the PCIe ExpressFabric to allocate its accelerators to the server node on the same rack for a non-sensitive CDI task, and move them back into a secure enclave in response to the demand for confidential computing. Our design runs a thin TCB stack for security management on a security controller (SC), while leaving a large set of software (e.g., AI runtime, GPU driver, etc.) to the integrated microservers that operate enclaves. An enclaves is physically isolated from others through hardware and verified by the SC at its inception. Its microserver and computing units are restored to a secure state upon termination.We implemented HETEE on a real hardware system, and evaluated it with popular neural network inference and training tasks. Our evaluations show that HETEE can easily support the CDI tasks on the real-world scale and incurred a maximal throughput overhead of 2.17% for inference and 0.95% for training on ResNet152.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- ShEF: shielded enclaves for cloud FPGAsMark Zhao, Mingyu Gao, Christos KozyrakisASPLOS 2022 · 被引用 53 次
- Honeycomb: Secure and Efficient GPU Executions via Static ValidationHaohui Mai, Jiacheng Zhao, Hongren Zheng, Yiyang Zhao 等OSDI 2023 · 被引用 39 次
- StrongBox: A GPU TEE on Arm EndpointsYunjie Deng, Chenxu Wang, Shunchang Yu, Shiqing Liu 等CCS 2022 · 被引用 37 次
- Confidential Computing within an AI AcceleratorKapil Vaswani, Stavros Volos, Cédric Fournet, Antonio Nino Diaz 等USENIX ATC 2023 · 被引用 31 次
- GuardNN: secure accelerator architecture for privacy-preserving deep learningWeizhe Hua, Muhammad Umar, Zhiru Zhang, G. Edward SuhDAC 2022 · 被引用 28 次
它引用的顶会 Paper4
- Sanctum: Minimal Hardware Extensions for Strong Software IsolationVictor Costan, Ilia A. Lebedev, Srinivas DevadasUSENIX Security 2016 · 被引用 649 次
- Rendered Insecure: GPU Side Channel Attacks are PracticalHoda Naghibijouybari, Ajaya Neupane, Zhiyun Qian, Nael B. Abu-GhazalehCCS 2018 · 被引用 214 次
- A Formal Foundation for Secure Remote Execution of EnclavesPramod Subramanyan, Rohit Sinha, Ilia A. Lebedev, Srinivas Devadas 等CCS 2017 · 被引用 146 次
- SecTEE: A Software-based Approach to Secure Enclave Architecture Using TEEShijun Zhao, Qianying Zhang, Yu Qin, Wei Feng 等CCS 2019 · 被引用 95 次
相关 Paper
- ACAI: Protecting Accelerator Execution with Arm Confidential Computing ArchitectureSupraja Sridhara, Andrin Bertschi, Benedict Schlüter, Mark Kuhne 等USENIX Security 2024 · 被引用 36 次
- TEEM³: Core-Independent and Cooperating Trusted Execution EnvironmentsNils Asmussen, Sebastian Haas, Carsten Weinhold, Nicholas Gordon 等ASPLOS 2026
- TensorTEE: Unifying Heterogeneous TEE Granularity for Efficient Secure Collaborative Tensor ComputingHusheng Han, Xinyao Zheng, Yuanbo Wen, Yifan Hao 等ASPLOS 2024 · 被引用 12 次
- CRONUS: Fault-isolated, Secure and High-performance Heterogeneous Computing for Trusted Execution EnvironmentJianyu Jiang, Ji Qi, Tianxiang Shen, Xusheng Chen 等MICRO 2022 · 被引用 27 次
- HyperTEE: A Decoupled TEE Architecture with Secure Enclave ManagementYunkai Bai, Peinan Li, Yubiao Huang, Michael C. Huang 等MICRO 2024 · 被引用 5 次
