GPUReplay: a 50-KB GPU stack for client ML
Heejin Park, Felix Xiaozhu Lin
摘要
GPUReplay (GR) is a novel way for deploying GPU-accelerated computation on mobile and embedded devices. It addresses high complexity of a modern GPU stack for deployment ease and security. The idea is to record GPU executions on the full GPU stack ahead of time and replay the executions on new input at run time. We address key challenges towards making GR feasible, sound, and practical to use. The resultant replayer is a drop-in replacement of the original GPU stack. It is tiny (50 KB of executable), robust (replaying long executions without divergence), portable (running in a commodity OS, in TEE, and baremetal), and quick to launch (speeding up startup by up to two orders of magnitude). We show that GPUReplay works with a variety of integrated GPU hardware, GPU APIs, ML frameworks, and 33 neural network (NN) implementations for inference or training. The code is available at https://github.com/bakhi/GPUReplay.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- STI: Turbocharge NLP Inference at the Edge via Elastic PipeliningLiwei Guo, Wonkyo Choe, Felix Xiaozhu LinASPLOS 2023 · 被引用 25 次
- Minimum viable device drivers for ARM trustzoneLiwei Guo, Felix Xiaozhu LinEuroSys 2022 · 被引用 15 次
- ProvCam: A Camera Module with Self-Contained TCB for Producing Verifiable VideosYuxin (Myles) Liu, Zhihao Yao, Mingyi Chen, Ardalan Amiri Sani 等MobiCom 2024 · 被引用 7 次
- Program Environment FuzzingRuijie Meng, Gregory J. Duck, Abhik RoychoudhuryCCS 2024 · 被引用 6 次
- μUSB: Practical and Safe USB Driver Reuse for Arm TrustZoneXuankai Zhang, Sijin Li, Pei Meng, Meng Wang 等OSDI 2026
它引用的顶会 Paper5
- Telekine: Secure Computing with Cloud GPUsTyler Hunt, Zhipeng Jia, Vance Miller, Ariel Szekely 等NSDI 2020 · 被引用 108 次
- Nimble: Lightweight and Parallel GPU Task Scheduling for Deep LearningWoosuk Kwon, Gyeong-In Yu, Eunji Jeong, Byung-Gon ChunNeurIPS 2020 · 被引用 102 次
- Enabling Rack-scale Confidential Computing using Heterogeneous Trusted Execution EnvironmentJianping Zhu, Rui Hou, XiaoFeng Wang, Wenhao Wang 等S&P 2020 · 被引用 95 次
- Enabling Refinable Cross-Host Attack Investigation with Efficient Data Flow Tagging and TrackingYang Ji, Sangho Lee, Mattia Fazzini, Joey Allen 等USENIX Security 2018 · 被引用 70 次
- AvA: Accelerated Virtualization of AcceleratorsHangchen Yu, Arthur Michener Peters, Amogh Akshintala, Christopher J. RossbachASPLOS 2020 · 被引用 33 次
相关 Paper
- Safe and Practical GPU Computation in TrustZoneHeejin Park, Felix Xiaozhu LinEuroSys 2023 · 被引用 18 次
- GraCE: Unlocking CUDA Graphs with Compiler Support for ML WorkloadsAbhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava BasuOSDI 2026
- GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System ResearchMeng Wang, Gus Waldspurger, Naufal Ananda, Yuyang Huang 等VLDB 2025 · 被引用 1 次
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu 等ASPLOS 2026
- RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceUdit Gupta, Samuel Hsia, Jeff Zhang, Mark Wilkening 等MICRO 2021 · 被引用 31 次
