GPUReplay: a 50-KB GPU stack for client ML
Heejin Park, Felix Xiaozhu Lin
Abstract
GPUReplay (GR) is a novel way for deploying GPU-accelerated computation on mobile and embedded devices. It addresses high complexity of a modern GPU stack for deployment ease and security. The idea is to record GPU executions on the full GPU stack ahead of time and replay the executions on new input at run time. We address key challenges towards making GR feasible, sound, and practical to use. The resultant replayer is a drop-in replacement of the original GPU stack. It is tiny (50 KB of executable), robust (replaying long executions without divergence), portable (running in a commodity OS, in TEE, and baremetal), and quick to launch (speeding up startup by up to two orders of magnitude). We show that GPUReplay works with a variety of integrated GPU hardware, GPU APIs, ML frameworks, and 33 neural network (NN) implementations for inference or training. The code is available at https://github.com/bakhi/GPUReplay.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fadf08d3-cbf6-4443-876b-0d78ae9350adCited by top-tier papers7
- STI: Turbocharge NLP Inference at the Edge via Elastic PipeliningLiwei Guo, Wonkyo Choe, Felix Xiaozhu LinASPLOS 2023 · 25 citations
- Minimum viable device drivers for ARM trustzoneLiwei Guo, Felix Xiaozhu LinEuroSys 2022 · 15 citations
- ProvCam: A Camera Module with Self-Contained TCB for Producing Verifiable VideosYuxin (Myles) Liu, Zhihao Yao, Mingyi Chen, Ardalan Amiri Sani et al.MobiCom 2024 · 7 citations
- Program Environment FuzzingRuijie Meng, Gregory J. Duck, Abhik RoychoudhuryCCS 2024 · 6 citations
- μUSB: Practical and Safe USB Driver Reuse for Arm TrustZoneXuankai Zhang, Sijin Li, Pei Meng, Meng Wang et al.OSDI 2026
Builds on5
- Telekine: Secure Computing with Cloud GPUsTyler Hunt, Zhipeng Jia, Vance Miller, Ariel Szekely et al.NSDI 2020 · 108 citations
- Nimble: Lightweight and Parallel GPU Task Scheduling for Deep LearningWoosuk Kwon, Gyeong-In Yu, Eunji Jeong, Byung-Gon ChunNeurIPS 2020 · 102 citations
- Enabling Rack-scale Confidential Computing using Heterogeneous Trusted Execution EnvironmentJianping Zhu, Rui Hou, XiaoFeng Wang, Wenhao Wang et al.S&P 2020 · 95 citations
- Enabling Refinable Cross-Host Attack Investigation with Efficient Data Flow Tagging and TrackingYang Ji, Sangho Lee, Mattia Fazzini, Joey Allen et al.USENIX Security 2018 · 70 citations
- AvA: Accelerated Virtualization of AcceleratorsHangchen Yu, Arthur Michener Peters, Amogh Akshintala, Christopher J. RossbachASPLOS 2020 · 33 citations
Related papers
- Safe and Practical GPU Computation in TrustZoneHeejin Park, Felix Xiaozhu LinEuroSys 2023 · 18 citations
- GraCE: Unlocking CUDA Graphs with Compiler Support for ML WorkloadsAbhishek Ghosh, Ajay Nayak, Ashish Panwar, Arkaprava BasuOSDI 2026
- GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System ResearchMeng Wang, Gus Waldspurger, Naufal Ananda, Yuyang Huang et al.VLDB 2025 · 1 citation
- FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsZhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu et al.ASPLOS 2026
- RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceUdit Gupta, Samuel Hsia, Jeff Zhang, Mark Wilkening et al.MICRO 2021 · 31 citations
