LAIKA: Machine Learning-Assisted In-Kernel APU Acceleration
Haoming Zhuo, Dingding Li, Ronghua Lin, Yong Tang
2026Year
Abstract
The integration of machine learning (ML) into OS kernels is severely hampered by the high latency of offloading to discrete GPUs (dGPUs), where data transfers across the PCIe bus can consume over 93% of the total execution time. This paper argues that for many latency-sensitive kernel tasks, the solution is not a more powerful dGPU but a fundamental shift to an I/O-efficient architecture: the integrated GPU (iGPU) found in modern APUs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Towards a Machine Learning-Assisted Kernel with LAKEHenrique Fingler, Isha Tarte, Hangchen Yu, Ariel Szekely et al.ASPLOS 2023 · 19 citations
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody et al.ASPLOS 2026 · 1 citation
- ARK: GPU-driven Code Execution for Distributed Deep LearningChangho Hwang, KyoungSoo Park, Ran Shu, Xinyuan Qu et al.NSDI 2023 · 22 citations
- LithOS: An Operating System for Efficient Machine Learning on GPUsPatrick H. Coppock, Brian Zhang, Eliot H. Solomon, Vasilis Kypriotis et al.SOSP 2025 · 4 citations
- Deadline-Aware Offloading for High-Throughput AcceleratorsTsung Tai Yeh, Matthew D. Sinclair, Bradford M. Beckmann, Timothy G. RogersHPCA 2021 · 16 citations
