Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
Wei Zhao, Anand Jayarajan, Gennady Pekhimenko
Abstract
GPU underutilization is a significant concern in many production deep learning clusters, leading to prolonged job queues and increased operational expenses. A promising solution to this inefficiency is GPU sharing, which improves resource utilization by allowing multiple workloads to execute concurrently on a single GPU. However, deploying GPU sharing in production settings faces critical obstacles due to the limitations of existing mechanisms, including high integration costs, inadequate performance isolation, and limited application compatibility. To address these issues, we introduce Tally, a non-intrusive GPU sharing mechanism that provides robust performance isolation and comprehensive workload compatibility. The key to Tally's robust performance isolation capability lies in its fine-grained threadblock-level GPU kernel scheduling strategy, which allows the system to effectively mitigate interference caused by workload co-execution. We evaluate Tally on a diverse range of workloads and show that it incurs an average overhead of only 7.2% on the 99 𝑡ℎ -percentile latency of high-priority inference tasks when executed concurrently with best-effort training workloads, compared to 188.9% overhead exhibited by the state-of-the-art GPU sharing systems like TGS, while achieving over 80% of TGS's system throughput.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3043d149-d018-412f-a064-d96507d3cf3dCited by top-tier papers3
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory BallooningShan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li et al.OSDI 2026 · 33 citations
- Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingPrasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, Neeraja J. YadwadkarSOSP 2026
- NotebookOS: A Replicated Notebook Platform for Interactive Training with On-Demand GPUsBenjamin Carver, Jingyuan Zhang, Haoliang Wang, Kanak Mahadik et al.ASPLOS 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
Related papers
- Transparent GPU Sharing in Container Clouds for Deep Learning WorkloadsBingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu et al.NSDI 2023 · 112 citations
- KubeShare: A Framework to Manage GPUs as First-Class and Shared Resources in Container CloudTing-An Yeh, Hung-Hsin Chen, Jerry ChouHPDC 2020 · 56 citations
- Orion: Interference-aware, Fine-grained GPU Sharing for ML ApplicationsFoteini Strati, Xianzhe Ma, Ana KlimovicEuroSys 2024 · 96 citations
- µShare: Non-Intrusive Kernel Co-Locating on NVIDIA GPUsWenhao Huang, Zhaolin Duan, Laiping Zhao, Yuhao Zhang et al.HPCA 2026
- PipeSwitch: Fast Pipelined Context Switching for Deep Learning ApplicationsZhihao Bai, Zhen Zhang, Yibo Zhu, Xin JinOSDI 2020 · 152 citations
