Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
Yixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang, Zhihao Jia, Rashmi Vinayak
Abstract
This paper introduces Helix, a distributed system for high-throughput, low-latency large language model (LLM) serving in heterogeneous GPU clusters. The key idea behind Helix is to formulate inference computation of LLMs over heterogeneous GPUs and network connections as a max-flow problem on directed, weighted graphs, whose nodes represent GPU instances and edges capture both GPU and network heterogeneity through their capacities. Helix then uses a mixed integer linear programming (MILP) algorithm to discover highly optimized strategies to serve LLMs on heterogeneous GPUs. This approach allows Helix to jointly optimize model placement and request scheduling, two highly entangled tasks in heterogeneous LLM serving. Our evaluation on several heterogeneous clusters ranging from 24 to 42 GPU nodes shows that Helix improves serving throughput by up to 3.3x and reduces prompting and decoding latency by up to 66% and 24%, respectively, compared to existing approaches. Helix is available at https://github.com/Thesys-lab/Helix-ASPLOS25.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e20b4943-3b4a-4f61-bce2-7554b684db2bCited by top-tier papers20
- vAttention: Dynamic Memory Management for Serving LLMs without PagedAttentionRamya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee et al.ASPLOS 2025 · 38 citations
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 20 citations
- GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at ScaleTianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang et al.USENIX ATC 2025 · 19 citations
- L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsXiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang et al.NeurIPS 2025 · 16 citations
- Enabling Efficient GPU Communication over Multiple NICs with FuseLinkZhenghang Ren, Yuxuan Li, Zilong Wang, Xinyang Huang et al.OSDI 2025 · 10 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
Related papers
- HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous EnvironmentYouhe Jiang, Ran Yan, Binhang YuanICLR 2025
- Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic ParallelismZizhao Mo, Jianxiong Liao, Huanle Xu, Zhi Zhou et al.SC 2025 · 3 citations
- HexGen: Generative Inference of Large Language Model over Heterogeneous EnvironmentYouhe Jiang, Ran Yan, Xiaozhe Yao, Yang Zhou et al.ICML 2024 · 46 citations
- Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUsYouhe Jiang, Fangcheng Fu, Xiaozhe Yao, Guoliang He et al.ICML 2025
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah et al.ISCA 2024 · 282 citations
