HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
Youhe Jiang, Ran Yan, Binhang Yuan
Abstract
Disaggregating the prefill and decoding phases represents an effective new paradigm for generative inference of large language models (LLM), which eliminates prefill-decoding interference and optimizes resource allocation. However, it is still an open problem about how to deploy the disaggregated inference paradigm across a group of heterogeneous GPUs, which can be an economical alternative to deployment over homogeneous high-performance GPUs. Towards this end, we introduce HEXGEN-2, a distributed system for efficient and economical LLM serving on heterogeneous GPUs following the disaggregated paradigm. Built on top of HEXGEN, the core component of HEXGEN-2 is a scheduling algorithm that formalizes the allocation of disaggregated LLM inference computations and communications over heterogeneous GPUs and network connections as a constraint optimization problem. We leverage the graph partitioning and max-flow algorithms to co-optimize resource allocation, parallel strategies for distinct inference phases, and the efficiency of inter-phase key-value (KV) cache communications. We conduct extensive experiments to evaluate HEXGEN-2, i.e., on OPT (30B) and LLAMA-2 (70B) models in various real-world settings, the results reveal that HEXGEN-2 delivers up to a 2.0× and on average a 1.3× improvement in serving throughput, reduces the average inference latency by 1.5× compared with stateof-the-art systems given the same price budget, and achieves comparable inference performance with a 30% lower price budget.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f43b0394-35dc-4ebd-8517-069aa2e9a22aCited by top-tier papers8
- Cascadia: An Efficient Cascade Serving System for Large Language ModelsYouhe Jiang, Fangcheng Fu, Wanru Zhao, Stephan Rabanser et al.ICLR 2026 · 7 citations
- Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline MethodsWanru Zhao, Yihong Chen, Yuzhi Tang, Wentao Ma et al.ICLR 2026 · 4 citations
- OServe: Accelerating LLM Serving via Spatial-Temporal Workload OrchestrationYouhe Jiang, Fangcheng Fu, Taiyi Wang, Guoliang He et al.ICML 2026 · 4 citations
- UEP: Portable Expert-Parallel CommunicationZiming Mao, Yihan Zhang, Chihan Cui, Zhen Huang et al.OSDI 2026
- SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference AccelerationJintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu et al.ICLR 2025
Builds on16
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
Related papers
- HexGen: Generative Inference of Large Language Model over Heterogeneous EnvironmentYouhe Jiang, Ran Yan, Xiaozhe Yao, Yang Zhou et al.ICML 2024 · 46 citations
- HexGen-3: A Fully Disaggregated LLM Serving Framework with Fine-Grained Heterogeneous Resource AutoscalingYouhe Jiang, Wenshuang Li, You Peng, Jintao Zhang et al.ICML 2026
- Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-FlowYixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang et al.ASPLOS 2025 · 33 citations
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah et al.ISCA 2024 · 282 citations
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQLYou Peng, Youhe Jiang, Wenqi Jiang, Chen Wang et al.ICDE 2026
