Lune

ISCA2025顶会

Hybe: GPU-NPU Hybrid System for Efficient LLM Inference with Million-Token Context Window

Seungjae Moon, Junseo Cha, Hyunjun Park, Joo-Young Kim

2025年份
11被引次数
2顶会引用

摘要

The growth of context window size in large language model (LLM) inference poses a very distinct computational challenge of hardware inefficiency.The inefficiency arises from the computational imbalance during LLM inference between the compute-intensive prefill stage, and memory-intensive decode stage.The predominant inference hardware, GPU, boasts large number of cores to excel in the prefill stage, which processes the entire input context at once, but suffers from hardware underutilization in the decode stage, which iteratively generates one output token at a time.In conventional LLM, batching has been able to alleviate the underutilization by generating multiple tokens of different requests.However, batching becomes infeasible in models with large context windows over 100K tokens because the Key-Value (KV) activations dominate the physical memory capacity, surpassing the entire model size.In this paper, we propose Hybe, a GPU-NPU hybrid system for efficient LLM inference with a million-token context window.Hybe utilizes the preexisting GPU for the prefill stage and employs lightweight NPUs during the decode stage.Each NPU includes only the necessary computing resources to fully utilize the given memory bandwidth, thereby achieving maximum hardware efficiency.Furthermore, Hybe introduces fine-grained KV transmission, a kernel scheduling method that immediately offloads partial KV produced from the GPU to the NPU, which significantly reduces the KV memory required in the GPU.Lastly, Hybe scheduler applies stage-wise pipelining that dynamically assigns queued requests to idle hardware to minimize stalls.Hybe utilizes NVIDIA H100 GPU with inference-optimized vLLM library and implement Hybe NPU in 4nm process with equal HBM specification.Hybe achieves 2.1× speedup for Phi-3 with 100K-token context window and 3.9× energy efficiency for Llama-3 with 1M-token context window, over H100 GPUs with equal total device count.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖