Hybe: GPU-NPU Hybrid System for Efficient LLM Inference with Million-Token Context Window
Seungjae Moon, Junseo Cha, Hyunjun Park, Joo-Young Kim
Abstract
The growth of context window size in large language model (LLM) inference poses a very distinct computational challenge of hardware inefficiency.The inefficiency arises from the computational imbalance during LLM inference between the compute-intensive prefill stage, and memory-intensive decode stage.The predominant inference hardware, GPU, boasts large number of cores to excel in the prefill stage, which processes the entire input context at once, but suffers from hardware underutilization in the decode stage, which iteratively generates one output token at a time.In conventional LLM, batching has been able to alleviate the underutilization by generating multiple tokens of different requests.However, batching becomes infeasible in models with large context windows over 100K tokens because the Key-Value (KV) activations dominate the physical memory capacity, surpassing the entire model size.In this paper, we propose Hybe, a GPU-NPU hybrid system for efficient LLM inference with a million-token context window.Hybe utilizes the preexisting GPU for the prefill stage and employs lightweight NPUs during the decode stage.Each NPU includes only the necessary computing resources to fully utilize the given memory bandwidth, thereby achieving maximum hardware efficiency.Furthermore, Hybe introduces fine-grained KV transmission, a kernel scheduling method that immediately offloads partial KV produced from the GPU to the NPU, which significantly reduces the KV memory required in the GPU.Lastly, Hybe scheduler applies stage-wise pipelining that dynamically assigns queued requests to idle hardware to minimize stalls.Hybe utilizes NVIDIA H100 GPU with inference-optimized vLLM library and implement Hybe NPU in 4nm process with equal HBM specification.Hybe achieves 2.1× speedup for Phi-3 with 100K-token context window and 3.9× energy efficiency for Llama-3 with 1M-token context window, over H100 GPUs with equal total device count.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9bffdf9e-bdf2-4456-8d70-ea269e47b003Cited by top-tier papers2
- Early Silicon of Raptor: The First 3D-DRAM Accelerator for Generative InferencePrashant J. Nair, Ramyad Hadidi, Subramani Ganesh, Sangamesh Kodge et al.ISCA 2026 · 4 citations
- PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-Based Long-Context LLM Inference SystemHyucksung Kwon, Kyungmo Koo, Janghyeon Kim, Woongkyu Lee et al.HPCA 2026 · 3 citations
Related papers
- Fast On-device LLM Inference with NPUsDaliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu et al.ASPLOS 2025 · 38 citations
- PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceYufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang et al.ASPLOS 2025 · 44 citations
- POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM InferenceAditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter et al.ASPLOS 2025 · 24 citations
- Lincoln: Real-Time 50 100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash MemoryWeiyi Sun, Mingyu Gao, Zhaoshi Li, Aoyang Zhang et al.HPCA 2025 · 11 citations
- PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing SystemYintao He, Haiyu Mao, Christina Giannoula, Mohammad Sadrosadati et al.ASPLOS 2025 · 37 citations
