Lune

EuroSys2026顶会

PiLLM: Resource-Efficient LLM Inference Using Workload Prediction

Yunqian Fan, Shihao Bai, Ruihao Gong, Zaijun Wang, Rui Fan

2026年份

摘要

LLM inference demands substantial GPU resources, but its highly variable workload characteristics create efficiency challenges at both inter-GPU and intra-GPU levels. To meet Service Level Objectives (SLOs), existing systems typically overprovision resources in two ways: allocating excess GPUs to handle peak loads and reserving excessive memory per request to prevent out-of-memory during token generation. We introduce PiLLM (Predictable inference for LLMs), a system that addresses these inefficiencies through accurate workload prediction and dynamic resource allocation.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get e6326b4e-da03-4b9d-9799-c224566ff1f8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖