Lune

EuroSys2026Top-tier venue

PiLLM: Resource-Efficient LLM Inference Using Workload Prediction

Yunqian Fan, Shihao Bai, Ruihao Gong, Zaijun Wang, Rui Fan

2026Year

Abstract

LLM inference demands substantial GPU resources, but its highly variable workload characteristics create efficiency challenges at both inter-GPU and intra-GPU levels. To meet Service Level Objectives (SLOs), existing systems typically overprovision resources in two ways: allocating excess GPUs to handle peak loads and reserving excessive memory per request to prevent out-of-memory during token generation. We introduce PiLLM (Predictable inference for LLMs), a system that addresses these inefficiencies through accurate workload prediction and dynamic resource allocation.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get e6326b4e-da03-4b9d-9799-c224566ff1f8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines