Lune

SC2025Top-tier venue

MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM Serving

Dimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee, Neeraja J. Yadwadkar

2025Year
2Citations

Abstract

Large Language Models (LLMs) are becoming ubiquitous across industries, where applications demand they fulfill diverse user intents. However, developers currently face the challenge of manually exploring numerous deployment configurations-combinations of parallelism and compression techniques that impact resource usage, latency, cost, and accuracy-to meet these intents. Previous works automate configuration selection and deployment, but, they (a) rely on expensive profiling, and (b) suboptimally utilize the fragmented resource availability in multi-tenant GPU clusters, inflating operational costs for the provider. Moreover, none of these solutions tailors deployment configuration decisions to diverse user intents.

We present MaverIQ, an LLM inference serving system that exposes an intuitive interface allowing users to express their intent (e.g., minimize latency, meet cost target). MaverIQ automatically translates the intent into the best configuration and deploys it on behalf of the user, while minimizing the operational cost for the provider. To reduce profiling costs, MaverIQ introduces and observes a compact proxy of the LLM, called fingerprint, under a few configurations, and uses novel analytical models to extrapolate the observed fingerprint data to the full LLM. To optimize the cost for the provider, we leverage our key observation that, unlike training, distributing LLM layers across GPUs in an uneven manner has little impact on the overall inference latency. Our novel deployment algorithm exploits this insight to enable MaverIQ to utilize the fragmented cluster resources. Through rigorous empirical evaluation, we show that MaverIQ reduces profiling cost by 7-15× compared to state-of-the-art baselines. Lastly, across various LLMs, traces, and loads, MaverIQ best meets user intents while reducing operational cost for the provider by 3.8-8.3×. Our code is available at https://github.com/UT-SysML/MaverIQ.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 2836d017-d27f-4e2b-83f5-2c94c4708367

Builds on60

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines