Lune

SC2025顶会

MaverIQ: Fingerprint-Guided Extrapolation and Fragmentation-Aware Layering for Intent-Based LLM Serving

Dimitrios Liakopoulos, Prasoon Sinha, Tianrui Hu, Myungjin Lee, Neeraja J. Yadwadkar

2025年份
2被引次数

摘要

Large Language Models (LLMs) are becoming ubiquitous across industries, where applications demand they fulfill diverse user intents. However, developers currently face the challenge of manually exploring numerous deployment configurations-combinations of parallelism and compression techniques that impact resource usage, latency, cost, and accuracy-to meet these intents. Previous works automate configuration selection and deployment, but, they (a) rely on expensive profiling, and (b) suboptimally utilize the fragmented resource availability in multi-tenant GPU clusters, inflating operational costs for the provider. Moreover, none of these solutions tailors deployment configuration decisions to diverse user intents.

We present MaverIQ, an LLM inference serving system that exposes an intuitive interface allowing users to express their intent (e.g., minimize latency, meet cost target). MaverIQ automatically translates the intent into the best configuration and deploys it on behalf of the user, while minimizing the operational cost for the provider. To reduce profiling costs, MaverIQ introduces and observes a compact proxy of the LLM, called fingerprint, under a few configurations, and uses novel analytical models to extrapolate the observed fingerprint data to the full LLM. To optimize the cost for the provider, we leverage our key observation that, unlike training, distributing LLM layers across GPUs in an uneven manner has little impact on the overall inference latency. Our novel deployment algorithm exploits this insight to enable MaverIQ to utilize the fragmented cluster resources. Through rigorous empirical evaluation, we show that MaverIQ reduces profiling cost by 7-15× compared to state-of-the-art baselines. Lastly, across various LLMs, traces, and loads, MaverIQ best meets user intents while reducing operational cost for the provider by 3.8-8.3×. Our code is available at https://github.com/UT-SysML/MaverIQ.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 2836d017-d27f-4e2b-83f5-2c94c4708367

它引用的顶会 Paper60

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖