RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language Models
Yufei Li, Zexin Li, Wei Yang, Cong Liu
摘要
Recent advancements in language models (LMs) have gained substantial attentions on their capability to generate human-like responses. Though exhibiting a promising future for various applications such as conversation AI, these LMs face deployment challenges on various devices due to their extreme computational cost and unpredictable inference latency. Such varied inference latency, identified as a consequence of uncertainty intrinsic to the nature of language, can lead to computational inefficiency and degrade the overall performance of LMs, especially under high-traffic workloads. Unfortunately, the bandwidth of these uncertainty sources is extensive, complicating the prediction of latency and the effects emanating from such uncertainties. To understand and mitigate the impact of uncertainty on real-time response-demanding systems, we take the first step to comprehend, quantify and optimize these uncertainty-induced latency performance variations in LMs. Specifically, we present RT-LM, an uncertainty-aware resource management ecosystem for real-time inference of LMs. RT-LM innovatively quantifies how specific input uncertainties, recognized within the NLP community, adversely affect latency, often leading to an increased output length. Exploiting these insights, we devise a lightweight yet effective method to dynamically correlate input text uncertainties with output length at runtime. Utilizing this quantification as a latency heuristic, we integrate the uncertainty information into a system-level scheduler which explores several uncertainty-induced optimization opportunities, including uncertainty-aware prioritization, dynamic consolidation, and strategic CPU offloading. Quantitative experiments across five state-of-the-art LMs on two hardware platforms demonstrates that RT-LM can significantly reduce the average response time and improve throughput while incurring a rather small runtime overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DARIS: An Oversubscribed Spatio-Temporal Scheduler for Real-Time DNN Inference on GPUsAmir Fakhim Babaei, Thidapat ChantemDAC 2025 · 被引用 4 次
- Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsYufei Li, Zexin Li, Yinglun Zhu, Cong LiuRTSS 2025 · 被引用 4 次
- Nova: Real-Time Agentic Vision-Language Model Serving With Adaptive Cross-Stage ParallelizationYuhang Xu, Shengzhong Liu, Dong Zhang, Bingheng Yan 等RTSS 2025 · 被引用 1 次
- From Token to Token Pair: Efficient Prompt Compression for Large Language Models in Clinical PredictionMingcheng Zhu, Zhiyao Luo, Yu Liu, Tingting ZhuICML 2026
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao 等OSDI 2020 · 被引用 392 次
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 被引用 96 次
- LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN TasksWoosung Kang, Kilho Lee, Jinkyu Lee, Insik Shin 等RTSS 2021 · 被引用 68 次
相关 Paper
- Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer ModelsDeepak Narayanan, Keshav Santhanam, Peter Henderson, Rishi Bommasani 等NeurIPS 2023 · 被引用 14 次
- An Empirical Study of LLM Reasoning Ability Under Strict Output Length ConstraintYi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu 等EMNLP 2025 · 被引用 1 次
- Energy Considerations of Large Language Model Inference and Efficiency OptimizationsJared Fernandez, Clara Na, Vashisth Tiwari, Yonatan Bisk 等ACL 2025
- Relying on the Unreliable: The Impact of Language Models' Reluctance to Express UncertaintyKaitlyn Zhou, Jena D. Hwang, Xiang Ren, Maarten SapACL 2024
- ExeGPT: Constraint-Aware Resource Scheduling for LLM InferenceHyungjun Oh, Kihong Kim, Jaemin Kim, Sungkyun Kim 等ASPLOS 2024 · 被引用 41 次
