Inference in the Shadows: Taming Memory Bandwidth Contention in Mobile LLM Inference with Sereno
Tong Xin, Xinrui Shi, Mingkai Dong, Zeyu Mi
摘要
The proliferation of large language models (LLMs) on mobile devices introduces a new performance challenge: resource contention between compute-intensive inference and latency-sensitive foreground applications. We identify a severe and asymmetric interference where concurrent LLM inference substantially degrades foreground applications' quality-of-service (QoS)-increasing the aggregate jank rate (the fraction of frames that appear as visible stutters) by 153%. In contrast, LLM throughput degrades by only 1.01% and 1.64% during prefill and decode stages, respectively. This imbalance arises because the hardware prioritizes NPU memory traffic-originally to guarantee critical media tasks (e.g., video recording)-a privilege that best-effort LLM inference inherits, causing aggressive bandwidth contention.
To address this asymmetric degradation, we present SERENO, a foreground-QoS-friendly LLM inference framework that resolves bandwidth contention between foreground applications and background LLM inference, without hardware modification. SERENO repurposes speculative decoding to introduce fine-grained yield points for preemptible execution, letting the system detect memory contention and dynamically yield bandwidth to the foreground without losing inference progress. Extensive evaluations on commercial smartphones across diverse categories of popular applications demonstrate that SERENO reduces the foreground jank rate by up to 92.6% (58.5% on average) while boosting LLM throughput by up to 67.9% (26.4% on average). Compared with vanilla speculative decoding, SERENO can reduce the foreground jank rate by up to 72.1% while incurring only a 6.2% performance degradation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 被引用 424 次
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 被引用 347 次
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 被引用 290 次
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang 等NeurIPS 2024 · 被引用 245 次
相关 Paper
- Towards High-Goodput LLM Serving with Prefill-decode MultiplexingYukang Chen, Weihao Cui, Han Zhao, Ziyi Xu 等ASPLOS 2026 · 被引用 18 次
- LLMShare: Optimizing LLM Inference Serving with Hardware Architecture ExplorationHongduo Liu, Chen Bai, Peng Xu, Lihao Yin 等DAC 2025 · 被引用 2 次
- Progressive Mixed-Precision Decoding for Efficient LLM InferenceHao Mark Chen, Fuwen Tan, Alexandros Kouris, Royson Lee 等ICLR 2025 · 被引用 1 次
- AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative DecodingZikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro 等EuroSys 2026
- QoServe: Breaking the Silos of LLM Inference ServingKanishk Goel, Jayashree Mohan, Nipun Kwatra, Ravi Shreyas Anupindi 等ASPLOS 2026 · 被引用 3 次
