Inference in the Shadows: Taming Memory Bandwidth Contention in Mobile LLM Inference with Sereno
Tong Xin, Xinrui Shi, Mingkai Dong, Zeyu Mi
Abstract
The proliferation of large language models (LLMs) on mobile devices introduces a new performance challenge: resource contention between compute-intensive inference and latency-sensitive foreground applications. We identify a severe and asymmetric interference where concurrent LLM inference substantially degrades foreground applications' quality-of-service (QoS)-increasing the aggregate jank rate (the fraction of frames that appear as visible stutters) by 153%. In contrast, LLM throughput degrades by only 1.01% and 1.64% during prefill and decode stages, respectively. This imbalance arises because the hardware prioritizes NPU memory traffic-originally to guarantee critical media tasks (e.g., video recording)-a privilege that best-effort LLM inference inherits, causing aggressive bandwidth contention.
To address this asymmetric degradation, we present SERENO, a foreground-QoS-friendly LLM inference framework that resolves bandwidth contention between foreground applications and background LLM inference, without hardware modification. SERENO repurposes speculative decoding to introduce fine-grained yield points for preemptible execution, letting the system detect memory contention and dynamically yield bandwidth to the foreground without losing inference progress. Extensive evaluations on commercial smartphones across diverse categories of popular applications demonstrate that SERENO reduces the foreground jank rate by up to 92.6% (58.5% on average) while boosting LLM throughput by up to 67.9% (26.4% on average). Compared with vanilla speculative decoding, SERENO can reduce the foreground jank rate by up to 72.1% while incurring only a 6.2% performance degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 447b2c4f-9fa8-4654-8811-68bc6e34fd3eBuilds on26
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangICML 2024 · 424 citations
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 347 citations
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 290 citations
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJunyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang et al.NeurIPS 2024 · 245 citations
Related papers
- Towards High-Goodput LLM Serving with Prefill-decode MultiplexingYukang Chen, Weihao Cui, Han Zhao, Ziyi Xu et al.ASPLOS 2026 · 18 citations
- LLMShare: Optimizing LLM Inference Serving with Hardware Architecture ExplorationHongduo Liu, Chen Bai, Peng Xu, Lihao Yin et al.DAC 2025 · 2 citations
- Progressive Mixed-Precision Decoding for Efficient LLM InferenceHao Mark Chen, Fuwen Tan, Alexandros Kouris, Royson Lee et al.ICLR 2025 · 1 citation
- AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative DecodingZikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro et al.EuroSys 2026
- QoServe: Breaking the Silos of LLM Inference ServingKanishk Goel, Jayashree Mohan, Nipun Kwatra, Ravi Shreyas Anupindi et al.ASPLOS 2026 · 3 citations
