SMOOTH: Hardware-Assisted Fine-Grained On-Chip Memory Management for Efficient On-Device LLM Inference
Seulki Kim, Bokyeong Kim, Kyeonghyeon Ryu, Yeji Jung, Hwanjun Lee, Sungju Kim, Yunhyeong Jeon, Daehoon Kim
Abstract
The growing demand for running large language models (LLMs) directly on mobile devices has intensified the need for efficient on-device inference under stringent memory and bandwidth constraints. While compiler-level optimizations such as memory tiling and lifetime-based allocation improve on-chip SRAM utilization, they remain ineffective in addressing bursty memory traffic and fragmentation arising from the alternating compute- and I/O-bound phases of autoregressive decoding. This paper proposes SMOOTH, a hardware-assisted on-chip memory management framework that dynamically optimizes scratchpad usage at runtime. First, a fine-grained, block-based allocation and preloading scheme improves effective SRAM utilization and exploits idle memory bandwidth. Second, a hardware-driven early reclamation mechanism leverages buffer-level signals to promptly release unused memory blocks, enabling more aggressive and timely preloading. We implement SMOOTH in Verilog and integrate it into LLMCompass, an LLM-optimized extension of ScaleSim, for cycle-accurate evaluation. Experimental results demonstrate that SMOOTH reduces Time-to-First-Token (TTFT) by up to 59.2% and Time-to-Last-Token (TTLT) by up to 73.0% compared to prior baseline approaches on memory-constrained mobile SoCs, achieving average energy reductions of up to 51.2% compared to state-of-the-art baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 39b24025-45cf-45b7-8388-02739ca94e61Related papers
- Understand and Accelerate Memory Processing Pipeline for Large Language Model InferenceZifan He, Rui Ma, Yizhou Sun, Jason CongICML 2026
- Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context UnderstandingKai Liu, Zhan Su, Peijie Dong, Fengran Mo et al.ICLR 2026 · 3 citations
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 290 citations
- LLM in a flash: Efficient Large Language Model Inference with Limited MemoryKeivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, S. Khatamifard et al.ACL 2024 · 73 citations
- AiF: Accelerating On-Device LLM Inference Using In-Flash ProcessingJaeyong Lee, Hyeunjoo Kim, Sanghun Oh, Myoungjun Chun et al.ISCA 2025 · 14 citations
