AiF: Accelerating On-Device LLM Inference Using In-Flash Processing
Jaeyong Lee, Hyeunjoo Kim, Sanghun Oh, Myoungjun Chun, Myungsuk Kim, Jihong Kim
Abstract
While large language models (LLMs) achieve remarkable performance across diverse application domains, their substantial memory demands present challenges, especially on personal devices with limited DRAM capacity.Recent LLM inference engines have introduced SSD offloading for model parameters to reduce memory footprint.However, the highly memory-bound nature of on-device LLMs makes inference speed heavily dependent on read bandwidth, leading to significant performance degradation due to the limited bandwidth of SSDs.In this paper, we propose an in-flash processing solution for on-device LLM, called Accelerator-in-Flash (AiF), which integrates matrix-vector multiplication (GEMV) operations directly into flash chips.By enabling in-flash GEMV operations, AiF leverages the high internal bandwidth of flash chips without being constrained by the limited external bandwidth.Building on this core structure, AiF employs two novel flash read techniques that were specifically optimized for reading LLM parameters stored in flash memory.AiF achieves a 4x boost in internal read bandwidth during inference with minimal implementation overhead, thanks to its streamlined error correction process.Evaluations on eight real-world LLMs reveal that AiF provides a 14.6x throughput improvement compared to baseline SSD offloading schemes.Furthermore, AiF surpasses in-memory inference, delivering 1.4x higher throughput with a significantly reduced memory footprint.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2e46e7f0-5fb9-4674-8868-c11c275f73f9Cited by top-tier papers2
- FastTTS: Accelerating Test-Time Scaling for Edge LLM ReasoningHao Mark Chen, Zhiwen Mo, Guanxi Lu, Shuang Liang et al.ASPLOS 2026 · 1 citation
- V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache RetrievalDonghyuk Kim, Sejeong Yang, Wonjin Shin, Joo-Young KimHPCA 2026
Related papers
- LLM in a flash: Efficient Large Language Model Inference with Limited MemoryKeivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, S. Khatamifard et al.ACL 2024 · 73 citations
- Lincoln: Real-Time 50 100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash MemoryWeiyi Sun, Mingyu Gao, Zhaoshi Li, Aoyang Zhang et al.HPCA 2025 · 11 citations
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 2 citations
- InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM InferenceXiurui Pan, Endian Li, Qiao Li, Shengwen Liang et al.HPCA 2025 · 22 citations
- FACIL: Flexible DRAM Address Mapping for SoC-PIM Cooperative On-device LLM InferenceSeong Hoon Seo, Junghoon Kim, Donghyun Lee, Seonah Yoo et al.HPCA 2025 · 7 citations
