AiF: Accelerating On-Device LLM Inference Using In-Flash Processing
Jaeyong Lee, Hyeunjoo Kim, Sanghun Oh, Myoungjun Chun, Myungsuk Kim, Jihong Kim
摘要
While large language models (LLMs) achieve remarkable performance across diverse application domains, their substantial memory demands present challenges, especially on personal devices with limited DRAM capacity.Recent LLM inference engines have introduced SSD offloading for model parameters to reduce memory footprint.However, the highly memory-bound nature of on-device LLMs makes inference speed heavily dependent on read bandwidth, leading to significant performance degradation due to the limited bandwidth of SSDs.In this paper, we propose an in-flash processing solution for on-device LLM, called Accelerator-in-Flash (AiF), which integrates matrix-vector multiplication (GEMV) operations directly into flash chips.By enabling in-flash GEMV operations, AiF leverages the high internal bandwidth of flash chips without being constrained by the limited external bandwidth.Building on this core structure, AiF employs two novel flash read techniques that were specifically optimized for reading LLM parameters stored in flash memory.AiF achieves a 4x boost in internal read bandwidth during inference with minimal implementation overhead, thanks to its streamlined error correction process.Evaluations on eight real-world LLMs reveal that AiF provides a 14.6x throughput improvement compared to baseline SSD offloading schemes.Furthermore, AiF surpasses in-memory inference, delivering 1.4x higher throughput with a significantly reduced memory footprint.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- FastTTS: Accelerating Test-Time Scaling for Edge LLM ReasoningHao Mark Chen, Zhiwen Mo, Guanxi Lu, Shuang Liang 等ASPLOS 2026 · 被引用 1 次
- V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache RetrievalDonghyuk Kim, Sejeong Yang, Wonjin Shin, Joo-Young KimHPCA 2026
相关 Paper
- LLM in a flash: Efficient Large Language Model Inference with Limited MemoryKeivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, S. Khatamifard 等ACL 2024 · 被引用 73 次
- Lincoln: Real-Time 50 100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash MemoryWeiyi Sun, Mingyu Gao, Zhaoshi Li, Aoyang Zhang 等HPCA 2025 · 被引用 11 次
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 被引用 2 次
- InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM InferenceXiurui Pan, Endian Li, Qiao Li, Shengwen Liang 等HPCA 2025 · 被引用 22 次
- FACIL: Flexible DRAM Address Mapping for SoC-PIM Cooperative On-device LLM InferenceSeong Hoon Seo, Junghoon Kim, Donghyun Lee, Seonah Yoo 等HPCA 2025 · 被引用 7 次
