Lune

DAC2025顶会

PIMPAL: Accelerating LLM Inference on Edge Devices via In-DRAM Arithmetic Lookup

Yoonho Jang, Hyeongjun Cho, Yesin Ryu, Jungrae Kim, Seokin Hong

2025年份
6被引次数

摘要

Deploying Large Language Models (LLMs) on edge devices poses significant challenges due to their high computational and memory demands. In particular, General MatrixVector Multiplication (GEMV), a key operation in LLM inference, is highly memory-intensive, making it difficult to accelerate using conventional edge computing systems. While Processing-in-memory (PIM) architectures have emerged as a promising solution to this challenge, they often suffer from high area overhead or restricted computational precision. This paper proposes PIMPAL (Processing-In-Memory architecture with Parallel Arithmetic Lookup), a cost-effective PIM architecture leveraging LookUp Table (LUT)-based computation for GEMV acceleration in sLLMs (small LLMs). By replacing traditional arithmetic operations with parallel in-DRAM LUT lookups, PIMPAL significantly reduces area overhead while maintaining high performance. PIMPAL introduces three key innovations: (1) it divides DRAM bank subarrays into compute blocks for parallel LUT processing; (2) it employs Localityaware Compute Mapping (LCM) to reduce row activations by maximizing LUT access locality; and (3) it enables multi-precision computations through a LUT Aggregation (LAG) mechanism that combines results from multiple small LUTs. Experimental results show that PIMPAL achieves up to 17.8x17.8 x higher performance than previous LUT-based PIM designs and reduces area overhead by 40%40 \% compared to conventional processing unit-based PIM designs.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get cb92ea52-b89f-4b1e-a155-eac57d255414

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖