FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
Gunho Park, Hyeokjun Kwon, Jiwoo Kim, Jeongin Bae, Baeseong Park, Dongsoo Lee, Youngjoo Lee
摘要
Weight-only quantization has emerged as a promising solution to the deployment challenges of large language models (LLMs). However, it necessitates FP-INT operations, which make implementation on general-purpose hardware like GPUs difficult. In this paper, we propose FIGLUT, an efficient look-up table (LUT)-based GEMM accelerator architecture. Instead of performing traditional arithmetic operations, FIGLUT retrieves precomputed values from an LUT based on weight patterns, significantly reducing the computational complexity. We also introduce a novel LUT design that addresses the limitations of conventional memory architectures. To further improve LUT-based operations, we propose a half-size LUT combined with a dedicated decoding and multiplexing unit. FIGLUT efficiently supports different bit precisions and quantization methods using a single fixed hardware configuration. For the same 3-bit weight precision, FIGLUT demonstrates 59% higher TOPS/W and 20% lower perplexity than state-of-the-art accelerator design. When targeting the same perplexity, FIGLUT achieves higher TOPS/W by performing 2.4-bit operations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMsGunho Park, Jeongin Bae, Beomseok Kwon, Byeongwook Kim 等ICLR 2026 · 被引用 8 次
- -LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical FormatsYuzong Chen, Chao Fang, Xilai Dai, Yuheng Wu 等ISCA 2026 · 被引用 4 次
- CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMsGunho Park, Jeongin Bae, Byeongwook Kim, Baeseong Park 等NeurIPS 2025 · 被引用 3 次
- AxCore: A Quantization-Aware Approximate GEMM Unit for LLM InferenceJiaxiang Zou, Yonghao Chen, Xingyu Chen, Chenxi Xu 等MICRO 2025 · 被引用 2 次
- EVA: Accelerating LLM Decoding via an Efficient Vector Quantization ArchitectureBowen Duan, Cong Guo, Chiyue Wei, Haoxuan Shan 等ISCA 2026 · 被引用 2 次
它引用的顶会 Paper17
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu 等NeurIPS 2022 · 被引用 816 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali 等DAC 2021 · 被引用 325 次
相关 Paper
- LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceZhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng 等ISCA 2025 · 被引用 17 次
- Omni-LUT: Energy-Efficient LUT-Based Accelerator with Hardware-Aware KV Cache QuantizationCheng-Han Tsai, Kuan-Chen Chou, Yu-Hsin Wang, Chieh-Dun Wen 等ISCA 2026
- FIGNA: Integer Unit-Based Accelerator Design for FP-INT GEMM Preserving Numerical AccuracyJaeyong Jang, Yulhwa Kim, Juheun Lee, Jae-Joon KimHPCA 2024 · 被引用 41 次
- GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language ModelsPengxiang Zhao, Xiaoming YuanICML 2025
- OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference AccelerationXueying Wu, Baijun Zhou, Zhihui Gao, Yuzhe Fu 等ISCA 2026
