LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
Zhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng, Shijie Cao, Lingxiao Ma, Naifeng Jing, Ting Cao, Jilong Xue, Fan Yang, Mao Yang
Abstract
Large Language Model (LLM) inference becomes resource-intensive, prompting a shift toward low-bit model weights to reduce the memory footprint and improve efficiency. Such low-bit LLMs necessitate the mixed-precision matrix multiplication (mpGEMM), an important yet underexplored operation involving the multiplication of lower-precision weights with higher-precision activations. Off-theshelf hardware does not support this operation natively, leading to indirect, thus inefficient, dequantization-based implementations.
In this paper, we study the lookup table (LUT)-based approach for mpGEMM and find that a conventional LUT implementation fails to achieve the promised gains. To unlock the full potential of LUT-based mpGEMM, we propose LUT Tensor Core, a softwarehardware co-design for low-bit LLM inference. LUT Tensor Core differentiates itself from conventional LUT designs through: 1) software-based optimizations to minimize table precompute overhead and weight reinterpretation to reduce table storage; 2) a LUTbased Tensor Core hardware design with an elongated tiling shape to maximize table reuse and a bit-serial design to support diverse precision combinations in mpGEMM; 3) a new instruction set and compilation optimizations for LUT-based mpGEMM. LUT Tensor Core significantly outperforms existing pure software LUT implementations and achieves a 1.44× improvement in compute density and energy efficiency compared to previous state-of-the-art LUT-based accelerators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cd61ced-dd7f-4e86-a0b8-7daf0d9eaa19Cited by top-tier papers8
- -LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical FormatsYuzong Chen, Chao Fang, Xilai Dai, Yuheng Wu et al.ISCA 2026 · 4 citations
- CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMsGunho Park, Jeongin Bae, Byeongwook Kim, Baeseong Park et al.NeurIPS 2025 · 3 citations
- PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMsRuokai Yin, Yuhang Li, Priyadarshini PandaDAC 2025 · 1 citation
- V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache RetrievalDonghyuk Kim, Sejeong Yang, Wonjin Shin, Joo-Young KimHPCA 2026
- CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-ExpertsXiangyang Yin, Xingyu Liu, Tianhua Xia, BO BAO et al.ICLR 2026
Builds on46
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on EdgeJianyu Wei, Shijie Cao, Ting Cao, Lingxiao Ma et al.EuroSys 2025 · 30 citations
- Omni-LUT: Energy-Efficient LUT-Based Accelerator with Hardware-Aware KV Cache QuantizationCheng-Han Tsai, Kuan-Chen Chou, Yu-Hsin Wang, Chieh-Dun Wen et al.ISCA 2026
- FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up TablesGunho Park, Hyeokjun Kwon, Jiwoo Kim, Jeongin Bae et al.HPCA 2025 · 10 citations
- LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language ModelsGunho Park, Baeseong Park, Minsub Kim, Sungjae Lee et al.ICLR 2024 · 134 citations
- AxCore: A Quantization-Aware Approximate GEMM Unit for LLM InferenceJiaxiang Zou, Yonghao Chen, Xingyu Chen, Chenxi Xu et al.MICRO 2025 · 2 citations
