LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table Lookup
Xiaohu Tang, Yang Wang, Ting Cao, Li Lyna Zhang, Qi Chen, Deng Cai, Yunxin Liu, Mao Yang
摘要
On-device Deep Neural Network (DNN) inference consumes significant computing resources and development efforts. To alleviate that, we propose LUT-NN, the first system to empower inference by table lookup, to reduce inference cost. LUT-NN learns the typical features for each operator, named centroid, and precompute the results for these centroids to save in lookup tables. During inference, the results of the closest centroids with the inputs can be read directly from the table, as the approximated outputs without computations.
LUT-NN integrates two major novel techniques: (1) differentiable centroid learning through backpropagation, which adapts three levels of approximation to minimize the accuracy impact by centroids; (2) table lookup inference execution, which comprehensively considers different levels of
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on EdgeJianyu Wei, Shijie Cao, Ting Cao, Lingxiao Ma 等EuroSys 2025 · 被引用 30 次
- PIM-DL: Expanding the Applicability of Commodity DRAM-PIMs for Deep Learning via Algorithm-System Co-OptimizationCong Li, Zhe Zhou, Yang Wang, Fan Yang 等ASPLOS 2024 · 被引用 27 次
- Bitnet.cpp: Efficient Edge Inference for Ternary LLMsJinheng Wang, Hansong Zhou, Ting Song, Shijie Cao 等ACL 2025 · 被引用 17 次
- LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning AcceleratorGuoyu Li, Shengyu Ye, Chunyun Chen, Yang Wang 等HPCA 2025 · 被引用 7 次
- Phases, Modalities, Spatial and Temporal Locality: Domain Specific ML Prefetcher for Accelerating Graph AnalyticsPengmiao Zhang, Rajgopal Kannan, Viktor K. PrasannaSC 2023 · 被引用 5 次
它引用的顶会 Paper11
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- And the Bit Goes Down: Revisiting the Quantization of Neural NetworksPierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham 等ICLR 2020 · 被引用 157 次
- ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationCong Guo, Chen Zhang, Jingwen Leng, Zihan Liu 等MICRO 2022 · 被引用 109 次
- Differentiable Product Quantization for End-to-End Embedding CompressionTing Chen, Lala Li, Yizhou SunICML 2020 · 被引用 81 次
- AsyMo: scalable and efficient deep-learning inference on asymmetric mobile CPUsManni Wang, Shaohua Ding, Ting Cao, Yunxin Liu 等MobiCom 2021 · 被引用 69 次
相关 Paper
- NN-LUT: neural approximation of non-linear operations for efficient transformer inferenceJoonsang Yu, Junki Park, Seongmin Park, Minsoo Kim 等DAC 2022 · 被引用 63 次
- LoCaLUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIMJunguk Hong, Changmin Shin, Sukjin Kim, Si Ung Noh 等HPCA 2026
- LookupFFN: Making Transformers Compute-lite for CPU inferenceZhanpeng Zeng, Michael Davies, Pranav Pulijala, Karthikeyan Sankaralingam 等ICML 2023 · 被引用 12 次
- Practical Single-Image Super-Resolution Using Look-Up TableYounghyun Jo, Seon Joo KimCVPR 2021
- LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up TablesXunpeng Yi, Yibing Zhang, Xinyu Xiang, Qinglong Yan 等ICCV 2025 · 被引用 11 次
