SC2021Top-tier venue
APNN-TC: accelerating arbitrary precision neural networks on ampere GPU tensor cores
Boyuan Feng, Yuke Wang, Tong Geng, Ang Li, Yufei Ding
Abstract
Over the years, accelerating neural networks with quantization has been widely studied. Unfortunately, prior efforts with diverse precisions (e.g., 1-bit weights and 2-bit activations) are usually restricted by limited precision support on GPUs (e.g., int1 and int4). To break such restrictions, we introduce the first Arbitrary Precision Neural Network framework (APNN-TC) 1 to fully exploit quantization benefits on Ampere GPU Tensor Cores. Specifically, APNN-TC first incorporates a novel emulation algorithm to support arbitrary short bit-width computation with int1 compute primitives and XOR/AND Boolean operations. Second, APNN-TC integrates arbitrary precision layer designs to efficiently map our emulation algorithm to Tensor Cores with novel batching strategies and specialized memory organization. Third, APNN-TC embodies a novel arbitrary precision NN design to minimize memory access across layers and further improve performance. Extensive evaluations show that APNN-TC can achieve significant speedup over CUT-LASS kernels and various NN models, such as ResNet and VGG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca0a48ce-3c85-4a8d-a9b9-a8530ffcc95cCited by top-tier papers13
- TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPUShengyu Fan, Zhiwei Wang, Weizhi Xu, Rui Hou et al.HPCA 2023 · 90 citations
- DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector MultiplicationYuechen Lu, Weifeng LiuSC 2023 · 37 citations
- Efficient Quantized Sparse Matrix Operations on Tensor CoresShigang Li, Kazuki Osawa, Torsten HoeflerSC 2022 · 27 citations
- DREW: Efficient Winograd CNN Inference with Deep ReuseRuofan Wu, Feng Zhang, Jiawei Guan, Zhen Zheng et al.WWW 2022 · 20 citations
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
Builds on6
- AutoCompress: An Automatic DNN Structured Pruning Framework for Ultra-High Compression RatesNing Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang et al.AAAI 2020 · 204 citations
- PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesXiaolong Ma, Fu-Ming Guo, Wei Niu, Xue Lin et al.AAAI 2020 · 201 citations
- Searching for Low-Bit Weights in Quantized Neural NetworksZhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu et al.NeurIPS 2020 · 103 citations
- ChrEn: Cherokee-English Machine Translation for Endangered Language RevitalizationShiyue Zhang, Benjamin Frey, Mohit BansalEMNLP 2020 · 21 citations
- Boosting Deep Neural Network Efficiency with Dual-Module InferenceLiu Liu, Lei Deng, Zhaodong Chen, Yuke Wang et al.ICML 2020 · 9 citations
Related papers
- QGTC: accelerating quantized graph neural networks via GPU tensor coreYuke Wang, Boyuan Feng, Yufei DingPPoPP 2022 · 47 citations
- ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language ModelsChao Zeng, Songwei Liu, Yusheng Xie, Hong Liu et al.AAAI 2025 · 24 citations
- Efficient Design Space Exploration for Sparse Mixed Precision Neural ArchitecturesKrishna Teja Chitty-Venkata, Murali Emani, Venkatram Vishwanath, Arun K. SomaniHPDC 2022 · 6 citations
- HAWQ-V3: Dyadic Neural Network QuantizationZhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami et al.ICML 2021 · 240 citations
- Quant-LLM: Accelerating the Serving of Large Language Models via FP6-Centric Algorithm-System Co-Design on Modern GPUsHaojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen et al.USENIX ATC 2024 · 27 citations
