Range-Invariant Approximation of Non-Linear Operations for Efficient BERT Fine-Tuning
Janghyeon Kim, Janghwan Lee, Jungwook Choi, JeongHo Han, Sangheon Lee
摘要
This paper proposes a range-invariant approximation of non-linear operations for training computations of Transformer-based large language models. The proposed method decomposes the approximation into the scaling and the range-invariant resolution for LUT approximation, covering diverse data ranges of non-linear operations with drastically reduced LUT entries during task-dependent BERT fine-tuning. We demonstrate that the proposed method robustly approximates all the non-linear operations of BERT without score degradation on challenging GLUE benchmarks using only a single-entry LUT, facilitating 52% area savings in hardware implementation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Genetic Quantization-Aware Approximation for Non-Linear Operations in TransformersPingcheng Dong, Yonghao Tan, Dong Zhang, Tianwei Ni 等DAC 2024 · 被引用 12 次
- NLI : Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs InferenceJiangyong Yu, Xiaomeng Han, Xing Hu, Chen Xu 等ICLR 2026 · 被引用 2 次
相关 Paper
- NN-LUT: neural approximation of non-linear operations for efficient transformer inferenceJoonsang Yu, Junki Park, Seongmin Park, Minsoo Kim 等DAC 2022 · 被引用 63 次
- MA-BERT: Towards Matrix Arithmetic-only BERT Inference by Eliminating Complex Non-Linear FunctionsNeo Wei Ming, Zhehui Wang, Cheng Liu, Rick Siow Mong Goh 等ICLR 2023
- I-BERT: Integer-only BERT QuantizationSehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney 等ICML 2021 · 被引用 439 次
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 被引用 126 次
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang 等EMNLP 2020 · 被引用 147 次
