Enabling fast uncertainty estimation: accelerating bayesian transformers via algorithmic and hardware optimizations
Hongxiang Fan, Martin Ferianc, Wayne Luk
摘要
Quantifying the uncertainty of neural networks (NNs) has been required by many safety-critical applications such as autonomous driving or medical diagnosis. Recently, Bayesian transformers have demonstrated their capabilities in providing high-quality uncertainty estimates paired with excellent accuracy. However, their real-time deployment is limited by the compute-intensive attention mechanism that is core to the transformer architecture, and the repeated Monte Carlo sampling to quantify the predictive uncertainty. To address these limitations, this paper accelerates Bayesian transformers via both algorithmic and hardware optimizations. On the algorithmic level, an evolutionary algorithm (EA)-based framework is proposed to exploit the sparsity in Bayesian transformers and ease their computational workload. On the hardware level, we demonstrate that the sparsity brings hardware performance improvement on our optimized CPU and GPU implementations. An adaptable hardware architecture is also proposed to accelerate Bayesian transformers on an FPGA. Extensive experiments demonstrate that the EA-based framework, together with hardware optimizations, reduce the latency of Bayesian transformers by up to 13, 12 and 20 times on CPU, GPU and FPGA platforms respectively, while achieving higher algorithmic performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 被引用 412 次
- Specifying Weight Priors in Bayesian Deep Neural Networks with Empirical BayesRanganath Krishnan, Mahesh Subedar, Omesh TickooAAAI 2020 · 被引用 65 次
- High-Performance FPGA-based Accelerator for Bayesian Neural NetworksHongxiang Fan, Martin Ferianc, Miguel Rodrigues, Hongyu Zhou 等DAC 2021 · 被引用 31 次
- Fast-BCNN: Massive Neuron Skipping in Bayesian Convolutional Neural NetworksQiyu Wan, Xin FuMICRO 2020 · 被引用 26 次
相关 Paper
- When Monte-Carlo Dropout Meets Multi-Exit: Optimizing Bayesian Neural Networks on FPGAHongxiang Fan, Mark Chen, Liam Castelli, Zhiqiang Que 等DAC 2023 · 被引用 5 次
- Calibrating Transformers via Sparse Gaussian ProcessesWenlong Chen, Yingzhen LiICLR 2023
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai 等ACL 2020 · 被引用 215 次
- Towards Uncertainty-aware Robotic Perception via Mixed-signal BNN Engine Leveraging Probabilistic Quantum TunnelingLikai Pei, Yu Zhou, Xingtian Wang, Xueji Zhao 等DAC 2025
- SIKA-GP: Accelerating Gaussian Process Inference with Sparse Inducing Kernel Approximations for Bayesian Deep LearningWenyuan Zhao, Rui Tuo, Chao TianICML 2026
