Lune

DAC2023Top-tier venue

When Monte-Carlo Dropout Meets Multi-Exit: Optimizing Bayesian Neural Networks on FPGA

Hongxiang Fan, Mark Chen, Liam Castelli, Zhiqiang Que, He Li, Kenneth Long, Wayne Luk

2023Year
5Citations
1Top-tier citations

Abstract

Bayesian Neural Networks (BayesNNs) have demonstrated their capability of providing calibrated prediction for safety-critical applications such as medical imaging and autonomous driving. However, the high algorithmic complexity and the poor hardware performance of BayesNNs hinder their deployment in real-life applications. To bridge this gap, this paper proposes a novel multi-exit Monte-Carlo Dropout (MCD)-based BayesNN that achieves well-calibrated predictions with low algorithmic complexity. To further reduce the barrier to adopting BayesNNs, we propose a transformation framework that can generate FPGA-based accelerators for multi-exit MCD-based BayesNNs. Several novel optimization techniques are introduced to improve hardware performance. Our experiments demonstrate that our auto-generated accelerator achieves higher energy efficiency than CPU, GPU, and other state-of-the-art hardware implementations. Our code is publicly available at: https://github.com/os-hxfan/BayesNN FPGA.git

• A novel multi-exit MCD-based BayesNN with better calibration ability than conventional MCD-based BayesNN, and higher computational efficiency and flexibility over traditional deep ensembles.

• A design framework for transforming non-BayesNN models to multi-exit BayesNN hardware accelerators with high hardware performance and energy efficiency.

• Various optimization strategies including spatial-temporal mapping and algorithm-hardware co-exploration for performance improvement.

A. Bayesian Neural Networks

BayesNNs are able to achieve robustness against overfitting and to provide the estimation of their model uncertainty by means of Bayesian inference. Instead of capturing point-wise weight values like non-BayesNNs, BayesNNs are trained to learn the distribution of the weights. The Bayes rule is adopted in learning the distribution p(w|D) for the weights w with respect to training data D. It is, however, computationally intractable to calculate the posterior

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 8b2d93a3-bbe9-44c2-8321-089d23a83c97

Cited by top-tier papers1

Ask how each one uses it

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines