Training Binary Neural Networks using the Bayesian Learning Rule

Xiangming Meng; Roman Bachmann; Mohammad Emtiyaz Khan

2020 ICML ICML 2020

Training Binary Neural Networks using the Bayesian Learning Rule

Abstract

Neural networks with binary weights are computation-efficient and hardware-friendly, but their training is challenging because it involves a discrete optimization problem. Surprisingly, ignoring the discrete nature of the problem and using gradient-based methods, such as the Straight-Through Estimator, still works well in practice. This raises the question: are there principled approaches which justify such methods? In this paper, we propose such an approach using the Bayesian learning rule. The rule, when applied to estimate a Bernoulli distribution over the binary weights, results in an algorithm which justifies some of the algorithmic choices made by the previous approaches. The algorithm not only obtains state-of-the-art performance, but also enables uncertainty estimation and continual learning to avoid catastrophic forgetting. Our work provides a principled approach for training binary neural networks which also justifies and extends existing approaches.

🌉 Interdisciplinary Bridge — Deep Learning and Machine Learning

🧭 Keyword Pioneer — bernoulli distribution

🐣 Hot Topic Early Bird — continual learning

🐝 Cross-Pollinator — Artificial Intelligence, Computer Vision, Data Science & Analytics, Deep Learning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Speech & Audio

Authors

Xiangming Meng , Roman Bachmann , Mohammad Emtiyaz Khan

Topics

Machine Learning > Optimization & Theory > Bayesian Inference Deep Learning > Architectures > Neural Networks Machine Learning > Application Areas > Model Compression Machine Learning > Bayesian & Probabilistic > Bayesian Inference Deep Learning > Optimization & Theory > Neural Network Optimization Deep Learning > Optimization & Theory > Optimization Machine Learning > Learning Types > Uncertainty Quantification

Keywords

model compression continual learning variational inference uncertainty estimation weight quantization binary neural network bayesian learning rule bernoulli distribution

Download PDF

Related papers

Correlation Clustering with Asymmetric Classification Errors 2020

Learning Portable Representations for High-Level Planning 2020

Proving the Lottery Ticket Hypothesis: Pruning is All You Need 2020

Minimax Pareto Fairness: A Multi Objective Perspective 2020

DeepMatch: Balancing Deep Covariate Representations for Causal Inference Using Adversarial Training 2020