Revisiting Structured Dropout

Yiren Zhao; Oluwatomisin Dada; Robert Mullins; Xitong Gao

2023 ACML ACML 2023

Revisiting Structured Dropout

Abstract

Large neural networks are often overparameterised and prone to overfitting, Dropout is a widely used regularization technique to combat overfitting and improve model generalization. However, unstructured Dropout is not always effective for specific network architectures and this has led to the formation of multiple structured Dropout approaches to improve model performance and, sometimes, reduce the computational resources required for inference. In this work, we revisit structured Dropout comparing different Dropout approaches on natural language processing and computer vision tasks for multiple state-of-the-art networks. Additionally, we devise an approach to structured Dropout we call \textbf{\emph{ProbDropBlock}} which drops contiguous blocks from feature maps with a probability given by the normalized feature salience values. We find that, with a simple scheduling strategy, the proposed approach to structured Dropout consistently improves model performance compared to baselines and other Dropout approaches on a diverse range of tasks and models. In particular, we show \textbf{\emph{ProbDropBlock}} improves RoBERTa finetuning on MNLI by $0.22%$, and training of ResNet50 on ImageNet by $0.28%$.

🌉 Interdisciplinary Bridge — Deep Learning and Machine Learning

🧭 Keyword Pioneer — feature salience

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Speech & Audio

Authors

Yiren Zhao , Oluwatomisin Dada , Robert Mullins , Xitong Gao

Topics

Machine Learning > Optimization & Theory > Neural Network Optimization Deep Learning > Techniques > Normalization

Keywords

neural network regularization model generalization structured dropout feature salience

Download PDF

Related papers

How GAN Generators can Invert Networks in Real-Time 2023

ProtoDiffusion: Classifier-Free Diffusion Guidance with Prototype Learning 2023

BarlowRL: Barlow Twins for Data-Efficient Reinforcement Learning 2023

Enhancing Cross-Category Learning in Recommendation Systems with Multi-Layer Embedding Training 2023

Deep Representation Learning for Prediction of Temporal Event Sets in the Continuous Time Domain 2023