Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Beyond BatchNorm: Towards a General Understanding of Normalization in Deep Learning

Jul 08, 2021

Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka

Figure 1 for Beyond BatchNorm: Towards a General Understanding of Normalization in Deep Learning

Figure 2 for Beyond BatchNorm: Towards a General Understanding of Normalization in Deep Learning

Figure 3 for Beyond BatchNorm: Towards a General Understanding of Normalization in Deep Learning

Figure 4 for Beyond BatchNorm: Towards a General Understanding of Normalization in Deep Learning

Share this with someone who'll enjoy it:

Abstract:Inspired by BatchNorm, there has been an explosion of normalization layers for deep neural networks (DNNs). However, these alternative normalization layers have seen minimal use, partially due to a lack of guiding principles that can help identify when these layers can serve as a replacement for BatchNorm. To address this problem, we take a theoretical approach, generalizing the known beneficial mechanisms of BatchNorm to several recently proposed normalization techniques. Our generalized theory leads to the following set of principles: (i) similar to BatchNorm, activations-based normalization layers can prevent exponential growth of activations in ResNets, but parametric layers require explicit remedies; (ii) use of GroupNorm can ensure informative forward propagation, with different samples being assigned dissimilar activations, but increasing group size results in increasingly indistinguishable activations for different samples, explaining slow convergence speed in models with LayerNorm; (iii) small group sizes result in large gradient norm in earlier layers, hence explaining training instability issues in Instance Normalization and illustrating a speed-stability tradeoff in GroupNorm. Overall, our analysis reveals a unified set of mechanisms that underpin the success of normalization methods in deep learning, providing us with a compass to systematically explore the vast design space of DNN normalization layers.

View paper on

Share this with someone who'll enjoy it:

Title:Beyond BatchNorm: Towards a General Understanding of Normalization in Deep Learning

Paper and Code