Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:On the impact of activation and normalization in obtaining isometric embeddings at initialization

May 28, 2023

Amir Joudaki, Hadi Daneshmand, Francis Bach

Figure 1 for On the impact of activation and normalization in obtaining isometric embeddings at initialization

Figure 2 for On the impact of activation and normalization in obtaining isometric embeddings at initialization

Figure 3 for On the impact of activation and normalization in obtaining isometric embeddings at initialization

Figure 4 for On the impact of activation and normalization in obtaining isometric embeddings at initialization

Share this with someone who'll enjoy it:

Abstract:In this paper, we explore the structure of the penultimate Gram matrix in deep neural networks, which contains the pairwise inner products of outputs corresponding to a batch of inputs. In several architectures it has been observed that this Gram matrix becomes degenerate with depth at initialization, which dramatically slows training. Normalization layers, such as batch or layer normalization, play a pivotal role in preventing the rank collapse issue. Despite promising advances, the existing theoretical results (i) do not extend to layer normalization, which is widely used in transformers, (ii) can not characterize the bias of normalization quantitatively at finite depth. To bridge this gap, we provide a proof that layer normalization, in conjunction with activation layers, biases the Gram matrix of a multilayer perceptron towards isometry at an exponential rate with depth at initialization. We quantify this rate using the Hermite expansion of the activation function, highlighting the importance of higher order ($\ge 2$) Hermite coefficients in the bias towards isometry.

View paper on

Share this with someone who'll enjoy it:

Title:On the impact of activation and normalization in obtaining isometric embeddings at initialization

Paper and Code