Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks

May 26, 2022

Zhiwei Bai, Tao Luo, Zhi-Qin John Xu, Yaoyu Zhang

Figure 1 for Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks

Figure 2 for Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks

Figure 3 for Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks

Figure 4 for Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks

Share this with someone who'll enjoy it:

Abstract:Unraveling the general structure underlying the loss landscapes of deep neural networks (DNNs) is important for the theoretical study of deep learning. Inspired by the embedding principle of DNN loss landscape, we prove in this work an embedding principle in depth that loss landscape of an NN "contains" all critical points of the loss landscapes for shallower NNs. Specifically, we propose a critical lifting operator that any critical point of a shallower network can be lifted to a critical manifold of the target network while preserving the outputs. Through lifting, local minimum of an NN can become a strict saddle point of a deeper NN, which can be easily escaped by first-order methods. The embedding principle in depth reveals a large family of critical points in which layer linearization happens, i.e., computation of certain layers is effectively linear for the training inputs. We empirically demonstrate that, through suppressing layer linearization, batch normalization helps avoid the lifted critical manifolds, resulting in a faster decay of loss. We also demonstrate that increasing training data reduces the lifted critical manifold thus could accelerate the training. Overall, the embedding principle in depth well complements the embedding principle (in width), resulting in a complete characterization of the hierarchical structure of critical points/manifolds of a DNN loss landscape.

View paper on

Share this with someone who'll enjoy it:

Title:Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks

Paper and Code