Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:The effective noise of Stochastic Gradient Descent

Dec 20, 2021

Francesca Mignacco, Pierfrancesco Urbani

Figure 1 for The effective noise of Stochastic Gradient Descent

Figure 2 for The effective noise of Stochastic Gradient Descent

Figure 3 for The effective noise of Stochastic Gradient Descent

Figure 4 for The effective noise of Stochastic Gradient Descent

Share this with someone who'll enjoy it:

Abstract:Stochastic Gradient Descent (SGD) is the workhorse algorithm of deep learning technology. At each step of the training phase, a mini batch of samples is drawn from the training dataset and the weights of the neural network are adjusted according to the performance on this specific subset of examples. The mini-batch sampling procedure introduces a stochastic dynamics to the gradient descent, with a non-trivial state-dependent noise. We characterize the stochasticity of SGD and a recently-introduced variant, persistent SGD, in a prototypical neural network model. In the under-parametrized regime, where the final training error is positive, the SGD dynamics reaches a stationary state and we define an effective temperature from the fluctuation-dissipation theorem, computed from dynamical mean-field theory. We use the effective temperature to quantify the magnitude of the SGD noise as a function of the problem parameters. In the over-parametrized regime, where the training error vanishes, we measure the noise magnitude of SGD by computing the average distance between two replicas of the system with the same initialization and two different realizations of SGD noise. We find that the two noise measures behave similarly as a function of the problem parameters. Moreover, we observe that noisier algorithms lead to wider decision boundaries of the corresponding constraint satisfaction problem.

* 7 pages + appendix, 5 figures

View paper on

Share this with someone who'll enjoy it:

Title:The effective noise of Stochastic Gradient Descent

Paper and Code