Opening book details…
Can I read Is Stochastic Gradient Descent Effective? A PDE Perspective on Machine Learning processes on EtoBox?
Is Stochastic Gradient Descent Effective? A PDE Perspective on Machine Learning processes by Barbieri, Davide; Bonforte, Matteo; Ibarrondo, Peio is a scholarly article available to read on EtoBox.
What is Is Stochastic Gradient Descent Effective? A PDE Perspective on Machine Learning processes about?
In this paper we analyze the behaviour of the stochastic gradient descent (SGD), a widely used method in supervised learning for optimizing neural network weights via a minimization of non-convex loss functions. Since the pioneering work of E, Li and Tai (2017), the underlying structure of such processes can be understood via parabolic PDEs of Fokker-Planck type, which are at the core of our analysis. Even if Fokker-Planck equations have a long history and a extensive literature, almost nothing is known when the potential is non-convex or when the diffusion matrix is degenerate, and this is the main difficulty that we face in our analysis. We identify two different regimes: in the initial phase of SGD, the loss function drives the weights to concentrate around the nearest local minimum. We refer to this phase as the drift regime and we provide quantitative estimates on this concentration phenomenon. Next, we introduce the diffusion regime, where stochastic fluctuations help the learning process to escape suboptimal local minima. We analyze the Mean Exit Time (MET) and prove upper and lower bounds of the MET. Finally, we address the asymptotic convergence of SGD, for a non-convex co
- Author
- Barbieri, Davide; Bonforte, Matteo; Ibarrondo, Peio
- Published
- 2025
- Language
- EN