Loading...

Over-parameterized Neural Networks: Convergence Analysis and Generalization Bounds

Tinati, Mohammad | 2022

1387 Viewed
  1. Type of Document: M.Sc. Thesis
  2. Language: Farsi
  3. Document No: 55156 (05)
  4. University: Sharif University of Technolog
  5. Department: Electrical Engineering
  6. Advisor(s): Maddah Ali, Mohammad Ali; Motahari, Abolfazl
  7. Abstract:
  8. Despite its extraordinary empirical achievements, the theoretical foundation of modern Machine Learning, and in particular deep neural networks (DNN), is still a mystery. In this thesis, we have studied the effect of optimization algorithms on the generalization properties for shallow neural networks. Particularly, we have focused on the implicit biases these optimization procedures, specifically dropout, deal with. As an example for this implicit bias, classical results had shown that for linear regression, in the interpolation regime, gradient descent, among all the possible solutions, converges to the minimum L2-norm interpolation. Due to the complex nature of the neural networks optimization problem, such results require a more complicated analysis. Our first attempt to understand this implicit bias was to investigate the dynamics in the so-called “lazy” regime, where the neurons’ weights stay in the close vicinity of the initialization. Inspired by the recent line of work on the Neural Tangent Kernel (NTK), we took a functional analysis approach and tracked the effect of dropout on the output function. In this lazy regime the output function behaves similarly to the linear approximation of the function at the initialization—where the data is represented by a fixed feature space and its analysis is much simpler. In this thesis, first, we derived the partial differential equation that governs the evolution of parameters in the presence and absence of dropout. Then, by comparing the two, we showed how dropout technique effectively behaves as an implicit regularization for the loss function in the lazy regime.Having investigated the dynamics in the lazy regime, we learned that this regime might not be able to capture the representation learning of neural networks—the process by which layers extract the important features of the dataset. Thus, we have investigated the effect of dropout on two-layer networks in the mean-field regime where the neurons can move far away from the initialization point and make representation learning possible. Inspired by the Jordan-Kinderlherer-Otto scheme, our approach is to analyze dropout effect on the space of measures that represent shallow networks, while their topology is defined by the Wasserstein metric. In particular, we have studied the dynamics of noisy stochastic gradient descent in the presence of dropout.
  9. Keywords:
  10. Nonconvex Optimization ; Mean-Field Theory ; Generalization Error ; Over-Parameterized Neural Networks ; Neural Tangent Kernel

 Digital Object List

 Bookmark

...see more