The Mathematical Theory of Neural Network-based Machine Learning

-
Weinan E, Princeton University

The task of supervised learning is to approximate a function using a given set ofÌýdata. In low dimensions, its mathematical theory has been established in classicalÌýnumerical analysis and approximation theory in which the function spaces ofÌýinterest (the Sobolev or Besov spaces), the order of the error and the convergenceÌýrate of the gradient-based algorithms are all well-understood.ÌýDirect extension of such a theory to high dimensionsÌýleads to estimates that suffer from theÌýcurse of dimensionality as well as degeneracy in the over-parametrized regime.

In this talk, we attempt to put forward a unified mathematical framework forÌýanalyzing neural network-based machine learning in high dimension (and theÌýover-parametrized regime). We illustrate this framework using kernel methods,Ìýshallow network models and deep network models. For each of these methods, we identify the right function spaces (for whichÌýthe optimal complexity estimates andÌýdirect and inverse approximation theorems hold), prove optimal a priori generalizationÌýerror estimates and study the behavior of gradient decent dynamics.

The talk is based mostly on joint work with Chao Ma, Lei Wu as well as Qingcan Wang.
Ìý