{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scaling-description-of-generalization-with","title":"Scaling description of generalization with number of parameters in deep learning","arxiv_id":"1901.01608","date":"2019-01-06","proceeding":null,"authors":["Mario Geiger","Arthur Jacot","Stefano Spigler","Franck Gabriel","Levent Sagun","Stéphane d'Ascoli","Giulio Biroli","Clément Hongler","Matthieu Wyart"],"abstract":"Supervised deep learning involves the training of neural networks with a large number $N$ of parameters. For large enough $N$, in the so-called over-parametrized regime, one can essentially fit the training data points. Sparsity-based arguments would suggest that the generalization error increases as $N$ grows past a certain threshold $N^{*}$. Instead, empirical studies have shown that in the over-parametrized regime, generalization error keeps decreasing with $N$. We resolve this paradox through a new framework. We rely on the so-called Neural Tangent Kernel, which connects large neural nets to kernel methods, to show that the initialization causes finite-size random fluctuations $\\|f_{N}-\\bar{f}_{N}\\|\\sim N^{-1/4}$ of the neural net output function $f_{N}$ around its expectation $\\bar{f}_{N}$. These affect the generalization error $\\epsilon_{N}$ for classification: under natural assumptions, it decays to a plateau value $\\epsilon_{\\infty}$ in a power-law fashion $\\sim N^{-1/2}$. This description breaks down at a so-called jamming transition $N=N^{*}$. At this threshold, we argue that $\\|f_{N}\\|$ diverges. This result leads to a plausible explanation for the cusp in test error known to occur at $N^{*}$. Our results are confirmed by extensive empirical observations on the MNIST and CIFAR image datasets. Our analysis finally suggests that, given a computational envelope, the smallest generalization error is obtained using several networks of intermediate sizes, just beyond $N^{*}$, and averaging their outputs.","url_abs":"https://arxiv.org/abs/1901.01608v5","url_pdf":"https://arxiv.org/pdf/1901.01608v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scaling-description-of-generalization-with","repo_url":"https://github.com/glouppe/info8010-deep-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.01608","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}