{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/asymptotic-analysis-of-two-layer-neural","title":"Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure","arxiv_id":"2503.00856","date":"2025-03-02","proceeding":null,"authors":["Samet Demir","Zafer Dogan"],"abstract":"In this work, we study the training and generalization performance of two-layer neural networks (NNs) after one gradient descent step under structured data modeled by Gaussian mixtures. While previous research has extensively analyzed this model under isotropic data assumption, such simplifications overlook the complexities inherent in real-world datasets. Our work addresses this limitation by analyzing two-layer NNs under Gaussian mixture data assumption in the asymptotically proportional limit, where the input dimension, number of hidden neurons, and sample size grow with finite ratios. We characterize the training and generalization errors by leveraging recent advancements in Gaussian universality. Specifically, we prove that a high-order polynomial model performs equivalent to the nonlinear neural networks under certain conditions. The degree of the equivalent model is intricately linked to both the \"data spread\" and the learning rate employed during one gradient step. Through extensive simulations, we demonstrate the equivalence between the original model and its polynomial counterpart across various regression and classification tasks. Additionally, we explore how different properties of Gaussian mixtures affect learning outcomes. Finally, we illustrate experimental results on Fashion-MNIST classification, indicating that our findings can translate to realistic data.","url_abs":"https://arxiv.org/abs/2503.00856v3","url_pdf":"https://arxiv.org/pdf/2503.00856v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"asymptotic-analysis-of-two-layer-neural","repo_url":"https://github.com/KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2503.00856","atlas_url":"https://app.syntology.ai/?focus=2503.00856","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.00856"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":1,"unverified":3},"by_repo_kind":{"official":{"samples":4,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"3a1c3d5bae9c495b","entry":"load_celebA","repo":"KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data","repo_kind":"official","path":"utils.py","file_url":"https://github.com/KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data/blob/HEAD/utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3a1c3d5bae9c495b"}},{"code_sha256_prefix":"17a13fda48fac95a","entry":"dataloader","repo":"KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data","repo_kind":"official","path":"dataloader.py","file_url":"https://github.com/KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data/blob/HEAD/dataloader.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"17a13fda48fac95a"}},{"code_sha256_prefix":"3ceb0ad81b0b7777","entry":"load_mnist","repo":"KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data","repo_kind":"official","path":"utils.py","file_url":"https://github.com/KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data/blob/HEAD/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3ceb0ad81b0b7777"}},{"code_sha256_prefix":"eb49535f1d6b6115","entry":"save_images","repo":"KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data","repo_kind":"official","path":"utils.py","file_url":"https://github.com/KU-MLIP/2-Layer-NNs-with-Gaussian-Mixtures-Data/blob/HEAD/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"eb49535f1d6b6115"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}