{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-neural-networks-extrapolate-from","title":"How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks","arxiv_id":"2009.11848","date":"2020-09-24","proceeding":"ICLR 2021 1","authors":["Keyulu Xu","Mozhi Zhang","Jingling Li","Simon S. Du","Ken-ichi Kawarabayashi","Stefanie Jegelka"],"abstract":"We study how neural networks trained by gradient descent extrapolate, i.e., what they learn outside the support of the training distribution. Previous works report mixed empirical results when extrapolating with neural networks: while feedforward neural networks, a.k.a. multilayer perceptrons (MLPs), do not extrapolate well in certain simple tasks, Graph Neural Networks (GNNs) -- structured networks with MLP modules -- have shown some success in more complex tasks. Working towards a theoretical explanation, we identify conditions under which MLPs and GNNs extrapolate well. First, we quantify the observation that ReLU MLPs quickly converge to linear functions along any direction from the origin, which implies that ReLU MLPs do not extrapolate most nonlinear functions. But, they can provably learn a linear target function when the training distribution is sufficiently \"diverse\". Second, in connection to analyzing the successes and limitations of GNNs, these results suggest a hypothesis for which we provide theoretical and empirical evidence: the success of GNNs in extrapolating algorithmic tasks to new data (e.g., larger graphs or edge weights) relies on encoding task-specific non-linearities in the architecture or features. Our theoretical analysis builds on a connection of over-parameterized networks to the neural tangent kernel. Empirically, our theory holds across different training settings.","url_abs":"https://arxiv.org/abs/2009.11848v5","url_pdf":"https://arxiv.org/pdf/2009.11848v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-neural-networks-extrapolate-from","repo_url":"https://github.com/jinglingli/nn-extrapolate/blob/master/README.md","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"how-neural-networks-extrapolate-from","repo_url":"https://github.com/Xiaohui9607/RFF_pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"how-neural-networks-extrapolate-from","repo_url":"https://github.com/jinglingli/nn-extrapolate","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[{"method_slug":"relu","method_name":"ReLU"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2009.11848","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2009.11848"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jinglingli/nn-extrapolate","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jinglingli/nn-extrapolate/blob/master/README.md","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Xiaohui9607/RFF_pytorch","reach":{"status":"ok"}}],"summary":{"ran_honours":1,"ran_violates":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"e570468b6f81a89c","entry":"mape","repo":"jinglingli/nn-extrapolate","repo_kind":"official","path":"feedforward/MLPs.py","file_url":"https://github.com/jinglingli/nn-extrapolate/blob/HEAD/feedforward/MLPs.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e570468b6f81a89c"}},{"code_sha256_prefix":"83be324d300a4045","entry":"square","repo":"jinglingli/nn-extrapolate","repo_kind":"official","path":"feedforward/MLPs.py","file_url":"https://github.com/jinglingli/nn-extrapolate/blob/HEAD/feedforward/MLPs.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"83be324d300a4045"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}