{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/stable-architectures-for-deep-neural-networks","title":"Stable Architectures for Deep Neural Networks","arxiv_id":"1705.03341","date":"2017-05-09","proceeding":null,"authors":["Eldad Haber","Lars Ruthotto"],"abstract":"Deep neural networks have become invaluable tools for supervised machine\nlearning, e.g., classification of text or images. While often offering superior\nresults over traditional techniques and successfully expressing complicated\npatterns in data, deep architectures are known to be challenging to design and\ntrain such that they generalize well to new data. Important issues with deep\narchitectures are numerical instabilities in derivative-based learning\nalgorithms commonly called exploding or vanishing gradients. In this paper we\npropose new forward propagation techniques inspired by systems of Ordinary\nDifferential Equations (ODE) that overcome this challenge and lead to\nwell-posed learning problems for arbitrarily deep networks.\n  The backbone of our approach is our interpretation of deep learning as a\nparameter estimation problem of nonlinear dynamical systems. Given this\nformulation, we analyze stability and well-posedness of deep learning and use\nthis new understanding to develop new network architectures. We relate the\nexploding and vanishing gradient phenomenon to the stability of the discrete\nODE and present several strategies for stabilizing deep learning for very deep\nnetworks. While our new architectures restrict the solution space, several\nnumerical experiments show their competitiveness with state-of-the-art\nnetworks.","url_abs":"http://arxiv.org/abs/1705.03341v3","url_pdf":"http://arxiv.org/pdf/1705.03341v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"stable-architectures-for-deep-neural-networks","repo_url":"https://github.com/ClaraGalimberti/HamiltonianNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"stable-architectures-for-deep-neural-networks","repo_url":"https://github.com/DecodEPFL/HamiltonianNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"stable-architectures-for-deep-neural-networks","repo_url":"https://github.com/TheoryDev/Deep-neural-network-training-optimisation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"stable-architectures-for-deep-neural-networks","repo_url":"https://github.com/xtractopen/meganet.m","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"stable-architectures-for-deep-neural-networks","repo_url":"https://github.com/MindCode-4/code-13/tree/main/stable-sam","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"parameter-estimation","task_name":"parameter estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1705.03341","atlas_url":"https://app.syntology.ai/?focus=1705.03341","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1705.03341"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/TheoryDev/Deep-neural-network-training-optimisation","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ClaraGalimberti/HamiltonianNet","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-13/tree/main/stable-sam","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/xtractopen/meganet.m","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/DecodEPFL/HamiltonianNet","reach":null}],"summary":{"ran_honours":1},"by_repo_kind":{"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"b72ce5c71654e88c","entry":"test","repo":"DecodEPFL/HamiltonianNet","repo_kind":"listed","path":"examples/run_MNIST.py","file_url":"https://github.com/DecodEPFL/HamiltonianNet/blob/HEAD/examples/run_MNIST.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"b72ce5c71654e88c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}