{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-diminishing-returns-of-width-for","title":"On the Diminishing Returns of Width for Continual Learning","arxiv_id":"2403.06398","date":"2024-03-11","proceeding":null,"authors":["Etash Guha","Vihan Lakshman"],"abstract":"While deep neural networks have demonstrated groundbreaking performance in various settings, these models often suffer from \\emph{catastrophic forgetting} when trained on new tasks in sequence. Several works have empirically demonstrated that increasing the width of a neural network leads to a decrease in catastrophic forgetting but have yet to characterize the exact relationship between width and continual learning. We design one of the first frameworks to analyze Continual Learning Theory and prove that width is directly related to forgetting in Feed-Forward Networks (FFN). Specifically, we demonstrate that increasing network widths to reduce forgetting yields diminishing returns. We empirically verify our claims at widths hitherto unexplored in prior studies where the diminishing returns are clearly observed as predicted by our theory.","url_abs":"https://arxiv.org/abs/2403.06398v3","url_pdf":"https://arxiv.org/pdf/2403.06398v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-diminishing-returns-of-width-for","repo_url":"https://github.com/vihan-lakshman/diminishing-returns-wide-continual-learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"continual-learning","task_name":"Continual Learning"},{"task_slug":"learning-theory","task_name":"Learning Theory"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2403.06398","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2403.06398"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vihan-lakshman/diminishing-returns-wide-continual-learning","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"efdce8167e519f21","entry":"get_rotated_gtsrb","repo":"vihan-lakshman/diminishing-returns-wide-continual-learning","repo_kind":"official","path":"data_utils.py","file_url":"https://github.com/vihan-lakshman/diminishing-returns-wide-continual-learning/blob/HEAD/data_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"efdce8167e519f21"}},{"code_sha256_prefix":"6226df0913be22a5","entry":"get_rotated_mnist","repo":"vihan-lakshman/diminishing-returns-wide-continual-learning","repo_kind":"official","path":"data_utils.py","file_url":"https://github.com/vihan-lakshman/diminishing-returns-wide-continual-learning/blob/HEAD/data_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6226df0913be22a5"}},{"code_sha256_prefix":"e0d56369bb1e68c8","entry":"get_rotated_svhn","repo":"vihan-lakshman/diminishing-returns-wide-continual-learning","repo_kind":"official","path":"data_utils.py","file_url":"https://github.com/vihan-lakshman/diminishing-returns-wide-continual-learning/blob/HEAD/data_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e0d56369bb1e68c8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}