{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/an-empirical-analysis-of-forgetting-in-pre","title":"An Empirical Analysis of Forgetting in Pre-trained Models with Incremental Low-Rank Updates","arxiv_id":"2405.18069","date":"2024-05-28","proceeding":null,"authors":["Albin Soutif--Cormerais","Simone Magistri","Joost Van de Weijer","Andew D. Bagdanov"],"abstract":"Broad, open source availability of large pretrained foundation models on the internet through platforms such as HuggingFace has taken the world of practical deep learning by storm. A classical pipeline for neural network training now typically consists of finetuning these pretrained network on a small target dataset instead of training from scratch. In the case of large models this can be done even on modest hardware using a low rank training technique known as Low-Rank Adaptation (LoRA). While Low Rank training has already been studied in the continual learning setting, existing works often consider storing the learned adapter along with the existing model but rarely attempt to modify the weights of the pretrained model by merging the LoRA with the existing weights after finishing the training of each task. In this article we investigate this setting and study the impact of LoRA rank on the forgetting of the pretraining foundation task and on the plasticity and forgetting of subsequent ones. We observe that this rank has an important impact on forgetting of both the pretraining and downstream tasks. We also observe that vision transformers finetuned in that way exhibit a sort of ``contextual'' forgetting, a behaviour that we do not observe for residual networks and that we believe has not been observed yet in previous continual learning works.","url_abs":"https://arxiv.org/abs/2405.18069v2","url_pdf":"https://arxiv.org/pdf/2405.18069v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"an-empirical-analysis-of-forgetting-in-pre","repo_url":"https://github.com/AlbinSou/lora_cl_analysis","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"continual-learning","task_name":"Continual Learning"}],"methods":[{"method_slug":"adapter","method_name":"Adapter"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2405.18069","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2405.18069"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/AlbinSou/lora_cl_analysis","reach":{"status":"ok"}}],"summary":{"ran":3,"unverified":2},"by_repo_kind":{"official":{"samples":5,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"3d960338e27ef18d","entry":"create_lora_config","repo":"AlbinSou/lora_cl_analysis","repo_kind":"official","path":"experiments/lora_forget.py","file_url":"https://github.com/AlbinSou/lora_cl_analysis/blob/HEAD/experiments/lora_forget.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3d960338e27ef18d"}},{"code_sha256_prefix":"20ecdb599baec0a6","entry":"extract_kwargs","repo":"AlbinSou/lora_cl_analysis","repo_kind":"official","path":"src/toolkit/utils.py","file_url":"https://github.com/AlbinSou/lora_cl_analysis/blob/HEAD/src/toolkit/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"20ecdb599baec0a6"}},{"code_sha256_prefix":"a89041eb0b09cf89","entry":"process_cars","repo":"AlbinSou/lora_cl_analysis","repo_kind":"official","path":"src/factories/benchmark_factory.py","file_url":"https://github.com/AlbinSou/lora_cl_analysis/blob/HEAD/src/factories/benchmark_factory.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a89041eb0b09cf89"}},{"code_sha256_prefix":"1f0b9095eeb0bd7b","entry":"convert_pil","repo":"AlbinSou/lora_cl_analysis","repo_kind":"official","path":"src/factories/flowers.py","file_url":"https://github.com/AlbinSou/lora_cl_analysis/blob/HEAD/src/factories/flowers.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1f0b9095eeb0bd7b"}},{"code_sha256_prefix":"f1ef2ea33696e018","entry":"create_default_args","repo":"AlbinSou/lora_cl_analysis","repo_kind":"official","path":"src/toolkit/utils.py","file_url":"https://github.com/AlbinSou/lora_cl_analysis/blob/HEAD/src/toolkit/utils.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f1ef2ea33696e018"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}