{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-much-can-we-forget-about-data","title":"How Much Can We Forget about Data Contamination?","arxiv_id":"2410.03249","date":"2024-10-04","proceeding":null,"authors":["Sebastian Bordt","Suraj Srinivas","Valentyn Boreiko","Ulrike Von Luxburg"],"abstract":"The leakage of benchmark data into the training data has emerged as a significant challenge for evaluating the capabilities of large language models (LLMs). In this work, we challenge the common assumption that small-scale contamination renders benchmark evaluations invalid. First, we experimentally quantify the magnitude of benchmark overfitting based on scaling along three dimensions: The number of model parameters (up to 1.6B), the number of times an example is seen (up to 144), and the number of training tokens (up to 40B). If model and data follow the Chinchilla scaling laws, minor contamination indeed leads to overfitting. At the same time, even 144 times of contamination can be forgotten if the training data is scaled beyond five times Chinchilla, a regime characteristic of many modern LLMs. Continual pre-training of OLMo-7B corroborates these results. Next, we study the impact of the weight decay parameter on example forgetting, showing that empirical forgetting occurs faster than the cumulative weight decay. This allows us to gauge the degree of example forgetting in large-scale training runs, indicating that many LLMs, including Lllama 3 405B, have forgotten the data seen at the beginning of training.","url_abs":"https://arxiv.org/abs/2410.03249v3","url_pdf":"https://arxiv.org/pdf/2410.03249v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-much-can-we-forget-about-data","repo_url":"https://github.com/tml-tuebingen/forgetting-contamination","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[{"method_slug":"adamw","method_name":"AdamW"},{"method_slug":"chinchilla","method_name":"Chinchilla"},{"method_slug":"llama","method_name":"LLaMA"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.03249","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.03249"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tml-tuebingen/forgetting-contamination","reach":null}],"summary":{"ran_violates":1,"unverified":1},"by_repo_kind":{"official":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"ec3eea7a63ed06ec","entry":"has_subdirs","repo":"tml-tuebingen/forgetting-contamination","repo_kind":"official","path":"llm.c/create_contaminated_dataset.py","file_url":"https://github.com/tml-tuebingen/forgetting-contamination/blob/HEAD/llm.c/create_contaminated_dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ec3eea7a63ed06ec"}},{"code_sha256_prefix":"59549b8961f14328","entry":"is_instance_of","repo":"tml-tuebingen/forgetting-contamination","repo_kind":"official","path":"evaluation/evalutils.py","file_url":"https://github.com/tml-tuebingen/forgetting-contamination/blob/HEAD/evaluation/evalutils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"59549b8961f14328"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}