{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-stability-of-iterative-retraining-of","title":"On the Stability of Iterative Retraining of Generative Models on their own Data","arxiv_id":"2310.00429","date":"2023-09-30","proceeding":null,"authors":["Quentin Bertrand","Avishek Joey Bose","Alexandre Duplessis","Marco Jiralerspong","Gauthier Gidel"],"abstract":"Deep generative models have made tremendous progress in modeling complex data, often exhibiting generation quality that surpasses a typical human's ability to discern the authenticity of samples. Undeniably, a key driver of this success is enabled by the massive amounts of web-scale data consumed by these models. Due to these models' striking performance and ease of availability, the web will inevitably be increasingly populated with synthetic content. Such a fact directly implies that future iterations of generative models will be trained on both clean and artificially generated data from past models. In this paper, we develop a framework to rigorously study the impact of training generative models on mixed datasets -- from classical training on real data to self-consuming generative models trained on purely synthetic data. We first prove the stability of iterative training under the condition that the initial generative models approximate the data distribution well enough and the proportion of clean training data (w.r.t. synthetic data) is large enough. We empirically validate our theory on both synthetic and natural images by iteratively training normalizing flows and state-of-the-art diffusion models on CIFAR10 and FFHQ.","url_abs":"https://arxiv.org/abs/2310.00429v5","url_pdf":"https://arxiv.org/pdf/2310.00429v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-stability-of-iterative-retraining-of","repo_url":"https://github.com/qb3/gen_models_dont_go_mad","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"},{"method_slug":"normalizing-flows","method_name":"Normalizing Flows"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2310.00429","atlas_url":"https://app.syntology.ai/?focus=2310.00429","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.00429"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/qb3/gen_models_dont_go_mad","reach":{"status":"ok"}}],"summary":{"ran":1,"unverified":2},"by_repo_kind":{"official":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"7b2ae902263515a1","entry":"train","repo":"qb3/gen_models_dont_go_mad","repo_kind":"official","path":"iterative_retraining.py","file_url":"https://github.com/qb3/gen_models_dont_go_mad/blob/HEAD/iterative_retraining.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7b2ae902263515a1"}},{"code_sha256_prefix":"001b3e9ab12e8bb4","entry":"generate_ddpm","repo":"qb3/gen_models_dont_go_mad","repo_kind":"official","path":"iterative_retraining.py","file_url":"https://github.com/qb3/gen_models_dont_go_mad/blob/HEAD/iterative_retraining.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"001b3e9ab12e8bb4"}},{"code_sha256_prefix":"d475af2f9c6efc0c","entry":"train_ddpm","repo":"qb3/gen_models_dont_go_mad","repo_kind":"official","path":"iterative_retraining.py","file_url":"https://github.com/qb3/gen_models_dont_go_mad/blob/HEAD/iterative_retraining.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d475af2f9c6efc0c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}