{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/machine-learning-models-that-remember-too","title":"Machine Learning Models that Remember Too Much","arxiv_id":"1709.07886","date":"2017-09-22","proceeding":null,"authors":["Congzheng Song","Thomas Ristenpart","Vitaly Shmatikov"],"abstract":"Machine learning (ML) is becoming a commodity. Numerous ML frameworks and\nservices are available to data holders who are not ML experts but want to train\npredictive models on their data. It is important that ML models trained on\nsensitive inputs (e.g., personal images or documents) not leak too much\ninformation about the training data.\n  We consider a malicious ML provider who supplies model-training code to the\ndata holder, does not observe the training, but then obtains white- or\nblack-box access to the resulting model. In this setting, we design and\nimplement practical algorithms, some of them very similar to standard ML\ntechniques such as regularization and data augmentation, that \"memorize\"\ninformation about the training dataset in the model yet the model is as\naccurate and predictive as a conventionally trained model. We then explain how\nthe adversary can extract memorized information from the model.\n  We evaluate our techniques on standard ML tasks for image classification\n(CIFAR10), face recognition (LFW and FaceScrub), and text analysis (20\nNewsgroups and IMDB). In all cases, we show how our algorithms create models\nthat have high predictive power yet allow accurate extraction of subsets of\ntheir training data.","url_abs":"http://arxiv.org/abs/1709.07886v1","url_pdf":"http://arxiv.org/pdf/1709.07886v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"machine-learning-models-that-remember-too","repo_url":"https://github.com/csong27/ml-model-remember","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"face-recognition","task_name":"Face Recognition"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.07886","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}