{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generalization-in-adaptive-data-analysis-and","title":"Generalization in Adaptive Data Analysis and Holdout Reuse","arxiv_id":"1506.02629","date":"2015-06-08","proceeding":"NeurIPS 2015 12","authors":["Cynthia Dwork","Vitaly Feldman","Moritz Hardt","Toniann Pitassi","Omer Reingold","Aaron Roth"],"abstract":"Overfitting is the bane of data analysts, even when data are plentiful.\nFormal approaches to understanding this problem focus on statistical inference\nand generalization of individual analysis procedures. Yet the practice of data\nanalysis is an inherently interactive and adaptive process: new analyses and\nhypotheses are proposed after seeing the results of previous ones, parameters\nare tuned on the basis of obtained results, and datasets are shared and reused.\nAn investigation of this gap has recently been initiated by the authors in\n(Dwork et al., 2014), where we focused on the problem of estimating\nexpectations of adaptively chosen functions.\n  In this paper, we give a simple and practical method for reusing a holdout\n(or testing) set to validate the accuracy of hypotheses produced by a learning\nalgorithm operating on a training set. Reusing a holdout set adaptively\nmultiple times can easily lead to overfitting to the holdout set itself. We\ngive an algorithm that enables the validation of a large number of adaptively\nchosen hypotheses, while provably avoiding overfitting. We illustrate the\nadvantages of our algorithm over the standard use of the holdout set via a\nsimple synthetic experiment.\n  We also formalize and address the general problem of data reuse in adaptive\ndata analysis. We show how the differential-privacy based approach given in\n(Dwork et al., 2014) is applicable much more broadly to adaptive data analysis.\nWe then show that a simple approach based on description length can also be\nused to give guarantees of statistical validity in adaptive settings. Finally,\nwe demonstrate that these incomparable approaches can be unified via the notion\nof approximate max-information that we introduce.","url_abs":"http://arxiv.org/abs/1506.02629v2","url_pdf":"http://arxiv.org/pdf/1506.02629v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generalization-in-adaptive-data-analysis-and","repo_url":"https://github.com/DIDSR/ThresholdoutAUC","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"holdout-set","task_name":"Holdout Set"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1506.02629","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}