{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-interaction-effects-between-prediction","title":"On the Interaction Effects Between Prediction and Clustering","arxiv_id":"1807.06713","date":"2018-07-18","proceeding":null,"authors":["Matt Barnes","Artur Dubrawski"],"abstract":"Machine learning systems increasingly depend on pipelines of multiple\nalgorithms to provide high quality and well structured predictions. This paper\nargues interaction effects between clustering and prediction (e.g.\nclassification, regression) algorithms can cause subtle adverse behaviors\nduring cross-validation that may not be initially apparent. In particular, we\nfocus on the problem of estimating the out-of-cluster (OOC) prediction loss\ngiven an approximate clustering with probabilistic error rate $p_0$.\nTraditional cross-validation techniques exhibit significant empirical bias in\nthis setting, and the few attempts to estimate and correct for these effects\nare intractable on larger datasets. Further, no previous work has been able to\ncharacterize the conditions under which these empirical effects occur, and if\nthey do, what properties they have. We precisely answer these questions by\nproviding theoretical properties which hold in various settings, and prove that\nexpected out-of-cluster loss behavior rapidly decays with even minor clustering\nerrors. Fortunately, we are able to leverage these same properties to construct\nhypothesis tests and scalable estimators necessary for correcting the problem.\nEmpirical results on benchmark datasets validate our theoretical results and\ndemonstrate how scaling techniques provide solutions to new classes of\nproblems.","url_abs":"http://arxiv.org/abs/1807.06713v2","url_pdf":"http://arxiv.org/pdf/1807.06713v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-interaction-effects-between-prediction","repo_url":"https://github.com/mbarnes1/B3","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"prediction","task_name":"Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}