{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/model-assisted-variable-clustering-minimax","title":"Model Assisted Variable Clustering: Minimax-optimal Recovery and Algorithms","arxiv_id":"1508.01939","date":"2015-08-08","proceeding":null,"authors":["Florentina Bunea","Christophe Giraud","Xi Luo","Martin Royer","Nicolas Verzelen"],"abstract":"Model-based clustering defines population level clusters relative to a model\nthat embeds notions of similarity. Algorithms tailored to such models yield\nestimated clusters with a clear statistical interpretation. We take this view\nhere and introduce the class of G-block covariance models as a background model\nfor variable clustering. In such models, two variables in a cluster are deemed\nsimilar if they have similar associations will all other variables. This can\narise, for instance, when groups of variables are noise corrupted versions of\nthe same latent factor. We quantify the difficulty of clustering data generated\nfrom a G-block covariance model in terms of cluster proximity, measured with\nrespect to two related, but different, cluster separation metrics. We derive\nminimax cluster separation thresholds, which are the metric values below which\nno algorithm can recover the model-defined clusters exactly, and show that they\nare different for the two metrics. We therefore develop two algorithms, COD and\nPECOK, tailored to G-block covariance models, and study their\nminimax-optimality with respect to each metric. Of independent interest is the\nfact that the analysis of the PECOK algorithm, which is based on a corrected\nconvex relaxation of the popular K-means algorithm, provides the first\nstatistical analysis of such algorithms for variable clustering. Additionally,\nwe contrast our methods with another popular clustering method, spectral\nclustering, specialized to variable clustering, and show that ensuring exact\ncluster recovery via this method requires clusters to have a higher separation,\nrelative to the minimax threshold. Extensive simulation studies, as well as our\ndata analyses, confirm the applicability of our approach.","url_abs":"http://arxiv.org/abs/1508.01939v5","url_pdf":"http://arxiv.org/pdf/1508.01939v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"model-assisted-variable-clustering-minimax","repo_url":"https://github.com/martinroyer/pecok","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}