{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/selective-clustering-annotated-using-modes-of","title":"Selective Clustering Annotated using Modes of Projections","arxiv_id":"1807.10328","date":"2018-07-26","proceeding":null,"authors":["Evan Greene","Greg Finak","Raphael Gottardo"],"abstract":"Selective clustering annotated using modes of projections (SCAMP) is a new\nclustering algorithm for data in $\\mathbb{R}^p$. SCAMP is motivated from the\npoint of view of non-parametric mixture modeling. Rather than maximizing a\nclassification likelihood to determine cluster assignments, SCAMP casts\nclustering as a search and selection problem. One consequence of this problem\nformulation is that the number of clusters is $\\textbf{not}$ a SCAMP tuning\nparameter. The search phase of SCAMP consists of finding sub-collections of the\ndata matrix, called candidate clusters, that obey shape constraints along each\ncoordinate projection. An extension of the dip test of Hartigan and Hartigan\n(1985) is developed to assist the search. Selection occurs by scoring each\ncandidate cluster with a preference function that quantifies prior belief about\nthe mixture composition. Clustering proceeds by selecting candidates to\nmaximize their total preference score. SCAMP concludes by annotating each\nselected cluster with labels that describe how cluster-level statistics compare\nto certain dataset-level quantities. SCAMP can be run multiple times on a\nsingle data matrix. Comparison of annotations obtained across iterations\nprovides a measure of clustering uncertainty. Simulation studies and\napplications to real data are considered. A C++ implementation with R interface\nis $\\href{https://github.com/RGLab/scamp}{available\\ online}$.","url_abs":"http://arxiv.org/abs/1807.10328v1","url_pdf":"http://arxiv.org/pdf/1807.10328v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"selective-clustering-annotated-using-modes-of","repo_url":"https://github.com/RGLab/scamp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}