{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/high-dimensional-cluster-analysis-with-the","title":"High-dimensional cluster analysis with the Masked EM Algorithm","arxiv_id":"1309.2848","date":"2013-09-11","proceeding":null,"authors":["Shabnam N. Kadir","Dan F. M. Goodman","Kenneth D. Harris"],"abstract":"Cluster analysis faces two problems in high dimensions: first, the `curse of\ndimensionality' that can lead to overfitting and poor generalization\nperformance; and second, the sheer time taken for conventional algorithms to\nprocess large amounts of high-dimensional data. In many applications, only a\nsmall subset of features provide information about the cluster membership of\nany one data point, however this informative feature subset may not be the same\nfor all data points. Here we introduce a `Masked EM' algorithm for fitting\nmixture of Gaussians models in such cases. We show that the algorithm performs\nclose to optimally on simulated Gaussian data, and in an application of `spike\nsorting' of high channel-count neuronal recordings.","url_abs":"http://arxiv.org/abs/1309.2848v1","url_pdf":"http://arxiv.org/pdf/1309.2848v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"high-dimensional-cluster-analysis-with-the","repo_url":"https://github.com/klusta-team/klustakwik","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"spike-sorting","task_name":"Spike Sorting"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}