{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/efficient-mixture-model-for-clustering-of","title":"Efficient mixture model for clustering of sparse high dimensional binary data","arxiv_id":"1707.03157","date":"2017-07-11","proceeding":null,"authors":["Marek Śmieja","Krzysztof Hajto","Jacek Tabor"],"abstract":"In this paper we propose a mixture model, SparseMix, for clustering of sparse\nhigh dimensional binary data, which connects model-based with centroid-based\nclustering. Every group is described by a representative and a probability\ndistribution modeling dispersion from this representative. In contrast to\nclassical mixture models based on EM algorithm, SparseMix:\n  -is especially designed for the processing of sparse data,\n  -can be efficiently realized by an on-line Hartigan optimization algorithm,\n  -is able to automatically reduce unnecessary clusters.\n  We perform extensive experimental studies on various types of data, which\nconfirm that SparseMix builds partitions with higher compatibility with\nreference grouping than related methods. Moreover, constructed representatives\noften better reveal the internal structure of data.","url_abs":"http://arxiv.org/abs/1707.03157v1","url_pdf":"http://arxiv.org/pdf/1707.03157v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"efficient-mixture-model-for-clustering-of","repo_url":"https://github.com/hajtos/SparseMIX","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}