{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bayesian-cluster-enumeration-criterion-for","title":"Bayesian Cluster Enumeration Criterion for Unsupervised Learning","arxiv_id":"1710.07954","date":"2017-10-22","proceeding":null,"authors":["Freweyni K. Teklehaymanot","Michael Muma","Abdelhak M. Zoubir"],"abstract":"We derive a new Bayesian Information Criterion (BIC) by formulating the\nproblem of estimating the number of clusters in an observed data set as\nmaximization of the posterior probability of the candidate models. Given that\nsome mild assumptions are satisfied, we provide a general BIC expression for a\nbroad class of data distributions. This serves as a starting point when\nderiving the BIC for specific distributions. Along this line, we provide a\nclosed-form BIC expression for multivariate Gaussian distributed variables. We\nshow that incorporating the data structure of the clustering problem into the\nderivation of the BIC results in an expression whose penalty term is different\nfrom that of the original BIC. We propose a two-step cluster enumeration\nalgorithm. First, a model-based unsupervised learning algorithm partitions the\ndata according to a given set of candidate models. Subsequently, the number of\nclusters is determined as the one associated with the model for which the\nproposed BIC is maximal. The performance of the proposed two-step algorithm is\ntested using synthetic and real data sets.","url_abs":"http://arxiv.org/abs/1710.07954v3","url_pdf":"http://arxiv.org/pdf/1710.07954v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bayesian-cluster-enumeration-criterion-for","repo_url":"https://github.com/FreTekle/Bayesian-Cluster-Enumeration","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}