{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-impact-of-random-models-on-clustering","title":"The Impact of Random Models on Clustering Similarity","arxiv_id":"1701.06508","date":"2017-01-23","proceeding":null,"authors":["Alexander J. Gates","Yong-Yeol Ahn"],"abstract":"Clustering is a central approach for unsupervised learning. After clustering\nis applied, the most fundamental analysis is to quantitatively compare\nclusterings. Such comparisons are crucial for the evaluation of clustering\nmethods as well as other tasks such as consensus clustering. It is often argued\nthat, in order to establish a baseline, clustering similarity should be\nassessed in the context of a random ensemble of clusterings. The prevailing\nassumption for the random clustering ensemble is the permutation model in which\nthe number and sizes of clusters are fixed. However, this assumption does not\nnecessarily hold in practice; for example, multiple runs of K-means clustering\nreturns clusterings with a fixed number of clusters, while the cluster size\ndistribution varies greatly. Here, we derive corrected variants of two\nclustering similarity measures (the Rand index and Mutual Information) in the\ncontext of two random clustering ensembles in which the number and sizes of\nclusters vary. In addition, we study the impact of one-sided comparisons in the\nscenario with a reference clustering. The consequences of different random\nmodels are illustrated using synthetic examples, handwriting recognition, and\ngene expression data. We demonstrate that the choice of random model can have a\ndrastic impact on the ranking of similar clustering pairs, and the evaluation\nof a clustering method with respect to a random baseline; thus, the choice of\nrandom clustering model should be carefully justified.","url_abs":"http://arxiv.org/abs/1701.06508v2","url_pdf":"http://arxiv.org/pdf/1701.06508v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-impact-of-random-models-on-clustering","repo_url":"https://github.com/ajgates42/clusim","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"clustering-ensemble","task_name":"Clustering Ensemble"},{"task_slug":"handwriting-recognition","task_name":"Handwriting Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1701.06508","atlas_url":"https://app.syntology.ai/?focus=1701.06508","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}