{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sparse-convex-clustering","title":"Sparse Convex Clustering","arxiv_id":"1601.04586","date":"2016-01-18","proceeding":null,"authors":["Binhuan Wang","Yilong Zhang","Will Wei Sun","Yixin Fang"],"abstract":"Convex clustering, a convex relaxation of k-means clustering and hierarchical\nclustering, has drawn recent attentions since it nicely addresses the\ninstability issue of traditional nonconvex clustering methods. Although its\ncomputational and statistical properties have been recently studied, the\nperformance of convex clustering has not yet been investigated in the\nhigh-dimensional clustering scenario, where the data contains a large number of\nfeatures and many of them carry no information about the clustering structure.\nIn this paper, we demonstrate that the performance of convex clustering could\nbe distorted when the uninformative features are included in the clustering. To\novercome it, we introduce a new clustering method, referred to as Sparse Convex\nClustering, to simultaneously cluster observations and conduct feature\nselection. The key idea is to formulate convex clustering in a form of\nregularization, with an adaptive group-lasso penalty term on cluster centers.\nIn order to optimally balance the tradeoff between the cluster fitting and\nsparsity, a tuning criterion based on clustering stability is developed. In\ntheory, we provide an unbiased estimator for the degrees of freedom of the\nproposed sparse convex clustering method. Finally, the effectiveness of the\nsparse convex clustering is examined through a variety of numerical experiments\nand a real data application.","url_abs":"http://arxiv.org/abs/1601.04586v4","url_pdf":"http://arxiv.org/pdf/1601.04586v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sparse-convex-clustering","repo_url":"https://github.com/elong0527/scvxclustr","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"feature-selection","task_name":"feature selection"}],"methods":[{"method_slug":"k-means-clustering","method_name":"k-Means Clustering"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}