{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fast-discrete-distribution-clustering-using","title":"Fast Discrete Distribution Clustering Using Wasserstein Barycenter with Sparse Support","arxiv_id":"1510.00012","date":"2015-09-30","proceeding":null,"authors":["Jianbo Ye","Panruo Wu","James Z. Wang","Jia Li"],"abstract":"In a variety of research areas, the weighted bag of vectors and the histogram\nare widely used descriptors for complex objects. Both can be expressed as\ndiscrete distributions. D2-clustering pursues the minimum total within-cluster\nvariation for a set of discrete distributions subject to the\nKantorovich-Wasserstein metric. D2-clustering has a severe scalability issue,\nthe bottleneck being the computation of a centroid distribution, called\nWasserstein barycenter, that minimizes its sum of squared distances to the\ncluster members. In this paper, we develop a modified Bregman ADMM approach for\ncomputing the approximate discrete Wasserstein barycenter of large clusters. In\nthe case when the support points of the barycenters are unknown and have low\ncardinality, our method achieves high accuracy empirically at a much reduced\ncomputational cost. The strengths and weaknesses of our method and its\nalternatives are examined through experiments, and we recommend scenarios for\ntheir respective usage. Moreover, we develop both serial and parallelized\nversions of the algorithm. By experimenting with large-scale data, we\ndemonstrate the computational efficiency of the new methods and investigate\ntheir convergence properties and numerical stability. The clustering results\nobtained on several datasets in different domains are highly competitive in\ncomparison with some widely used methods in the corresponding areas.","url_abs":"http://arxiv.org/abs/1510.00012v4","url_pdf":"http://arxiv.org/pdf/1510.00012v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fast-discrete-distribution-clustering-using","repo_url":"https://github.com/bobye/d2_kmeans","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"fast-discrete-distribution-clustering-using","repo_url":"https://github.com/bobye/WBC_Matlab","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"}],"methods":[{"method_slug":"admm","method_name":"ADMM"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1510.00012","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}