{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scalable-and-flexible-clustering-of-grouped","title":"Scalable and Flexible Clustering of Grouped Data via Parallel and Distributed Sampling in Versatile Hierarchical Dirichlet Processes","arxiv_id":null,"date":"2020-08-04","proceeding":"Uncertainty in Artificial Intelligence 2020 8","authors":["Dinari Or","Freifeld Oren"],"abstract":"Adaptive clustering of grouped data is often\r\ndone via the Hierarchical Dirichlet Process Mixture Model (HDPMM). That approach, however, is limited in its flexibility and usually does\r\nnot scale well. As a remedy, we propose another, but closely related, hierarchical Bayesian\r\nnonparametric framework. Our main contributions are as follows. 1) a new model, called the\r\nVersatile HDPMM (vHDPMM), with two possible settings: full and reduced. While the latter\r\nis akin to the HDPMM’s setting, the former\r\nsupports not only global features (as HDPMM\r\ndoes) but also local ones. 2) An effective mechanism for detecting global features. 3) A new\r\nsampler that addresses the challenges posed\r\nby the vHDPMM and, in the reduced setting,\r\nscales better than HDPMM samplers. 4) An\r\nefficient, distributed, and easily-modifiable implementation that offers more flexibility (even\r\nin the reduced setting) than publicly-available\r\nHDPMM implementations. Finally, we show\r\nthe utility of the approach in applications such\r\nas image cosegmentation, visual topic modeling, and clustering with missing data.","url_abs":"http://auai.org/uai2020/accepted.php#paper115","url_pdf":"http://auai.org/uai2020/proceedings/115_main_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scalable-and-flexible-clustering-of-grouped","repo_url":"https://github.com/BGU-CS-VIL/VersatileHDPMixtureModels","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"scalable-and-flexible-clustering-of-grouped","repo_url":"https://github.com/BGU-CS-VIL/VersatileHDPMixtureModels.jl","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}