{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sparse-partially-collapsed-mcmc-for-parallel","title":"Sparse Partially Collapsed MCMC for Parallel Inference in Topic Models","arxiv_id":"1506.03784","date":"2015-06-11","proceeding":null,"authors":["Måns Magnusson","Leif Jonsson","Mattias Villani","David Broman"],"abstract":"Topic models, and more specifically the class of Latent Dirichlet Allocation\n(LDA), are widely used for probabilistic modeling of text. MCMC sampling from\nthe posterior distribution is typically performed using a collapsed Gibbs\nsampler. We propose a parallel sparse partially collapsed Gibbs sampler and\ncompare its speed and efficiency to state-of-the-art samplers for topic models\non five well-known text corpora of differing sizes and properties. In\nparticular, we propose and compare two different strategies for sampling the\nparameter block with latent topic indicators. The experiments show that the\nincrease in statistical inefficiency from only partial collapsing is smaller\nthan commonly assumed, and can be more than compensated by the speedup from\nparallelization and sparsity on larger corpora. We also prove that the\npartially collapsed samplers scale well with the size of the corpus. The\nproposed algorithm is fast, efficient, exact, and can be used in more modeling\nsituations than the ordinary collapsed sampler.","url_abs":"http://arxiv.org/abs/1506.03784v3","url_pdf":"http://arxiv.org/pdf/1506.03784v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sparse-partially-collapsed-mcmc-for-parallel","repo_url":"https://github.com/lejon/PartiallyCollapsedLDA","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"topic-models","task_name":"Topic Models"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}