{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dense-distributions-from-sparse-samples","title":"Dense Distributions from Sparse Samples: Improved Gibbs Sampling Parameter Estimators for LDA","arxiv_id":"1505.02065","date":"2015-05-08","proceeding":null,"authors":["Yannis Papanikolaou","James R. Foulds","Timothy N. Rubin","Grigorios Tsoumakas"],"abstract":"We introduce a novel approach for estimating Latent Dirichlet Allocation\n(LDA) parameters from collapsed Gibbs samples (CGS), by leveraging the full\nconditional distributions over the latent variable assignments to efficiently\naverage over multiple samples, for little more computational cost than drawing\na single additional collapsed Gibbs sample. Our approach can be understood as\nadapting the soft clustering methodology of Collapsed Variational Bayes (CVB0)\nto CGS parameter estimation, in order to get the best of both techniques. Our\nestimators can straightforwardly be applied to the output of any existing\nimplementation of CGS, including modern accelerated variants. We perform\nextensive empirical comparisons of our estimators with those of standard\ncollapsed inference algorithms on real-world data for both unsupervised LDA and\nPrior-LDA, a supervised variant of LDA for multi-label classification. Our\nresults show a consistent advantage of our approach over traditional CGS under\nall experimental conditions, and over CVB0 inference in the majority of\nconditions. More broadly, our results highlight the importance of averaging\nover multiple samples in LDA parameter estimation, and the use of efficient\ncomputational techniques to do so.","url_abs":"http://arxiv.org/abs/1505.02065v6","url_pdf":"http://arxiv.org/pdf/1505.02065v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dense-distributions-from-sparse-samples","repo_url":"https://github.com/ypapanik/cgs_p","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"multi-label-classification-2","task_name":"MUlTI-LABEL-ClASSIFICATION"},{"task_slug":"multi-label-classification","task_name":"Multi-Label Classification"},{"task_slug":"parameter-estimation","task_name":"parameter estimation"}],"methods":[{"method_slug":"lda","method_name":"LDA"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}