{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fast-flexible-models-for-discovering-topic","title":"Fast, Flexible Models for Discovering Topic Correlation across Weakly-Related Collections","arxiv_id":"1508.04562","date":"2015-08-19","proceeding":"EMNLP 2015 9","authors":["Jingwei Zhang","Aaron Gerow","Jaan Altosaar","James Evans","Richard Jean So"],"abstract":"Weak topic correlation across document collections with different numbers of\ntopics in individual collections presents challenges for existing\ncross-collection topic models. This paper introduces two probabilistic topic\nmodels, Correlated LDA (C-LDA) and Correlated HDP (C-HDP). These address\nproblems that can arise when analyzing large, asymmetric, and potentially\nweakly-related collections. Topic correlations in weakly-related collections\ntypically lie in the tail of the topic distribution, where they would be\noverlooked by models unable to fit large numbers of topics. To efficiently\nmodel this long tail for large-scale analysis, our models implement a parallel\nsampling algorithm based on the Metropolis-Hastings and alias methods (Yuan et\nal., 2015). The models are first evaluated on synthetic data, generated to\nsimulate various collection-level asymmetries. We then present a case study of\nmodeling over 300k documents in collections of sciences and humanities research\nfrom JSTOR.","url_abs":"http://arxiv.org/abs/1508.04562v1","url_pdf":"http://arxiv.org/pdf/1508.04562v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fast-flexible-models-for-discovering-topic","repo_url":"https://github.com/iceboal/correlated-lda","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"topic-models","task_name":"Topic Models"}],"methods":[{"method_slug":"lda","method_name":"LDA"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}