{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scaling-up-dynamic-topic-models","title":"Scaling up Dynamic Topic Models","arxiv_id":"1602.06049","date":"2016-02-19","proceeding":null,"authors":["Arnab Bhadury","Jianfei Chen","Jun Zhu","Shixia Liu"],"abstract":"Dynamic topic models (DTMs) are very effective in discovering topics and\ncapturing their evolution trends in time series data. To do posterior inference\nof DTMs, existing methods are all batch algorithms that scan the full dataset\nbefore each update of the model and make inexact variational approximations\nwith mean-field assumptions. Due to a lack of a more scalable inference\nalgorithm, despite the usefulness, DTMs have not captured large topic dynamics.\n  This paper fills this research void, and presents a fast and parallelizable\ninference algorithm using Gibbs Sampling with Stochastic Gradient Langevin\nDynamics that does not make any unwarranted assumptions. We also present a\nMetropolis-Hastings based $O(1)$ sampler for topic assignments for each word\ntoken. In a distributed environment, our algorithm requires very little\ncommunication between workers during sampling (almost embarrassingly parallel)\nand scales up to large-scale applications. We are able to learn the largest\nDynamic Topic Model to our knowledge, and learned the dynamics of 1,000 topics\nfrom 2.6 million documents in less than half an hour, and our empirical results\nshow that our algorithm is not only orders of magnitude faster than the\nbaselines but also achieves lower perplexity.","url_abs":"http://arxiv.org/abs/1602.06049v1","url_pdf":"http://arxiv.org/pdf/1602.06049v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scaling-up-dynamic-topic-models","repo_url":"https://github.com/alexismailov2/FastDTM","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"time-series-1","task_name":"Time Series"},{"task_slug":"time-series","task_name":"Time Series Analysis"},{"task_slug":"topic-models","task_name":"Topic Models"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}