{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/latent-dirichlet-allocation","title":"Latent Dirichlet Allocation","arxiv_id":null,"date":"2003-01-01","proceeding":null,"authors":["David M. Blei","Andrew Y. Ng","Michael I. Jordan"],"abstract":"We describe latent Dirichlet allocation (LDA), a generative probabilistic model for collections of\r\ndiscrete data such as text corpora. LDA is a three-level hierarchical Bayesian model, in which each\r\nitem of a collection is modeled as a finite mixture over an underlying set of topics. Each topic is, in\r\nturn, modeled as an infinite mixture over an underlying set of topic probabilities. In the context of\r\ntext modeling, the topic probabilities provide an explicit representation of a document. We present\r\nefficient approximate inference techniques based on variational methods and an EM algorithm for\r\nempirical Bayes parameter estimation. We report results in document modeling, text classification,\r\nand collaborative filtering, comparing to a mixture of unigrams model and the probabilistic LSI\r\nmodel.","url_abs":"https://dl.acm.org/doi/10.5555/944919.944937#d7400906e1","url_pdf":"https://www.jmlr.org/papers/volume3/blei03a/blei03a.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"latent-dirichlet-allocation","repo_url":"https://github.com/annasanchez27/DD2434-advanced-ml","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"latent-dirichlet-allocation","repo_url":"https://github.com/vrjkmr/arxiv-topic","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"collaborative-filtering","task_name":"Collaborative Filtering"},{"task_slug":"text-categorization","task_name":"Text Categorization"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"topic-models","task_name":"Topic Models"},{"task_slug":"parameter-estimation","task_name":"parameter estimation"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[{"method_slug":"stochastic-gradient-variational-bayes","method_name":"Stochastic Gradient Variational Bayes"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}