{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/metalda-a-topic-model-that-efficiently","title":"MetaLDA: a Topic Model that Efficiently Incorporates Meta information","arxiv_id":"1709.06365","date":"2017-09-19","proceeding":null,"authors":["He Zhao","Lan Du","Wray Buntine","Gang Liu"],"abstract":"Besides the text content, documents and their associated words usually come\nwith rich sets of meta informa- tion, such as categories of documents and\nsemantic/syntactic features of words, like those encoded in word embeddings.\nIncorporating such meta information directly into the generative process of\ntopic models can improve modelling accuracy and topic quality, especially in\nthe case where the word-occurrence information in the training data is\ninsufficient. In this paper, we present a topic model, called MetaLDA, which is\nable to leverage either document or word meta information, or both of them\njointly. With two data argumentation techniques, we can derive an efficient\nGibbs sampling algorithm, which benefits from the fully local conjugacy of the\nmodel. Moreover, the algorithm is favoured by the sparsity of the meta\ninformation. Extensive experiments on several real world datasets demonstrate\nthat our model achieves comparable or improved performance in terms of both\nperplexity and topic quality, particularly in handling sparse texts. In\naddition, compared with other models using meta information, our model runs\nsignificantly faster.","url_abs":"http://arxiv.org/abs/1709.06365v1","url_pdf":"http://arxiv.org/pdf/1709.06365v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"metalda-a-topic-model-that-efficiently","repo_url":"https://github.com/ethanhezhao/MetaLDA","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"topic-models","task_name":"Topic Models"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1709.06365","atlas_url":"https://app.syntology.ai/?focus=1709.06365","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}