{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/short-text-topic-modeling-techniques","title":"Short Text Topic Modeling Techniques, Applications, and Performance: A Survey","arxiv_id":"1904.07695","date":"2019-04-13","proceeding":null,"authors":["Qiang Jipeng","Qian Zhenyu","Li Yun","Yuan Yunhao","Wu Xindong"],"abstract":"Analyzing short texts infers discriminative and coherent latent topics that\nis a critical and fundamental task since many real-world applications require\nsemantic understanding of short texts. Traditional long text topic modeling\nalgorithms (e.g., PLSA and LDA) based on word co-occurrences cannot solve this\nproblem very well since only very limited word co-occurrence information is\navailable in short texts. Therefore, short text topic modeling has already\nattracted much attention from the machine learning research community in recent\nyears, which aims at overcoming the problem of sparseness in short texts. In\nthis survey, we conduct a comprehensive review of various short text topic\nmodeling techniques proposed in the literature. We present three categories of\nmethods based on Dirichlet multinomial mixture, global word co-occurrences, and\nself-aggregation, with example of representative approaches in each category\nand analysis of their performance on various tasks. We develop the first\ncomprehensive open-source library, called STTM, for use in Java that integrates\nall surveyed algorithms within a unified interface, benchmark datasets, to\nfacilitate the expansion of new methods in this research field. Finally, we\nevaluate these state-of-the-art methods on many real-world datasets and compare\ntheir performance against one another and versus long text topic modeling\nalgorithm.","url_abs":"http://arxiv.org/abs/1904.07695v1","url_pdf":"http://arxiv.org/pdf/1904.07695v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"short-text-topic-modeling-techniques","repo_url":"https://github.com/qiang2100/STTM","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"survey","task_name":"Survey"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.07695","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}