{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-label-dataless-text-classification-with","title":"Multi-label Dataless Text Classification with Topic Modeling","arxiv_id":"1711.01563","date":"2017-11-05","proceeding":null,"authors":["Daochen Zha","Chenliang Li"],"abstract":"Manually labeling documents is tedious and expensive, but it is essential for\ntraining a traditional text classifier. In recent years, a few dataless text\nclassification techniques have been proposed to address this problem. However,\nexisting works mainly center on single-label classification problems, that is,\neach document is restricted to belonging to a single category. In this paper,\nwe propose a novel Seed-guided Multi-label Topic Model, named SMTM. With a few\nseed words relevant to each category, SMTM conducts multi-label classification\nfor a collection of documents without any labeled document. In SMTM, each\ncategory is associated with a single category-topic which covers the meaning of\nthe category. To accommodate with multi-labeled documents, we explicitly model\nthe category sparsity in SMTM by using spike and slab prior and weak smoothing\nprior. That is, without using any threshold tuning, SMTM automatically selects\nthe relevant categories for each document. To incorporate the supervision of\nthe seed words, we propose a seed-guided biased GPU (i.e., generalized Polya\nurn) sampling procedure to guide the topic inference of SMTM. Experiments on\ntwo public datasets show that SMTM achieves better classification accuracy than\nstate-of-the-art alternatives and even outperforms supervised solutions in some\nscenarios.","url_abs":"http://arxiv.org/abs/1711.01563v1","url_pdf":"http://arxiv.org/pdf/1711.01563v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-label-dataless-text-classification-with","repo_url":"https://github.com/WHUIR/SMTM","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":null,"task_name":"GPU"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"multi-label-classification-2","task_name":"MUlTI-LABEL-ClASSIFICATION"},{"task_slug":"multi-label-classification","task_name":"Multi-Label Classification"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}