{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hitr-hierarchical-topic-model-re-estimation","title":"HiTR: Hierarchical Topic Model Re-estimation for Measuring Topical Diversity of Documents","arxiv_id":"1810.05436","date":"2018-10-12","proceeding":null,"authors":["Hosein Azarbonyad","Mostafa Dehghani","Tom Kenter","Maarten Marx","Jaap Kamps","Maarten de Rijke"],"abstract":"A high degree of topical diversity is often considered to be an important\ncharacteristic of interesting text documents. A recent proposal for measuring\ntopical diversity identifies three distributions for assessing the diversity of\ndocuments: distributions of words within documents, words within topics, and\ntopics within documents. Topic models play a central role in this approach and,\nhence, their quality is crucial to the efficacy of measuring topical diversity.\nThe quality of topic models is affected by two causes: generality and impurity\nof topics. General topics only include common information of a background\ncorpus and are assigned to most of the documents. Impure topics contain words\nthat are not related to the topic. Impurity lowers the interpretability of\ntopic models. Impure topics are likely to get assigned to documents\nerroneously. We propose a hierarchical re-estimation process aimed at removing\ngenerality and impurity. Our approach has three re-estimation components: (1)\ndocument re-estimation, which removes general words from the documents; (2)\ntopic re-estimation, which re-estimates the distribution over words of each\ntopic; and (3) topic assignment re-estimation, which re-estimates for each\ndocument its distributions over topics. For measuring topical diversity of text\ndocuments, our HiTR approach improves over the state-of-the-art measured on\nPubMed dataset.","url_abs":"http://arxiv.org/abs/1810.05436v1","url_pdf":"http://arxiv.org/pdf/1810.05436v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hitr-hierarchical-topic-model-re-estimation","repo_url":"https://github.com/HoseinAzarbonyad/HiTR","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"topic-models","task_name":"Topic Models"}],"methods":[{"method_slug":"interpretability","method_name":"Interpretability"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}