{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/short-text-hashing-improved-by-integrating","title":"Short Text Hashing Improved by Integrating Multi-Granularity Topics and Tags","arxiv_id":"1503.02801","date":"2015-03-10","proceeding":null,"authors":["Jiaming Xu","Bo Xu","Guanhua Tian","Jun Zhao","Fangyuan Wang","Hong-Wei Hao"],"abstract":"Due to computational and storage efficiencies of compact binary codes,\nhashing has been widely used for large-scale similarity search. Unfortunately,\nmany existing hashing methods based on observed keyword features are not\neffective for short texts due to the sparseness and shortness. Recently, some\nresearchers try to utilize latent topics of certain granularity to preserve\nsemantic similarity in hash codes beyond keyword matching. However, topics of\ncertain granularity are not adequate to represent the intrinsic semantic\ninformation. In this paper, we present a novel unified approach for short text\nHashing using Multi-granularity Topics and Tags, dubbed HMTT. In particular, we\npropose a selection method to choose the optimal multi-granularity topics\ndepending on the type of dataset, and design two distinct hashing strategies to\nincorporate multi-granularity topics. We also propose a simple and effective\nmethod to exploit tags to enhance the similarity of related texts. We carry out\nextensive experiments on one short text dataset as well as on one normal text\ndataset. The results demonstrate that our approach is effective and\nsignificantly outperforms baselines on several evaluation metrics.","url_abs":"http://arxiv.org/abs/1503.02801v1","url_pdf":"http://arxiv.org/pdf/1503.02801v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"short-text-hashing-improved-by-integrating","repo_url":"https://github.com/jacoxu/short-text-hashing-HMTT","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}