{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sampled-weighted-min-hashing-for-large-scale","title":"Sampled Weighted Min-Hashing for Large-Scale Topic Mining","arxiv_id":"1509.01771","date":"2015-09-06","proceeding":null,"authors":["Gibran Fuentes-Pineda","Ivan Vladimir Meza-Ruiz"],"abstract":"We present Sampled Weighted Min-Hashing (SWMH), a randomized approach to\nautomatically mine topics from large-scale corpora. SWMH generates multiple\nrandom partitions of the corpus vocabulary based on term co-occurrence and\nagglomerates highly overlapping inter-partition cells to produce the mined\ntopics. While other approaches define a topic as a probabilistic distribution\nover a vocabulary, SWMH topics are ordered subsets of such vocabulary.\nInterestingly, the topics mined by SWMH underlie themes from the corpus at\ndifferent levels of granularity. We extensively evaluate the meaningfulness of\nthe mined topics both qualitatively and quantitatively on the NIPS (1.7 K\ndocuments), 20 Newsgroups (20 K), Reuters (800 K) and Wikipedia (4 M) corpora.\nAdditionally, we compare the quality of SWMH with Online LDA topics for\ndocument representation in classification.","url_abs":"http://arxiv.org/abs/1509.01771v2","url_pdf":"http://arxiv.org/pdf/1509.01771v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sampled-weighted-min-hashing-for-large-scale","repo_url":"https://github.com/gibranfp/Sampled-MinHashing","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"}],"methods":[{"method_slug":"lda","method_name":"LDA"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}