{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/automated-phrase-mining-from-massive-text","title":"Automated Phrase Mining from Massive Text Corpora","arxiv_id":"1702.04457","date":"2017-02-15","proceeding":null,"authors":["Jingbo Shang","Jialu Liu","Meng Jiang","Xiang Ren","Clare R. Voss","Jiawei Han"],"abstract":"As one of the fundamental tasks in text analysis, phrase mining aims at\nextracting quality phrases from a text corpus. Phrase mining is important in\nvarious tasks such as information extraction/retrieval, taxonomy construction,\nand topic modeling. Most existing methods rely on complex, trained linguistic\nanalyzers, and thus likely have unsatisfactory performance on text corpora of\nnew domains and genres without extra but expensive adaption. Recently, a few\ndata-driven methods have been developed successfully for extraction of phrases\nfrom massive domain-specific text. However, none of the state-of-the-art models\nis fully automated because they require human experts for designing rules or\nlabeling phrases.\n  Since one can easily obtain many quality phrases from public knowledge bases\nto a scale that is much larger than that produced by human experts, in this\npaper, we propose a novel framework for automated phrase mining, AutoPhrase,\nwhich leverages this large amount of high-quality phrases in an effective way\nand achieves better performance compared to limited human labeled phrases. In\naddition, we develop a POS-guided phrasal segmentation model, which\nincorporates the shallow syntactic information in part-of-speech (POS) tags to\nfurther enhance the performance, when a POS tagger is available. Note that,\nAutoPhrase can support any language as long as a general knowledge base (e.g.,\nWikipedia) in that language is available, while benefiting from, but not\nrequiring, a POS tagger. Compared to the state-of-the-art methods, the new\nmethod has shown significant improvements in effectiveness on five real-world\ndatasets across different domains and languages.","url_abs":"http://arxiv.org/abs/1702.04457v2","url_pdf":"http://arxiv.org/pdf/1702.04457v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"automated-phrase-mining-from-massive-text","repo_url":"https://github.com/shangjingbo1226/AutoPhrase","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"automated-phrase-mining-from-massive-text","repo_url":"https://github.com/darthbatman/simplified-ECON","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"automated-phrase-mining-from-massive-text","repo_url":"https://github.com/dmis-lab/gener","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"automated-phrase-mining-from-massive-text","repo_url":"https://github.com/shangjingbo1226/AutoNER","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"general-knowledge","task_name":"General Knowledge"},{"task_slug":"pos","task_name":"POS"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.04457","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}