{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/supervised-topical-key-phrase-extraction-of","title":"Supervised Topical Key Phrase Extraction of News Stories using Crowdsourcing, Light Filtering and Co-reference Normalization","arxiv_id":"1306.4886","date":"2013-06-20","proceeding":"LREC 2012 5","authors":["Luis Marujo","Anatole Gershman","Jaime Carbonell","Robert Frederking","João P. Neto"],"abstract":"Fast and effective automated indexing is critical for search and personalized\nservices. Key phrases that consist of one or more words and represent the main\nconcepts of the document are often used for the purpose of indexing. In this\npaper, we investigate the use of additional semantic features and\npre-processing steps to improve automatic key phrase extraction. These features\ninclude the use of signal words and freebase categories. Some of these features\nlead to significant improvements in the accuracy of the results. We also\nexperimented with 2 forms of document pre-processing that we call light\nfiltering and co-reference normalization. Light filtering removes sentences\nfrom the document, which are judged peripheral to its main content.\nCo-reference normalization unifies several written forms of the same named\nentity into a unique form. We also needed a \"Gold Standard\" - a set of labeled\ndocuments for training and evaluation. While the subjective nature of key\nphrase selection precludes a true \"Gold Standard\", we used Amazon's Mechanical\nTurk service to obtain a useful approximation. Our data indicates that the\nbiggest improvements in performance were due to shallow semantic features, news\ncategories, and rhetorical signals (nDCG 78.47% vs. 68.93%). The inclusion of\ndeeper semantic features such as Freebase sub-categories was not beneficial by\nitself, but in combination with pre-processing, did cause slight improvements\nin the nDCG scores.","url_abs":"http://arxiv.org/abs/1306.4886v1","url_pdf":"http://arxiv.org/pdf/1306.4886v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"supervised-topical-key-phrase-extraction-of","repo_url":"https://github.com/LIAAD/KeywordExtractor-Datasets","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1306.4886","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}