{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/simple-unsupervised-keyphrase-extraction","title":"Simple Unsupervised Keyphrase Extraction using Sentence Embeddings","arxiv_id":"1801.04470","date":"2018-01-13","proceeding":"CONLL 2018 10","authors":["Kamil Bennani-Smires","Claudiu Musat","Andreea Hossmann","Michael Baeriswyl","Martin Jaggi"],"abstract":"Keyphrase extraction is the task of automatically selecting a small set of\nphrases that best describe a given free text document. Supervised keyphrase\nextraction requires large amounts of labeled training data and generalizes very\npoorly outside the domain of the training data. At the same time, unsupervised\nsystems have poor accuracy, and often do not generalize well, as they require\nthe input document to belong to a larger corpus also given as input. Addressing\nthese drawbacks, in this paper, we tackle keyphrase extraction from single\ndocuments with EmbedRank: a novel unsupervised method, that leverages sentence\nembeddings. EmbedRank achieves higher F-scores than graph-based state of the\nart systems on standard datasets and is suitable for real-time processing of\nlarge amounts of Web data. With EmbedRank, we also explicitly increase coverage\nand diversity among the selected keyphrases by introducing an embedding-based\nmaximal marginal relevance (MMR) for new phrases. A user study including over\n200 votes showed that, although reducing the phrases' semantic overlap leads to\nno gains in F-score, our high diversity selection is preferred by humans.","url_abs":"http://arxiv.org/abs/1801.04470v3","url_pdf":"http://arxiv.org/pdf/1801.04470v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"simple-unsupervised-keyphrase-extraction","repo_url":"https://github.com/swisscom/ai-research-keyphrase-extraction","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"simple-unsupervised-keyphrase-extraction","repo_url":"https://github.com/AnzorGozalishvili/embedrank_serving","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"simple-unsupervised-keyphrase-extraction","repo_url":"https://github.com/AnzorGozalishvili/unsupervised_keyword_extraction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"keyphrase-extraction","task_name":"Keyphrase Extraction"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"sentence-embeddings","task_name":"Sentence Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1801.04470","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}