Browse State-of-the-Art › Keyword Extraction
Keyword Extraction
33 papers with code · 3 benchmarks · 7 datasets archive 2025-07-28
Keyword extraction is tasked with the automatic identification of terms that best describe the subject of a document (Source: Wikipedia).
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Inspec (5 rows) | Phraseformer(BERT, ExEm(ft)) | Phraseformer: Multimodal Key-phrase Extraction using Transformer... | — | — | Compare |
| SemEval 2010 Task 8 (5 rows) | Phraseformer(BERT, ExEm(ft)) | Phraseformer: Multimodal Key-phrase Extraction using Transformer... | — | — | Compare |
| SemEval-2017 Task-10 (5 rows) | Phraseformer(BERT, ExEm(ft)) | Phraseformer: Multimodal Key-phrase Extraction using Transformer... | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 33 papers with code (172 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Nov 2018 5 repositories listedCombination of the proposed graph construction and scoring methods leads to a novel, parameterless keyword extraction method (sCAKE) based on semantic connectivity of words in the document.
-
26 Sep 2019 2 repositories listedThis shows that the proposed method is independent of the domain, collection, and language of the training corpora.
-
19 Feb 2025 1 repository listedIn this paper, we propose a CPS data collection protocol and create a new CPS dataset, called PSCon, which assists product search through conversations with human-like language.
-
18 Dec 2024 1 repository listedRelying on the insight that real-world keyword detection often requires handling of diverse content, we propose a novel supervised keyword extraction approach based on the mixture of experts (MoE) technique.
-
28 Oct 2024 1 repository listedRecent studies have demonstrated that large language models (LLMs) exhibit significant biases in evaluation tasks, particularly in preferentially rating and favoring self-generated content.
-
23 Jul 2024 1 repository listedThe superior performance of re-identified over de-identified data suggests a shift towards methods that enhance utility and privacy by using dummy PHIs to perplex privacy attacks.
-
19 Jul 2024 1 repository listedThe task of keyword extraction is often an important initial step in unsupervised information extraction, forming the basis for tasks such as topic modeling or document classification.
-
9 Jul 2024 1 repository listedOur findings demonstrate that models fine-tuned using the \textbf{Task-Fine-Tune} methodology not only achieve superior performance on these specific tasks but also significantly outperform models with higher general…
-
8 Jul 2024 1 repository listedWe study open-world multi-label text classification under extremely weak supervision (XWS), where the user only provides a brief description for classification objectives without any labels or ground-truth label space.
-
27 May 2024 1 repository listedThis paper introduces KSW, a Khmer-specific approach to keyword extraction that leverages a specialized stop word dictionary.
-
10 Feb 2024 1 repository listedSecond, we propose a new hybrid architecture that merges the cascaded and parallel architectures of SpeechCLIP into a multi-task learning framework.
-
30 Mar 2023 1 repository listedExisting conversational models are handled by a database(DB) and API based systems.
-
19 Dec 2022 1 repository listedThe two important tasks to do this are keyword extraction and text summarization.
-
14 Nov 2022 1 repository listedDownstream training for keyword extractors is a lengthy process and requires a significant amount of data.
-
13 Nov 2022 1 repository listedHowever, we find that large-scale bidirectional training between image and text enables zero-shot image captioning.
-
22 Sep 2022 1 repository listedKeyword spotting (KWS) has become a hot topic in speech processing due to the rise of commercial applications based on voice command detection, such as voice assistants.
-
13 Nov 2021 1 repository listedThis paper presents an experimental analysis of similarity scores of keywords generated by different supervised and unsupervised automated keyword extraction algorithms with expert provided keywords from the Electric…
-
13 Oct 2021 1 repository listedIn this work, we propose a novel unsupervised embedding-based KPE approach, Masked Document Embedding Rank (MDERank), to address this problem by leveraging a mask strategy and ranking candidates by the similarity…
-
2 Jul 2021 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedIn this work, we present to the NLP community, and to the wider research community as a whole, an application for the diachronic analysis of research corpora.
-
16 Apr 2021 1 repository listed Syntology ran 0 of 6 samples · 6 unverifiedTerm weighting schemes are widely used in Natural Language Processing and Information Retrieval.
-
10 Apr 2021 1 repository listedFor evaluating the proposed method, seven datasets were used: Semeval2010, SemEval2017, Inspec, fao30, Thesis100, pak2018, and Wikinews, with results reported as Precision, Recall, and F- measure.
-
10 Feb 2021 1 repository listedKeyphrase extraction methods can provide insights into large collections of documents such as social media posts.
-
31 Jan 2021 1 repository listedKeyword extraction is the task of identifying words (or multi-word expressions) that best describe a given document and serve in news portals to link articles of similar topics.
-
4 Jan 2021 1 repository listedOur paper is among the first ones by our knowledge to propose a model and to create datasets for the task of "outline to story".
-
17 Dec 2020 1 repository listedComponent (a) directly generates a matching score of a candidate document for a query.
-
21 Aug 2020 1 repository listedKeyword extraction is an important document process that aims at finding a small set of terms that concisely describe a document's topics.
-
20 Mar 2020 1 repository listedWith growing amounts of available textual data, development of algorithms capable of automatic analysis, categorization and summarization of these data has become a necessity.
-
6 Jan 2020 1 repository listedKeyword extraction has received an increasing attention as an important research topic which can lead to have advancements in diverse applications such as document context categorization, text indexing and document…
-
15 Jul 2019 1 repository listedKeyword extraction is used for summarizing the content of a document and supports efficient document retrieval, and is as such an indispensable part of modern text-based systems.
-
1 Jun 2018 1 repository listedCorpus2graph is an open-source NLP-application-oriented tool that generates a word co-occurrence network from a large corpus.
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections