Browse State-of-the-Art › NER
NER
665 papers with code · 4 benchmarks · 27 datasets archive 2025-07-28
The named entity recognition (NER) involves identification of key information in the text and classification into a set of predefined categories. This includes standard entities in the text like Part of Speech (PoS) and entities like places, names etc...
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CoNLL 2003 (1 row) | Def2Vec | Def2Vec: Extensible Word Embeddings from Dictionary Definitions | code | — | Compare |
| InLegalNER (1 row) | opennyaiorg/en_legal_ner_trf | Named Entity Recognition in Indian court judgments | code | — | Compare |
| SuperMat (1 row) | superconductors-Scibert | Automatic extraction of materials and properties from... | code | — | Compare |
| The EMBO SourceData-NLP dataset (1 row) | BioLinkBERT-large | Integrating curation into scientific publishing to train AI models | code | — | Compare |
| DaNE (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
27 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 665 papers with code (1,729 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
31 Aug 2019 10 repositories listedThe pre-trained language models have achieved great successes in various natural language understanding (NLU) tasks due to its capacity to capture the deep contextualized information in text by pre-training on…
-
8 Dec 2020 5 repositories listedCross-domain named entity recognition (NER) models are able to cope with the scarcity issue of NER samples in target domains.
-
28 Feb 2020 4 repositories listedRecently, with the surge of transformers based models, language-specific BERT based models have proven to be very efficient at language understanding, provided they are pre-trained on a very large corpus.
-
7 Feb 2017 4 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Today when many practitioners run basic NLP on the entire web and large-volume traffic, faster methods are paramount to saving time and energy costs.
-
27 Nov 2024 3 repositories listedDesigning resilient Internet of Things (IoT) systems requires i) identification of IoT Critical Objects (ICOs) such as services, devices, and resources, ii) threat analysis, and iii) mitigation strategy selection.
-
24 Jun 2024 3 repositories listedThis paper introduces a novel, entity-aware metric, termed as Radiological Report (Text) Evaluation (RaTEScore), to assess the quality of medical reports generated by AI models.
-
22 May 2023 3 repositories listedIn this paper, we propose DiffusionNER, which formulates the named entity recognition task as a boundary-denoising diffusion process and thus generates named entities from noisy spans.
-
22 Feb 2023 3 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Over the last two decades, the development of the CoNLL-2003 named entity recognition (NER) dataset has helped enhance the capabilities of deep learning and natural language processing (NLP).
-
22 Sep 2020 3 repositories listedDespite efforts to distinguish three different evaluation setups (Bekoulis et al., 2018), numerous end-to-end Relation Extraction (RE) articles present unreliable performance comparison to previous work.
-
19 Mar 2020 3 repositories listedIn this paper, we use the pre-trained deep bidirectional network, BERT, to make a model for named entity recognition in Persian.
-
16 Mar 2020 3 repositories listedIn all 20 annotation sets of the concept-annotation task, our system outperforms the pipeline system reported as a baseline in the CRAFT shared task 2019.
-
13 Jan 2020 3 repositories listedIn this paper, we introduce the NER dataset from CLUE organization (CLUENER2020), a well-defined fine-grained dataset for named entity recognition in Chinese.
-
29 Aug 2019 3 repositories listedWe test the practical impacts of the deficiency on real-world NER datasets, OntoNotes 5.
-
5 May 2018 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We investigate a lattice-structured LSTM model for Chinese NER, which encodes a sequence of input characters as well as all potential words that match a lexicon.
-
13 Sep 2017 3 repositories listedIn this study, we develop a novel neural framework to extract abundant knowledge hidden in raw texts to empower the sequence labeling task.
-
1 Apr 2025 2 repositories listedWe find that the best-performing mLLM model significantly outperforms conventional state-of-the-art OCR models and other frontier mLLMs.
-
28 Sep 2024 2 repositories listedMedication Extraction and Mining play an important role in healthcare NLP research due to its practical applications in hospital settings, such as their mapping into standard clinical knowledge bases (SNOMED-CT, BNF,…
-
12 Sep 2024 2 repositories listedIn this paper, we introduce WhisperNER, a novel model that allows joint speech transcription and entity recognition.
-
20 Apr 2024 2 repositories listedWe test widely used NER toolkits and transformer models, including models using the pre-trained contextual models RoBERTa and ELECTRA, on three datasets: a commonly used British English newswire dataset, CoNLL 2003, a…
-
18 Mar 2024 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Streaming text generation has become a common way of increasing the responsiveness of language model powered applications, such as chat assistants.
-
18 Mar 2024 2 repositories listedOur findings provide insights into the differing performance of BERT- and T5-based models for the NER task.
-
26 Feb 2024 2 repositories listedIn the Named Entity Recognition (NER) task, recent advancements have seen the remarkable improvement of LLMs in a broad range of entity domains via instruction tuning, by adopting entity-centric schema.
-
16 Dec 2023 2 repositories listedDef2Vec introduces a novel paradigm for word embeddings, leveraging dictionary definitions to learn semantic representations.
-
15 Nov 2023 2 repositories listedWe introduce Universal NER (UNER), an open, community-driven project to develop gold-standard NER benchmarks in many languages.
-
14 Nov 2023 2 repositories listedNamed Entity Recognition (NER) is essential in various Natural Language Processing (NLP) applications.
-
17 Oct 2023 2 repositories listedHowever, BIO-tagging scheme relies on the correct order of model inputs, which is not guaranteed in real-world NER on scanned VrDs where text are recognized and arranged by OCR systems.
-
13 Oct 2023 2 repositories listedNatural language processing (NLP) has made significant progress for well-resourced languages such as English but lagged behind for low-resource languages like Setswana.
-
2 Oct 2023 2 repositories listedWe evaluate this approach with Label Supervised LLaMA (LS-LLaMA), based on LLaMA-2-7B, a relatively small-scale LLM, and can be finetuned on a single GeForce RTX4090 GPU.
-
7 Aug 2023 2 repositories listed Syntology ran 5 of 12 samples · 7 unverifiedInstruction tuning has proven effective for distilling LLMs into more cost-efficient models such as Alpaca and Vicuna.
-
28 Jun 2023 2 repositories listedConclusions: This study advances public health research by implementing a novel, systematic pipeline for curating symptom lexicons from social media data.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections