Home › Datasets › task › Named Entity Recognition (NER)

Named Entity Recognition (NER) datasets

archive 2025-07-28

130 datasets carry the task tag "Named Entity Recognition (NER)" (the task itself: Named Entity Recognition (NER)), ordered by the archive's paper count. Page 1 of 3: 48 shown of 130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Named Entity Recognition (NER) datasets 1–48 of 130

CoNLL-2003 is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition.
755 papers · 6 benchmarks
The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts.
276 papers · 3 benchmarks
OntoNotes 5.0 is a large corpus comprising various genres of text (news, conversational telephone speech, weblogs, usenet newsgroups, broadcast, talk shows) in three languages (English, Chinese, and Arabic) with structural information…
254 papers · 12 benchmarks
BC5CDR (BioCreative V CDR corpus)
BC5CDR corpus consists of 1500 PubMed articles with 4409 annotated chemicals, 5818 diseases and 3116 chemical-disease interactions.
191 papers · 4 benchmarks
The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
FUNSD (Form Understanding in Noisy Scanned Documents)
Form Understanding in Noisy Scanned Documents (FUNSD) comprises 199 real, fully annotated, scanned forms.
179 papers · 3 benchmarks
The NCBI Disease corpus consists of 793 PubMed abstracts, which are separated into training (593), development (100) and test (100) subsets.
154 papers · 3 benchmarks
SciERC dataset is a collection of 500 scientific abstract annotated with scientific entities, their relations, and coreference clusters.
134 papers · 7 benchmarks
BLUE (Biomedical Language Understanding Evaluation)
The BLUE benchmark consists of five different biomedicine text-mining tasks with ten corpora.
133 papers · 0 benchmarks
WNUT 2017 (WNUT 2017 Emerging and Rare entity recognition)
This shared task focuses on identifying unusual, previously-unseen entities in the context of emerging discussions.
127 papers · 2 benchmarks
The GENIA corpus is the primary collection of biomedical literature compiled and annotated within the scope of the GENIA project.
121 papers · 7 benchmarks
CORD (Consolidated Receipt Dataset for Post-OCR Parsing)
OCR is inevitably linked to NLP since its final output is in text.
100 papers · 1 benchmark
RadGraph (RadGraph: Extracting Clinical Entities and Relations from Radiology Reports)
RadGraph is a dataset of entities and relations in radiology reports based on our novel information extraction schema, consisting of 600 reports with 30K radiologist annotations and 221K reports with 10.5M automatically generated…
78 papers · 0 benchmarks
Few-NERD is a large-scale, fine-grained manually annotated named entity recognition dataset, which contains 8 coarse-grained types, 66 fine-grained types, 188,200 sentences, 491,711 entities, and 4,601,223 tokens.
77 papers · 3 benchmarks
WikiANN (PAN-X)
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
The shared task of CoNLL-2002 concerns language-independent named entity recognition.
70 papers · 3 benchmarks
ACE 2005 (ACE 2005 Multilingual Training Corpus)
ACE 2005 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2005 Automatic Content Extraction (ACE) technology evaluation.
65 papers · 8 benchmarks
CoNLL++ is a corrected version of the CoNLL03 NER dataset where 5.38% of the test sentences have been fixed.
56 papers · 2 benchmarks
MasakhaNER is a collection of Named Entity Recognition (NER) datasets for 10 different African languages.
56 papers · 1 benchmark
ACE 2004 (ACE 2004 Multilingual Training Corpus)
ACE 2004 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2004 Automatic Content Extraction (ACE) technology evaluation.
51 papers · 6 benchmarks
MedMentions is a new manually annotated resource for the recognition of biomedical concepts.
48 papers · 1 benchmark
MultiCoNER is a large multilingual dataset (11 languages) for Named Entity Recognition.
47 papers · 0 benchmarks
A dataset of financial agreements made public through U.S.
39 papers · 0 benchmarks
IPM NEL (Derczynski IPM Named Entity Linking)
This data is for the task of named entity recognition and linking/disambiguation over tweets.
32 papers · 1 benchmark
WikiCoref is an English corpus annotated for anaphoric relations, where all documents are from the English version of Wikipedia.
28 papers · 1 benchmark
BioRED is a first-of-its-kind biomedical relation extraction dataset with multiple entity types (e.g.
25 papers · 3 benchmarks
SLUE (Spoken Language Understanding Evaluation)
Spoken Language Understanding Evaluation (SLUE) is a suite of benchmark tasks for spoken language understanding evaluation.
22 papers · 3 benchmarks
KLUE (Korean Language Understanding Evaluation)
Korean Language Understanding Evaluation (KLUE) benchmark is a series of datasets to evaluate natural language understanding capability of Korean language models.
21 papers · 1 benchmark
LinCE (Linguistic Code-switching Evaluation Dataset)
A centralized benchmark for Linguistic Code-switching Evaluation (LinCE) that combines ten corpora covering four different code-switched language pairs (i.e., Spanish-English, Nepali-English, Hindi-English, and Modern Standard…
21 papers · 0 benchmarks
JNLPBA is a biomedical dataset that comes from the GENIA version 3.02 corpus (Kim et al., 2003).
20 papers · 2 benchmarks
XFUND (A Multilingual Form Understanding Benchmark)
XFUND is a multilingual form understanding benchmark dataset that includes human-labeled forms with key-value pairs in 7 languages (Chinese, Japanese, Spanish, French, Italian, German, Portuguese).
20 papers · 0 benchmarks
NNE is a dataset for Nested Named Entity Recognition in English Newswire
19 papers · 1 benchmark
WNUT 2016 NER (WNUT 2016 Twitter Named Entity Recognition)
19 papers · 1 benchmark
DWIE (Deutsche Welle corpus for Information Extraction)
The 'Deutsche Welle corpus for Information Extraction' (DWIE) is a multi-task dataset that combines four main Information Extraction (IE) annotation sub-tasks: (i) Named Entity Recognition (NER), (ii) Coreference Resolution, (iii) Relation…
18 papers · 5 benchmarks
Groningen Meaning Bank is a semantic resource that anyone can edit and that integrates various semantic phenomena, including predicate-argument structure, scope, tense, thematic roles, animacy, pronouns, and rhetorical relations.
18 papers · 0 benchmarks
IndicGLUE (Indic General Language Understanding Evaluation Benchmark)
We now introduce IndicGLUE, the Indic General Language Understanding Evaluation Benchmark, which is a collection of various NLP tasks as de- scribed below.
16 papers · 4 benchmarks
CrossNER is a cross-domain NER (Named Entity Recognition) dataset, a fully-labeled collection of NER data spanning over five diverse domains (Politics, Natural Science, Music, Literature, and Artificial Intelligence) with specialized…
15 papers · 1 benchmark
Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports.
14 papers · 3 benchmarks
Created by Smith et al.
13 papers · 2 benchmarks
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
This paper introduces the Broad Twitter Corpus (BTC), which is not only significantly bigger, but sampled across different regions, temporal periods, and types of Twitter users.
12 papers · 2 benchmarks
Earnings-21, a 39-hour corpus of earnings calls containing entity-dense speech from nine different financial sectors.
12 papers · 0 benchmarks
NCBI Datasets is a valuable resource that simplifies the process of gathering data from various NCBI databases.
11 papers · 0 benchmarks
CMeEE (Chinese Medical Named Entity Recognition Dataset)
Chinese Medical Named Entity Recognition, a dataset first released in CHIP20204, is used for CMeEE task.
10 papers · 1 benchmark
Polyglot-NER builds massive multilingual annotators with minimal human expertise and intervention.
10 papers · 0 benchmarks
GeoWebNews provides test/train examples and enable fine-grained Geotagging and Toponym Resolution (Geocoding).
9 papers · 0 benchmarks
2018 n2c2 (Track 2) - Adverse Drug Events and Medication Extraction (2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records)
Abstract Objective This article summarizes the preparation, organization, evaluation, and results of Track 2 of the 2018 National NLP Clinical Challenges shared task.
8 papers · 0 benchmarks
BB (Bacteria Biotope)
The Bacteria Biotope (BB) Task is part of the BioNLP Open Shared Tasks and meets the BioNLP-OST standards of quality, originality and data formats.
8 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.