Home › Datasets › task › Named Entity Recognition (NER)
Named Entity Recognition (NER) datasets
archive 2025-07-28
130 datasets carry the task tag "Named Entity Recognition (NER)" (the task itself: Named Entity Recognition (NER)), ordered by the archive's paper count. Page 2 of 3: 48 shown of 130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Named Entity Recognition (NER) datasets 49–96 of 130
CoNLL-2000 is a dataset for dividing text into syntactically related non-overlapping groups of words, so-called text chunking.
8 papers · 0 benchmarks
Peyma is a Persian NER dataset to train and test NER systems.
8 papers · 0 benchmarks
Winogender Schemas is a novel, Winograd schema-style set of minimal pair sentences that differ only by pronoun gender.
8 papers · 0 benchmarks
BC4CHEMD (BioCreative IV Chemical compound and drug name recognition)
Introduced by Krallinger et al.
6 papers · 1 benchmark
DaNE (Danish Dependency Treebank)
Danish Dependency Treebank (DaNE) is a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme.
6 papers · 5 benchmarks
The MEDIA French corpus is dedicated to semantic extraction from speech in a context of human/machine dialogues.
6 papers · 0 benchmarks
Naamapadam is a Named Entity Recognition (NER) dataset for the 11 major Indian languages from two language families.
6 papers · 0 benchmarks
PGR (Phenotype-Gene Relations)
Phenotype-Gene Relations (PGR) is a corpus that consists of 1712 abstracts, 5676 human phenotype annotations, 13835 gene annotations, and 4283 relations.
6 papers · 2 benchmarks
WNUT 2020 (WNUT-2020 Task 1 Overview: Extracting Entities and Relations from Wet Lab Protocols)
The training and development dataset for our task was taken from previous work on wet lab corpus (Kulkarni et al., 2018) that consists of from the 623 protocols.
6 papers · 2 benchmarks
WikiNEuRal is a high-quality automatically-generated dataset for Multilingual Named Entity Recognition.
6 papers · 0 benchmarks
legalNER is a corpus of 46545 annotated legal named entities mapped to 14 legal entity types.
6 papers · 0 benchmarks
AMALGUM (A Machine Annotated Lookalike of GUM)
AMALGUM is a machine annotated multilayer corpus following the same design and annotation layers as GUM, but substantially larger (around 4M tokens).
5 papers · 0 benchmarks
DaN+ is a new multi-domain corpus and annotation guidelines for Danish nested named entities (NEs) and lexical normalization to support research on cross-lingual cross-domain learning for a less-resourced language.
5 papers · 0 benchmarks
Europeana Newspapers consists of four datasets with 100 pages each for the languages Dutch, French, German (including Austrian) as part of the Europeana Newspapers project is expected to contribute to the further development and…
5 papers · 0 benchmarks
LINNAEUS is a general-purpose dictionary matching software, capable of processing multiple types of document formats in the biomedical domain (MEDLINE, PMC, BMC, OTMI, text, etc.).
5 papers · 1 benchmark
SoMeSci (Software Mentions in Scientific Articles)
Knowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling.
5 papers · 0 benchmarks
Species-800 is a corpus for species entities, which is based on manually annotated abstracts.
5 papers · 1 benchmark
The first NER dataset in the field of traffic, which is to extract the characteristics and attributes of the vehicle on the road.
4 papers · 2 benchmarks
Finer (Finnish News Corpus for Named Entity Recognition)
Finnish News Corpus for Named Entity Recognition (Finer) is a corpus that consists of 953 articles (193,742 word tokens) with six named entity classes (organization, location, person, product, event,and date).
4 papers · 0 benchmarks
LegalNERo (Romanian Named Entity Recognition in the Legal domain)
LegalNERo is a manually annotated corpus for named entity recognition in the Romanian legal domain.
4 papers · 1 benchmark
RONEC (Romanian Named Entity Corpus)
Romanian Named Entity Corpus is a named entity corpus for the Romanian language.
4 papers · 0 benchmarks
This dataset contains 1304 de-identified longitudinal medical records describing 296 patients.
4 papers · 1 benchmark
BUSTER (BUSiness Transaction Entity Recognition dataset.)
BUSiness Transaction Entity Recognition dataset.
3 papers · 0 benchmarks
We introduce FUNSD-r and CORD-r in Token Path Prediction, the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.
3 papers · 1 benchmark
COVID-Q consists of COVID-19 questions which have been annotated into a broad category (e.g.
3 papers · 0 benchmarks
DR.BENCH (Diagnostic Reasoning Benchmark for clinical natural language processing)
DR.BENCH is a dataset for developing and evaluating cNLP models with clinical diagnostic reasoning ability.
3 papers · 0 benchmarks
E-NER is a publicly available legal Named Entity Recognition (NER) data set.
3 papers · 0 benchmarks
We introduce FUNSD-r and CORD-r in Token Path Prediction, the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.
3 papers · 1 benchmark
KIND (Kessler Italian Named-entities Dataset)
KIND is an Italian dataset for Named-Entity Recognition.
3 papers · 0 benchmarks
LeNER-Br is a dataset for named entity recognition (NER) in Brazilian Legal Text.
3 papers · 2 benchmarks
Named Entity (NER) annotations of the Hebrew Treebank (Haaretz newspaper) corpus, including: morpheme and token level NER labels, nested mentions, and more.
3 papers · 3 benchmarks
An open, broad-coverage corpus for informal Persian named entity recognition was collected from Twitter.
3 papers · 0 benchmarks
PhoNERCOVID19 is a dataset for recognising COVID-19 related named entities in Vietnamese, consisting of 35K entities over 10K sentences.
3 papers · 1 benchmark
A vast amount of information in the biomedical domain is available as natural language free text.
3 papers · 0 benchmarks
ViMQ is a Vietnamese dataset of medical questions from patients with sentence-level and entity-level annotations for the Intent Classification and Named Entity Recognition tasks.
3 papers · 0 benchmarks
Full-text chemical identification and indexing in PubMed articles.
2 papers · 3 benchmarks
Named entities in Bavarian text Details: Siyao Peng, Zihang Sun, Huangyan Shan, Marie Kolm, Verena Blaschke, Ekaterina Artemova, and Barbara Plank.
2 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
CLUENER2020 is a well-defined fine-grained dataset for named entity recognition in Chinese.
2 papers · 0 benchmarks
Chinese Gigaword corpus consists of 2.2M of headline-document pairs of news stories covering over 284 months from two Chinese newspapers, namely the Xinhua News Agency of China (XIN) and the Central News Agency of Taiwan (CNA).
2 papers · 0 benchmarks
A test dataset that annotated articles in 2020 following the CoNLL-2003 NER task.
2 papers · 1 benchmark
Dataset of Legal Documents consists of court decisions from 2017 and 2018 were selected for the dataset, published online by the Federal Ministry of Justice and Consumer Protection.
2 papers · 0 benchmarks
InLegalNER is a corpus of 46545 annotated legal named entities mapped to 14 legal entity types.
2 papers · 1 benchmark
KazNERD is a dataset for Kazakh named entity recognition.
2 papers · 0 benchmarks
MobIE is a German-language dataset which is human-annotated with 20 coarse- and fine-grained entity types and entity linking information for geographically linkable entities.
2 papers · 0 benchmarks
PcMSP is a dataset annotated from 305 open access scientific articles for material science information extraction that simultaneously contains the synthesis sentences extracted from the experimental paragraphs, as well as the entity…
2 papers · 0 benchmarks
Data annotation The 1,073 full rare disease mention annotations (from 312 MIMIC-III discharge summaries) are in fullsetRDannMIMICIIIdisch.csv.
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.