Home › Datasets › task › NER

NER datasets

archive 2025-07-28

27 datasets carry the task tag "NER" (the task itself: NER), ordered by the archive's paper count. Page 1 of 1: 27 shown of 27. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

NER datasets 1–27 of 27

The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
CoNLL-2003 is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition.
755 papers · 6 benchmarks
MSRA CN NER (MSRA CN NER Dataset)
Simplified Chinese dataset for NER in The Third International Chinese Language Processing Bakeoff (2006), provided by Microsoft Research Asia (MSRA).
23 papers · 3 benchmarks
Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports.
14 papers · 3 benchmarks
BC4CHEMD (BioCreative IV Chemical compound and drug name recognition)
Introduced by Krallinger et al.
6 papers · 1 benchmark
DaNE (Danish Dependency Treebank)
Danish Dependency Treebank (DaNE) is a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme.
6 papers · 5 benchmarks
fine-grained location names extraction from disaster-related tweets
5 papers · 1 benchmark
FiNER-139 is comprised of 1.1M sentences annotated with eXtensive Business Reporting Language (XBRL) tags extracted from annual and quarterly reports of publicly-traded companies in the US.
4 papers · 0 benchmarks
The dataset used to pre-train NuNER from the NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data Contains AI-extracted entities, their concepts, and descriptions from the given text
4 papers · 0 benchmarks
A growing number of papers are published in the area of superconducting materials science.
NER
4 papers · 1 benchmark
We present the development of a Named Entity Recognition (NER) dataset for Tagalog.
3 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
FIJO (French Insurance Job Offer dataset)
This dataset was collected as part of the multidisciplinary project Femmes face aux défis de la transformation numérique : une étude de cas dans le secteur des assurances (Women Facing the Challenges of Digital Transformation: A Case Study…
2 papers · 0 benchmarks
InLegalNER is a corpus of 46545 annotated legal named entities mapped to 14 legal entity types.
2 papers · 1 benchmark
LEPISZCZE is an open-source comprehensive benchmark for Polish NLP and a continuous-submission leaderboard, concentrating public Polish datasets (existing and new) in specific tasks.
2 papers · 0 benchmarks
AISECKG (AISecKG: Knowledge Graph Dataset for Cybersecurity Education)
Cybersecurity education is exceptionally challenging as it involves learning the complex attacks; tools and developing critical problem-solving skills to defend the systems.
1 paper · 0 benchmarks
First HAREM (Primeiro HAREM)
HAREM, an initiative by Linguateca, boasts a Golden Collection—a meticulously curated repository of annotated Portuguese texts.
1 paper · 0 benchmarks
The dataset is taken from the First shared task on Information Extractor for Conversational Systems in Indian Languages (IECSIL) .
1 paper · 1 benchmark
The MiniHAREM, a reiteration of the 2005 evaluation, used the same methodology and platform.
1 paper · 0 benchmarks
RaTE-NER dataset is a large-scale, radiological named entity recognition (NER) dataset, including 13,235 manually annotated sentences from 1,816 reports within the MIMIC-IV database, that spans 9 imaging modalities and 23 anatomical…
1 paper · 0 benchmarks
TASTEset Recipe Dataset and Food Entities Recognition is a dataset for Named Entity Recognition (NER) which consists of 700 recipes with more than 13,000 entities to extract.
1 paper · 0 benchmarks
The EMBO SourceData-NLP dataset (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process.
1 paper · 1 benchmark
Based on RADDLE and SNIPS , we construct Noise-SF, which includes two different perturbation settings.
0 papers · 0 benchmarks
POPCORN (POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks)
POPCORN is a French dataset consisting of 400 validation texts and 400 training texts, all written and annotated manually.
0 papers · 0 benchmarks
This dataset was taken from the SIGARRA information system at the University of Porto (UP).
0 papers · 0 benchmarks
Second HAREM (Segundo HAREM)
The Second HAREM was an evaluation exercise in Portuguese Named Entity Recognition.
0 papers · 0 benchmarks
SourceData-NLP (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
Introduction: The scientific publishing landscape is expanding rapidly, creating challenges for researchers to stay up-to-date with the evolution of the literature.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.