Home › Datasets › task › Named Entity Recognition (NER)
Named Entity Recognition (NER) datasets
archive 2025-07-28
130 datasets carry the task tag "Named Entity Recognition (NER)" (the task itself: Named Entity Recognition (NER)), ordered by the archive's paper count. Page 3 of 3: 34 shown of 130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Named Entity Recognition (NER) datasets 97–130 of 130
SemClinBr (A multi‑institutional and multi‑specialty semantically annotated corpus for Portuguese clinical NLP tasks)
Background: The high volume of research focusing on extracting patient information from electronic health records (EHRs) has led to an increase in the demand for annotated corpora, which are a precious resource for both the development and…
2 papers · 1 benchmark
The pioNER corpus provides gold-standard and automatically generated named-entity datasets for the Armenian language.
2 papers · 0 benchmarks
Digital Edition: Essays from Hannah Arendt We have created a NER dataset from the digital edition "Sechs Essays" by Hannah Arendt.
1 paper · 0 benchmarks
The dataset contains a total of 253,070 records, with 18 features.
1 paper · 0 benchmarks
Business license datasets and source code for named entity recognition.
1 paper · 0 benchmarks
The CareerCoach 2022 gold standard is available for download in the NIF and JSON format, and draws upon documents from a corpus of over 99,000 education courses which have been retrieved from 488 different education providers.
1 paper · 0 benchmarks
The dataset contains two few-shot chemical fine-grained entity extraction datasets, based on human-annotated ChemNER+ and CHEMET.
1 paper · 0 benchmarks
Chinese Literature NER RE is a Discourse-Level Named Entity Recognition and Relation Extraction Dataset for Chinese Literature Text.
1 paper · 0 benchmarks
DiaKG is a high-quality Chinese dataset for Diabetes knowledge graph.
1 paper · 0 benchmarks
Dataset Card for ESG/DLT Named Entity Recognition Dataset This dataset contains named entities related to Distributed Ledger Technology (DLT) and Environmental, Social, and Governance (ESG) topics created to support research in these areas…
1 paper · 0 benchmarks
Financial Language Understanding Evaluation is an open-source comprehensive suite of benchmarks for the financial domain.
1 paper · 0 benchmarks
HAREM, an initiative by Linguateca, boasts a Golden Collection—a meticulously curated repository of annotated Portuguese texts.
1 paper · 0 benchmarks
A scholarly named entity recognition dataset with focus on machine learning models and datasets.
1 paper · 0 benchmarks
HengamCopus is a Persian corpus with temporal tags (BIO standard tagging scheme).
1 paper · 1 benchmark
This dataset releases a significantly sized standard-abiding Hindi NER dataset containing 109,146 sentences and 2,220,856 tokens, annotated with 3 collapsed tags (PER, LOC, ORG).
1 paper · 1 benchmark
HiNER-original (HiNER: A Large Hindi Named Entity Recognition Dataset)
This dataset releases a significantly sized standard-abiding Hindi NER dataset containing 109,146 sentences and 2,220,856 tokens, annotated with 11 tags.
1 paper · 1 benchmark
The dataset is taken from the First shared task on Information Extractor for Conversational Systems in Indian Languages (IECSIL) .
1 paper · 1 benchmark
LPSC (Planetary Science Data Set)
This data set contains annotated text versions of 1635 two-page abstracts published at the Lunar and Planetary Science Conference from 1998 to 2020 of relevance to four Mars missions.
1 paper · 2 benchmarks
MASC (Manually Annotated Sub-Corpus)
The Manually Annotated Sub-Corpus (MASC) consists of approximately 500,000 words of contemporary American English written and spoken data drawn from the Open American National Corpus (OANC).
1 paper · 0 benchmarks
MSNER (Multilingual Spoken Named Entity Recognition)
This dataset contains named entities annotations for European Parliament recordings in Dutch, French, German and Spanish.
1 paper · 0 benchmarks
Medical Case Report Corpus is a new corpus comprising annotations of medical entities in case reports, originating from PubMed Central's open access library.
1 paper · 0 benchmarks
The MiniHAREM, a reiteration of the 2005 evaluation, used the same methodology and platform.
1 paper · 0 benchmarks
ScienceExamCER is a collection of resources for studying explanation-centered inference, including explanation graphs for 1,680 questions, with 4,950 tablestore rows, and other analyses of the knowledge required to answer elementary and…
1 paper · 0 benchmarks
Digital Edition: Sturm Edition Source: Schrade, Torsten: „Startseite“, in: DER STURM.
1 paper · 0 benchmarks
SumeCzech-NER contains named entity annotations of SumeCzech 1.0, a Czech news-based summarization dataset.
1 paper · 0 benchmarks
TASTEset Recipe Dataset and Food Entities Recognition is a dataset for Named Entity Recognition (NER) which consists of 700 recipes with more than 13,000 entities to extract.
1 paper · 0 benchmarks
Twitter Cyberthreat Detection Dataset is a dataset that contains tweets from two sets of accounts related to cybersecurity.
1 paper · 0 benchmarks
UNER v1 adds an NER annotation layer to 18 datasets (primarily treebanks from UD) and covers 12 geneologically and ty- pologically diverse languages: Cebuano, Danish, German, English, Croatian, Portuguese, Russian, Slovak, Serbian,…
1 paper · 31 benchmarks
Spoken Named Entity Recognition (NER) aims to extracting named entities from speech and categorizing them into types like person, location, organization, etc.
1 paper · 0 benchmarks
POPCORN (POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks)
POPCORN is a French dataset consisting of 400 validation texts and 400 training texts, all written and annotated manually.
0 papers · 0 benchmarks
This dataset was taken from the SIGARRA information system at the University of Porto (UP).
0 papers · 0 benchmarks
Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources The STEM ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity…
0 papers · 0 benchmarks
The Second HAREM was an evaluation exercise in Portuguese Named Entity Recognition.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.