Home › Datasets › task › Nested Named Entity Recognition
Nested Named Entity Recognition datasets
archive 2025-07-28
14 datasets carry the task tag "Nested Named Entity Recognition" (the task itself: Nested Named Entity Recognition), ordered by the archive's paper count. Page 1 of 1: 14 shown of 14. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Nested Named Entity Recognition datasets 1–14 of 14
The GENIA corpus is the primary collection of biomedical literature compiled and annotated within the scope of the GENIA project.
121 papers · 7 benchmarks
ACE 2005 (ACE 2005 Multilingual Training Corpus)
ACE 2005 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2005 Automatic Content Extraction (ACE) technology evaluation.
65 papers · 8 benchmarks
ACE 2004 (ACE 2004 Multilingual Training Corpus)
ACE 2004 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2004 Automatic Content Extraction (ACE) technology evaluation.
51 papers · 6 benchmarks
NNE is a dataset for Nested Named Entity Recognition in English Newswire
19 papers · 1 benchmark
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
AMALGUM (A Machine Annotated Lookalike of GUM)
AMALGUM is a machine annotated multilayer corpus following the same design and annotation layers as GUM, but substantially larger (around 4M tokens).
5 papers · 0 benchmarks
DaN+ is a new multi-domain corpus and annotation guidelines for Danish nested named entities (NEs) and lexical normalization to support research on cross-lingual cross-domain learning for a less-resourced language.
5 papers · 0 benchmarks
The Chilean Waiting List corpus comprises de-identified referrals from the waiting list in Chilean public hospitals.
4 papers · 1 benchmark
LegalNERo (Romanian Named Entity Recognition in the Legal domain)
LegalNERo is a manually annotated corpus for named entity recognition in the Romanian legal domain.
4 papers · 1 benchmark
Named Entity (NER) annotations of the Hebrew Treebank (Haaretz newspaper) corpus, including: morpheme and token level NER labels, nested mentions, and more.
3 papers · 3 benchmarks
A vast amount of information in the biomedical domain is available as natural language free text.
3 papers · 0 benchmarks
CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024…
1 paper · 2 benchmarks
CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024…
1 paper · 2 benchmarks
CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024…
1 paper · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.