Home › Datasets › task › UIE

UIE datasets

archive 2025-07-28

13 datasets carry the task tag "UIE" (the task itself: UIE), ordered by the archive's paper count. Page 1 of 1: 13 shown of 13. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

UIE datasets 1–13 of 13

CoNLL-2003 is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition.
755 papers · 6 benchmarks
The New York Times Annotated Corpus contains over 1.8 million articles written and published by the New York Times between January 1, 1987 and June 19, 2007 with article metadata provided by the New York Times Newsroom, the New York Times…
262 papers · 9 benchmarks
OntoNotes 5.0 is a large corpus comprising various genres of text (news, conversational telephone speech, weblogs, usenet newsgroups, broadcast, talk shows) in three languages (English, Chinese, and Arabic) with structural information…
254 papers · 12 benchmarks
BC5CDR (BioCreative V CDR corpus)
BC5CDR corpus consists of 1500 PubMed articles with 4409 annotated chemicals, 5818 diseases and 3116 chemical-disease interactions.
191 papers · 4 benchmarks
The NCBI Disease corpus consists of 793 PubMed abstracts, which are separated into training (593), development (100) and test (100) subsets.
154 papers · 3 benchmarks
SciERC dataset is a collection of 500 scientific abstract annotated with scientific entities, their relations, and coreference clusters.
134 papers · 7 benchmarks
WNUT 2017 (WNUT 2017 Emerging and Rare entity recognition)
This shared task focuses on identifying unusual, previously-unseen entities in the context of emerging discussions.
127 papers · 2 benchmarks
The GENIA corpus is the primary collection of biomedical literature compiled and annotated within the scope of the GENIA project.
121 papers · 7 benchmarks
WikiANN (PAN-X)
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
ACE 2004 (ACE 2004 Multilingual Training Corpus)
ACE 2004 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2004 Automatic Content Extraction (ACE) technology evaluation.
51 papers · 6 benchmarks
Created by Smith et al.
13 papers · 2 benchmarks
The first NER dataset in the field of traffic, which is to extract the characteristics and attributes of the vehicle on the road.
4 papers · 2 benchmarks
SUIM-E (SGUIE-Net: Semantic attention guided underwater image enhancement with multi-scale perception)
Underwater Image Enhancement Dataset
UIE
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.