Home › Datasets › task › Relation Classification

Relation Classification datasets

archive 2025-07-28

23 datasets carry the task tag "Relation Classification" (the task itself: Relation Classification), ordered by the archive's paper count. Page 1 of 1: 23 shown of 23. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Relation Classification datasets 1–23 of 23

TACRED (The TAC Relation Extraction Dataset)
TACRED is a large-scale relation extraction dataset with 106,264 examples built over newswire and web text from the corpus used in the yearly TAC Knowledge Base Population (TAC KBP) challenges.
204 papers · 2 benchmarks
FewRel (Few-Shot Relation Classification Dataset)
The FewRel (Few-Shot Relation Classification Dataset) contains 100 relations and 70,000 instances from Wikipedia.
189 papers · 3 benchmarks
A more challenging task to investigate two aspects of few-shot relation classification models: (1) Can they adapt to a new domain with only a handful of instances?
38 papers · 0 benchmarks
MATRES (Multi-Axis Temporal RElations for Start-points)
This is the Multi-Axis Temporal RElations for Start-points (i.e., MATRES) dataset
16 papers · 2 benchmarks
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
CDCP (Cornell eRulemaking Corpus)
The Cornell eRulemaking Corpus – CDCP is an argument mining corpus annotated with argumentative structure information capturing the evaluability of arguments.
12 papers · 3 benchmarks
The Discovery datasets consists of adjacent sentence pairs (s1,s2) with a discourse marker (y) that occurred at the beginning of s2.
10 papers · 1 benchmark
The TACRED-Revisited dataset improves the crowd-sourced TACRED dataset for relation extraction by relabeling the dev and test sets using expert linguistic annotators.
7 papers · 1 benchmark
RELX is a benchmark dataset for cross-lingual relation classification in English, French, German, Spanish and Turkish.
6 papers · 0 benchmarks
CrossRE is a cross-domain benchmark for Relation Extraction (RE), which comprises six distinct text domains and includes multi-label annotations.
5 papers · 0 benchmarks
LabPics (LabPics Dataset for computer vision for autonomous chemistry labs and medical labs)
LabPics Chemistry Dataset Dataset for computer vision for materials segmentation and classification in chemistry labs, medical labs, and any setting where materials are handled inside containers.
5 papers · 0 benchmarks
DISRPT2021 (DISRPT2021 shared task on Discourse Unit Segmentation, Connective Detection and Discourse Relation Classification)
The DISRPT 2021 shared task, co-located with CODI 2021 at EMNLP, introduces the second iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the first iteration of a cross-formalism…
3 papers · 0 benchmarks
FREDo is a Few-Shot Document-Level Relation Extraction Benchmark based on DocRED and SciERC.
3 papers · 2 benchmarks
SupplyGraph (SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural Networks)
Graph Neural Networks (GNNs) have gained traction across different domains such as transportation, bio-informatics, language processing, and computer vision.
3 papers · 0 benchmarks
CORE (Company Relation Extraction)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
DRI Corpus (Dr. Inventor Multi-layer Scientific Corpus)
The Dr.
2 papers · 2 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
PcMSP is a dataset annotated from 305 open access scientific articles for material science information extraction that simultaneously contains the synthesis sentences extracted from the experimental paragraphs, as well as the entity…
2 papers · 0 benchmarks
The AbstRCT dataset consists of randomized controlled trials retrieved from the MEDLINE database via PubMed search.
1 paper · 2 benchmarks
The relational pattern similarity dataset is a new dataset upon the work of Zeichner et al.
1 paper · 0 benchmarks
533 parallel examples sampled from TACRED, translated into Russian and Korean (and 3 additional examples in Russian), accompanied with tranlsation of a list of trigger words collected for the different relations.
1 paper · 0 benchmarks
TFH_Annotated_Dataset (Thin_Film_head_relevant_Patent_Annotated_Dataset)
Dataset Introduction TFHAnnotatedDataset is an annotated patent dataset pertaining to thin film head technology in hard-disk.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.