Home › Datasets › task › Relation Extraction
Relation Extraction datasets
archive 2025-07-28
82 datasets carry the task tag "Relation Extraction" (the task itself: Relation Extraction), ordered by the archive's paper count. Page 2 of 2: 34 shown of 82. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Relation Extraction datasets 49–82 of 82
a dataset from A Hierarchical Framework for Relation Extraction with Reinforcement Learning
4 papers · 1 benchmark
Persian dataset for relation extraction, which is an expert-translated version of the "Semeval-2010-Task-8" dataset.
4 papers · 0 benchmarks
TimeBankPT is a corpus of Portuguese text with annotations about time.
4 papers · 1 benchmark
The FB15k-237-low dataset is a variation of the FB15k-237 dataset where relations with a low number of triplets are kept.
3 papers · 0 benchmarks
FOBIE (Focused Open Biological Information Extraction)
The Focused Open Biology Information Extraction (FOBIE) dataset aims to support IE from Computer-Aided Biomimetics.
3 papers · 0 benchmarks
FREDo is a Few-Shot Document-Level Relation Extraction Benchmark based on DocRED and SciERC.
3 papers · 2 benchmarks
ROOR is a reading order prediction (ROP) benchmark which annotates layout reading order as ordering relations.
3 papers · 1 benchmark
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
CORE (Company Relation Extraction)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
MobIE is a German-language dataset which is human-annotated with 20 coarse- and fine-grained entity types and entity linking information for geographically linkable entities.
2 papers · 0 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
NYT-H is a dataset for distantly-supervised relation extraction, in which DS-labelled training data is used and several annotators to label test data are hired.
2 papers · 0 benchmarks
A human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems.
2 papers · 0 benchmarks
https://github.com/dialogue-evaluation/RuSentNE-evaluation
2 papers · 0 benchmarks
X-WikiRE is a new, large-scale multilingual relation extraction dataset in which relation extraction is framed as a problem of reading comprehension to allow for generalization to unseen relations.
2 papers · 0 benchmarks
Chinese Literature NER RE is a Discourse-Level Named Entity Recognition and Relation Extraction Dataset for Chinese Literature Text.
1 paper · 0 benchmarks
This is the dataset used for classifying Gene-Disease relationship types from sentences.
1 paper · 1 benchmark
DiaKG is a high-quality Chinese dataset for Diabetes knowledge graph.
1 paper · 0 benchmarks
The FB1.5M dataset is a benchmark for Knowledge Graph Completion.
1 paper · 0 benchmarks
FinDKG: The Global Financial Dynamic Knowledge Graph Dataset FinDKG is an open-source dataset focused on creating a temporally-resolved Financial Dynamic Knowledge Graph.
1 paper · 0 benchmarks
KGRED (Knowledge-graph-enhanced relation extraction datasets--)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LPSC (Planetary Science Data Set)
This data set contains annotated text versions of 1635 two-page abstracts published at the Lunar and Planetary Science Conference from 1998 to 2020 of relevance to four Mars missions.
1 paper · 2 benchmarks
Medical Case Report Corpus is a new corpus comprising annotations of medical entities in case reports, originating from PubMed Central's open access library.
1 paper · 0 benchmarks
Multi-CrossRE is a broadest multi-lingual dataset for Relation Extraction (RE) including 26 languages in addition to English, and covering six text domains.
1 paper · 0 benchmarks
The Part-Whole Relations dataset is a dataset of semantic relations between entities.
1 paper · 0 benchmarks
The corpus contains review sentences mostly of products in electronics domain, annotated and segregated into 4 comparison categories.
1 paper · 1 benchmark
SOMD (SOftware Mention Detection)
The dataset contains the training and test data for the SOftware Mention Detection challenge.
1 paper · 0 benchmarks
THRED (Two-Hop Relation Extraction Dataset)
This is two-hop relation extraction dataset derived from WikiHop dataset [1].
1 paper · 0 benchmarks
Green family of datasets for emergent communications on relations.
1 paper · 0 benchmarks
533 parallel examples sampled from TACRED, translated into Russian and Korean (and 3 additional examples in Russian), accompanied with tranlsation of a list of trigger words collected for the different relations.
1 paper · 0 benchmarks
TurkQA consists of a selection of sentences from English Wikipedia articles, with questions and answers crowdsourced from workers on Amazon Mechanical Turk.
1 paper · 0 benchmarks
ARF (Artificial Relationships in Fiction)
Artificial Relationships in Fiction Dataset Description Artificial Relationships in Fiction (ARF) is a synthetically annotated dataset for Relation Extraction (RE) in fiction, created from a curated selection of literary texts sourced from…
0 papers · 0 benchmarks
POPCORN (POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks)
POPCORN is a French dataset consisting of 400 validation texts and 400 training texts, all written and annotated manually.
0 papers · 0 benchmarks
Dataset Introduction TFHAnnotatedDataset is an annotated patent dataset pertaining to thin film head technology in hard-disk.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.