Datasets › DocRED
DocRED
DocRED (Document-Level Relation Extraction Dataset) is a relation extraction dataset constructed from Wikipedia and Wikidata. Each document in the dataset is human-annotated with named entity mentions, coreference information, intra- and inter-sentence relations, and supporting evidence. DocRED requires reading multiple sentences in a document to extract entities and infer their relations by synthesizing all information of the document. Along with the human-annotated data, the dataset provides large-scale distantly supervised data.
DocRED contains 132,375 entities and 56,354 relational facts annotated on 5,053 Wikipedia documents. In addition to the human-annotated data, the dataset provides large-scale distantly supervised data over 101,873 documents.
Source: DocRED: A Large-Scale Document-Level Relation Extraction Dataset Image Source: DocRED: A Large-Scale Document-Level Relation Extraction Dataset
Benchmarks archive 2025-07-28
All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Relation Extraction | DocRED | DREEAM F1 67.53 | DREEAM: Guiding Attention with Evidence for Improving... | youmima/dreeam | 62 | Compare |
| Joint Entity and Relation Extraction | DocRED | REBEL+pretraining Relation F1 47.1 | REBEL: Relation Extraction By End-to-end Language generation | Babelscape/rebel | 6 | Compare |
| Document-level Closed Information Extraction | DocRED | REXEL Relation F1 27.96 | REXEL: An End-to-end Model for Document-Level Relation... | amazon-science/e2e-docie | 1 | Compare |
| Few-Shot Relation Classification | DocRED | DL-MNAV F1 (1-Doc) 7.05 | Few-Shot Document-Level Relation Extraction | nicpopovic/fredo | 1 | Compare |
Papers archive 2025-07-28
30 shown of 43 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 155. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
The full list of 43 is in the JSON twin.
Dataset loaders archive 2025-07-28
3 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Unknown
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- DocRED
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections