Home › Datasets › task › Coreference Resolution

Coreference Resolution datasets

archive 2025-07-28

45 datasets carry the task tag "Coreference Resolution" (the task itself: Coreference Resolution), ordered by the archive's paper count. Page 1 of 1: 45 shown of 45. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Coreference Resolution datasets 1–45 of 45

WSC (Winograd Schema Challenge)
The Winograd Schema Challenge was introduced both as an alternative to the Turing Test and as a test of a system’s ability to do commonsense reasoning.
361 papers · 2 benchmarks
OntoNotes 5.0 is a large corpus comprising various genres of text (news, conversational telephone speech, weblogs, usenet newsgroups, broadcast, talk shows) in three languages (English, Chinese, and Arabic) with structural information…
254 papers · 12 benchmarks
The CoNLL dataset is a widely used resource in the field of natural language processing (NLP).
187 papers · 35 benchmarks
WinoBias contains 3,160 sentences, split equally for development and test, created by researchers familiar with the project.
134 papers · 0 benchmarks
GAP is a gender-balanced dataset containing 8,908 coreference-labeled pairs of (ambiguous pronoun, antecedent name), sampled from Wikipedia and released by Google AI Language for the evaluation of coreference resolution in practical…
104 papers · 0 benchmarks
The CoNLL-2012 shared task involved predicting coreference in English, Chinese, and Arabic, using the final version, v5.0, of the OntoNotes corpus.
89 papers · 0 benchmarks
ECB+ (extension to the EventCorefBank)
The ECB+ corpus is an extension to the EventCorefBank (ECB, Bejan and Harabagiu, 2010).
76 papers · 0 benchmarks
GAP (GAP Benchmark Suite)
GAP is a graph processing benchmark suite with the goal of helping to standardize graph processing evaluations.
60 papers · 1 benchmark
Quoref is a QA dataset which tests the coreferential reasoning capability of reading comprehension systems.
50 papers · 0 benchmarks
English Web Treebank is a dataset containing 254,830 word-level tokens and 16,624 sentence-level tokens of webtext in 1174 files annotated for sentence- and word-level tokenization, part-of-speech, and syntactic structure.
42 papers · 0 benchmarks
xP3 is a multilingual dataset for multitask prompted finetuning.
34 papers · 0 benchmarks
WikiCoref is an English corpus annotated for anaphoric relations, where all documents are from the English version of Wikipedia.
28 papers · 1 benchmark
LitBank is an annotated dataset of 100 works of English-language fiction to support tasks in natural language processing and the computational humanities, described in more detail in the following publications: - David Bamman, Sejal Popat…
23 papers · 1 benchmark
A large-scale English dataset for coreference resolution.
20 papers · 1 benchmark
DWIE (Deutsche Welle corpus for Information Extraction)
The 'Deutsche Welle corpus for Information Extraction' (DWIE) is a multi-task dataset that combines four main Information Extraction (IE) annotation sub-tasks: (i) Named Entity Recognition (NER), (ii) Coreference Resolution, (iii) Relation…
18 papers · 5 benchmarks
MAP (Maybe Ambiguous Pronoun)
Maybe Ambiguous Pronoun is a dataset similar to GAP dataset, but without binary gender constraints.
14 papers · 0 benchmarks
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
ParCorFull (Parallel Corpus Annotated with Full Coreference)
ParCorFull is a parallel corpus annotated with full coreference chains that has been created to address an important problem that machine translation and other multilingual natural language processing (NLP) technologies face -- translation…
11 papers · 0 benchmarks
A large-scale collection of visually-grounded, task-oriented dialogues in English designed to investigate shared dialogue history accumulating during conversation.
11 papers · 0 benchmarks
CLEVR-Dialog is a large diagnostic dataset for studying multi-round reasoning in visual dialog.
10 papers · 0 benchmarks
OntoGUM is an OntoNotes-like coreference dataset converted from GUM, an English corpus covering 12 genres using deterministic rules.
9 papers · 1 benchmark
Consists of multiple sentences whose clues are arranged by difficulty (from obscure to obvious) and uniquely identify a well-known entity such as those found on Wikipedia.
8 papers · 1 benchmark
Winogender Schemas is a novel, Winograd schema-style set of minimal pair sentences that differ only by pronoun gender.
8 papers · 0 benchmarks
Composes sentence pairs (i.e., twin sentences).
7 papers · 0 benchmarks
BiPaR is a manually annotated bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support monolingual, multilingual and cross-lingual reading comprehension on novels.
6 papers · 0 benchmarks
GICoref (Gender Inclusive Coreference)
GICoref is a fully annotated coreference resolution dataset written by and about trans people.
6 papers · 0 benchmarks
VisPro dataset contains coreference annotation of 29,722 pronouns from 5,000 dialogues.
6 papers · 0 benchmarks
An unsupervised dataset for co-reference resolution.
6 papers · 0 benchmarks
AMALGUM (A Machine Annotated Lookalike of GUM)
AMALGUM is a machine annotated multilayer corpus following the same design and annotation layers as GUM, but substantially larger (around 4M tokens).
5 papers · 0 benchmarks
A large-scale evaluation set that provides human ratings for the plausibility of 10,000 SP pairs over five SP relations, covering 2,500 most frequent verbs, nouns, and adjectives in American English.
5 papers · 0 benchmarks
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
MultiReQA is a cross-domain evaluation for retrieval question answering models.
2 papers · 0 benchmarks
PoC (Points of correspondence)
A dataset containing the documents, source and fusion sentences, and human annotations of points of correspondence between sentences.
2 papers · 0 benchmarks
Comet is a dataset which contains 11.5k user-assistant dialogs (totalling 103k utterances), grounded in simulated personal memory graphs.
1 paper · 0 benchmarks
The DocRED Information Extraction (DocRED-IE) dataset extends the DocRED dataset for the Document-level Closed Information Extraction (DocIE) task.
1 paper · 6 benchmarks
MASC (Manually Annotated Sub-Corpus)
The Manually Annotated Sub-Corpus (MASC) consists of approximately 500,000 words of contemporary American English written and spoken data drawn from the Open American National Corpus (OANC).
1 paper · 0 benchmarks
Describe the Marmara Turkish Coreference Corpus, which is an annotation of the whole METU-Sabanci Turkish Treebank with mentions and coreference chains.
1 paper · 0 benchmarks
MuDoCo_QueryRewrite (The MuDoCo dataset with Query Rewrite Annotations)
Given an ongoing dialogue between a user and a dialogue assistant, for the user query, the model is required to predict both coreference links between the query and the dialogue context, and the self-contained rewritten user query that is…
1 paper · 0 benchmarks
SciCo (Scientific Concept Induction Corpus)
SciCo is an expert-annotated dataset for hierarchical CDCR (cross-document coreference resolution) for concepts in scientific papers, with the goal of jointly inferring coreference clusters and hierarchy between them.
1 paper · 0 benchmarks
This dataset consists of Winograd schemas that test coreference resolution systems' ability to differentiate singular vs plural they/them pronouns.
1 paper · 0 benchmarks
WinoPron is a novel dataset of Winogender-like template pairs in English, which fixes inconsistencies in Winogender Schemas and contains balanced template pairs for pronoun forms in 3 grammatical cases, which we find impacts performance…
1 paper · 0 benchmarks
ContraCAT (Contrastive Coreference Analytical Templates (for Machine Translation))
Current approaches to context-aware MT rely on a set of surface heuristics to translate pronouns, which break down when translations require real reasoning.
0 papers · 0 benchmarks
POPCORN (POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks)
POPCORN is a French dataset consisting of 400 validation texts and 400 training texts, all written and annotated manually.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.