Browse State-of-the-Art › coreference-resolution
coreference-resolution
215 papers with code · 0 benchmarks · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 0 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| DaCoref (0 rows) | no rows in the archive | — | — | ||
| DaNED (0 rows) | no rows in the archive | — | — | ||
| OntoNotes (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 215 papers with code (594 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
25 Apr 2018 4 repositories listedWe present an empirical study of gender bias in coreference resolution systems.
-
18 Apr 2018 4 repositories listedWe introduce a new benchmark, WinoBias, for coreference resolution focused on gender bias.
-
30 Jun 2021 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedExperiments with pre-trained models such as BERT are often based on a single checkpoint.
-
5 May 2024 2 repositories listedOur contributions are: 1) Created Supervised Fine-Tuning (SFT) training data in alpaca format, along with a set of Low-Rank Adaptation (LoRA) weights, and 2) Developed a method for acquiring high-quality data leveraging…
-
25 May 2022 2 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedWhile coreference resolution typically involves various linguistic challenges, recent models are based on a single pairwise scorer for all types of pairs.
-
3 Apr 2022 2 repositories listedIn this paper, we develop a sequence-to-sequence approach, seq2rel, that can learn the subtasks of DocRE (entity extraction, coreference resolution and relation extraction) end-to-end, replacing a pipeline of…
-
25 Feb 2022 2 repositories listedAnaphoric expressions, such as pronouns and referential descriptions, are situated with respect to the linguistic context of prior turns, as well as, the immediate visual environment.
-
20 Sep 2021 2 repositories listedWhile coreference resolution is defined independently of dataset domain, most models for performing coreference resolution do not transfer well to unseen domains.
-
18 Apr 2021 2 repositories listedDetermining coreference of concept mentions across multiple documents is a fundamental task in natural language understanding.
-
17 Apr 2021 2 repositories listedAcademic neural models for coreference resolution (coref) are typically trained on a single dataset, OntoNotes, and model improvements are benchmarked on that same dataset.
-
11 Apr 2021 2 repositories listedTo complement these resources and enhance future research, we present Wikipedia Event Coreference (WEC), an efficient methodology for gathering a large-scale dataset for cross-document event coreference from Wikipedia,…
-
14 Jan 2021 2 repositories listedWe propose a new framework, Translation between Augmented Natural Languages (TANL), to solve many structured prediction language tasks including joint entity and relation extraction, nested named entity recognition,…
-
26 Sep 2020 2 repositories listed Syntology ran 5 of 14 samples · 9 unverified · 3 pointer-only (licence)Second, the document-level multi-task annotations require the models to transfer information between entity mentions located in different parts of the document, as well as between different tasks, in a joint learning…
-
23 Sep 2020 2 repositories listedRecent evaluation protocols for Cross-document (CD) coreference resolution have often been inconsistent or lenient, leading to incomparable results across works and overestimation of performance.
-
30 Apr 2020 2 repositories listedWe study the potential synergy between two different NLP tasks, both confronting predicate lexical variability: identifying predicate paraphrases, and event coreference resolution.
-
3 Jun 2019 2 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedWe present the first challenge set and evaluation protocol for the analysis of gender bias in machine translation (MT).
-
1 Aug 2018 2 repositories listedTo the best of our knowledge, this is the first time that plural mentions are thoroughly analyzed for these two resolution tasks.
-
16 May 2025 1 repository listed Syntology ran 2 of 6 samples · 4 unverifiedPhrase grounding between images and their captions is a well-established task.
-
18 Apr 2025 1 repository listedCompared to the baseline of unshortened (long) contexts, our experiments on four Indic languages (Hindi, Tamil, Telugu, and Urdu) demonstrate that context-shortening techniques yield an average improvement of 4\% in…
-
14 Apr 2025 1 repository listedWith the rise of knowledge graph based retrieval-augmented generation (RAG) techniques such as GraphRAG and Pike-RAG, the role of knowledge graphs in enhancing the reasoning capabilities of large language models (LLMs)…
-
18 Feb 2025 1 repository listedIn this paper, we present the first dataset for the legal domain, LegalCore, which has been annotated with comprehensive event and event coreference information.
-
16 Dec 2024 1 repository listedThe automatic extraction of character networks from literary texts is generally carried out using natural language processing (NLP) cascading pipelines.
-
12 Nov 2024 1 repository listedThe benchmark also consists of a curated mixture of different mention types and corresponding entities, allowing for a fine-grained analysis of model performance.
-
22 Oct 2024 1 repository listedWhile coreference resolution is traditionally used as a component in individual document understanding, in this work we take a more global view and explore what can we learn about a domain from the set of all…
-
21 Oct 2024 1 repository listedThe paper presents an overview of the third edition of the shared task on multilingual coreference resolution, held as part of the CRAC 2024 workshop.
-
12 Oct 2024 1 repository listedChallenge sets such as the Winograd Schema Challenge (WSC) are used to benchmark systems' ability to resolve ambiguities in natural language.
-
4 Oct 2024 1 repository listedTransformer-based language models have shown an excellent ability to effectively capture and utilize contextual information.
-
3 Oct 2024 1 repository listedIn this third iteration of the shared task, a novel objective is to also predict empty nodes needed for zero coreference mentions (while the empty nodes were given on input in previous years).
-
22 Sep 2024 1 repository listedThis paper explores the challenges posed by nominal adjectives (NAs) in natural language processing (NLP) tasks, particularly in part-of-speech (POS) tagging.
-
9 Sep 2024 1 repository listedWhile measuring bias and robustness in coreference resolution are important goals, such measurements are only as good as the tools we use to measure them.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections