Browse › Natural Language Processing › Coreference Resolution › Winograd Schema Challenge

Coreference Resolution archive 2025-07-28

Winograd Schema Challenge Benchmark (Coreference Resolution)

82 rows 64 with code listed 1 metric Dataset page

Coreference resolution is the task of clustering mentions in text that refer to the same underlying real world entities.

Example:

               +-----------+
               |           |
I voted for Obama because he was most aligned with my values", she said.
 |                                                 |            |
 +-------------------------------------------------+------------+

"I", "my", and "she" belong to the same cluster and "Obama" and "he" belong to the same cluster.

The archive carries no text for this table; the description above is the archive's text for the task Coreference Resolution. archive 2025-07-28

Over time archive 2025-07-28

The chart needs JavaScript; the table below carries every value.

Direction inferred from the metric name, not from the archive: Accuracy (higher is better). Points are placed at the row's paper date; 82 of 82 rows carry one.

Results archive 2025-07-28

Archive rows end at the archive snapshot, 2025-07-28: no result published after that date is in this table. Rank is the archive's row order at that snapshot; not re-ranked here. Metric values are the archive's strings. Column headers sort the table in your browser; each row keeps its archive rank.

Paper Code Ran Syntology Report
1 PaLM 540B (fine-tuned) 100 – Paper Code 2022 30 of 37 ran · 7 unverified report
2 Vega v2 6B (KD-based prompt transfer) 98.6 – Paper – 2022 no code linked report
3 UL2 20B (fine-tuned) 98.1 – Paper Code 2022 0 of 16 ran · 16 unverified report
4 Turing NLR v5 XXL 5.4B (fine-tuned) 97.3 – Paper – 2022 no code linked report
5 ST-MoE-32B 269B (fine-tuned) 96.6 – Paper Code 2022 5 of 5 ran · 0 unverified report
6 DeBERTa-1.5B 95.9 – Paper Code 2020 4 of 13 ran · 9 unverified report
7 T5-XXL 11B (fine-tuned) 93.8 – Paper Code 2019 2 of 31 ran · 29 unverified report
8 ST-MoE-L 4.1B (fine-tuned) 93.3 – Paper Code 2022 5 of 5 ran · 0 unverified report
9 RoBERTa-WinoGrande 355M 90.1 – Paper Code 2019 linked, not harvested report
10 Flan-T5 XXL (zero -shot) 89.82 – Paper Code 2022 8 of 17 ran · 9 unverified report
11 PaLM 540B (5-shot) 89.5 – Paper Code 2022 30 of 37 ran · 7 unverified report
12 PaLM 540B (0-shot) 89.1 – Paper Code 2022 30 of 37 ran · 7 unverified report
13 PaLM 2-M (1-shot) 88.1 – Paper Code 2023 linked, not harvested report
14 PaLM 2-L (1-shot) 86.9 – Paper Code 2023 linked, not harvested report
15 FLAN 137B (prompt-tuned) 86.5 – Paper Code 2021 0 of 1 ran · 1 unverified report
16 PaLM 540B (1-shot) 86.3 – Paper Code 2022 30 of 37 ran · 7 unverified report
17 TTTTT 3B (fine-tuned) 84.6 – Paper – 2020 no code linked report
18 PaLM 2-S (1-shot) 84.6 – Paper Code 2023 linked, not harvested report
19 RoBERTa-DPR 355M 83.1 – Paper Code 2019 linked, not harvested report
20 FLAN 137B (zero-shot) 80.8 – Paper Code 2021 0 of 1 ran · 1 unverified report
21 GPT-3 175B (few-shot) 80.1 – Paper Code 2020 15 of 65 ran · 50 unverified report
22 RoBERTa-large + G-DAug-Inf 80 – Paper Code 2020 2 of 4 ran · 2 unverified report
23 UL2 20B (0-shot) 79.9 – Paper Code 2022 0 of 16 ran · 16 unverified report
24 ALBERT-xxlarge 235M 78.8 – Paper – 2021 no code linked report
25 Neo-6B (QA + WS) 77.9 – Paper Code 2022 2 of 2 ran · 0 unverified report
26 HNN 75.1 – Paper Code 2019 0 of 1 ran · 1 unverified report
27 Neo-6B (QA) 74.7 – Paper Code 2022 2 of 2 ran · 0 unverified report
28 RoBERTa-large 354M 73.9 – Paper – 2021 no code linked report
29 GPT-2-XL 1.5B 73.3 – Paper Code 2023 linked, not harvested report
30 BERTwiki 340M (fine-tuned on WSCR) 72.5 – Paper Code 2019 linked, not harvested report
31 BERT-SocialIQA 340M 72.5 – Paper Code 2019 0 of 10 ran · 10 unverified report
32 BERT-large 340M (fine-tuned on WSCR) 71.4 – Paper Code 2019 linked, not harvested report
33 GPT-2-XL 1.5B 70.7 – Paper Code 2019 linked, not harvested report
34 BERTwiki 340M (fine-tuned on half of WSCR) 70.3 – Paper Code 2019 linked, not harvested report
35 LaMini-GPT 1.5B 69.6 – Paper Code 2023 linked, not harvested report
36 GPT-2 Medium 774M (partial scoring) 69.2 – Paper Code 2018 linked, not harvested report
37 N-Grammer 343M 68.3 – Paper Code 2022 0 of 6 ran · 6 unverified report
38 AlexaTM 20B 68.3 – Paper Code 2022 1 of 1 ran · 0 unverified report
39 BERT-large 340M 67 – Paper Code 2019 0 of 10 ran · 10 unverified report
40 T5-Large 738M 66.7 – Paper Code 2023 linked, not harvested report
41 T0-3B (CoT fine-tuned) 66 – Paper Code 2023 linked, not harvested report
42 KiC-770M 65.40 – Paper – 2022 no code linked report
43 GPT-2 Medium 774M (full scoring) 64.5 – Paper Code 2018 linked, not harvested report
44 LaMini-F-T5 783M 64.1 – Paper Code 2023 linked, not harvested report
45 Ensemble of 14 LMs 63.7 – Paper Code 2018 0 of 7 ran · 7 unverified report
46 H3 125M (3-shot, rank classification) 63.5 – Paper Code 2022 7 of 15 ran · 8 unverified report
47 DSSM 63.0 – Paper – 2019 no code linked report
48 RoBERTa-base 125M 63 – Paper – 2021 no code linked report
49 Word-level CNN+LSTM (partial scoring) 62.6 – Paper Code 2018 0 of 7 ran · 7 unverified report
50 UDSSM-II (ensemble) 62.4 – Paper – 2019 no code linked report
51 BERT-base 110M (fine-tuned on WSCR) 62.3 – Paper Code 2019 linked, not harvested report
52 RoE-3B 62.21 – Paper Code 2023 3 of 3 ran · 0 unverified report
53 BERT-large 340M 62.0 – Paper Code 2018 204 of 659 ran · 455 unverified report
54 GPT-2 Small 117M (partial scoring) 61.5 – Paper Code 2018 linked, not harvested report
55 H3 125M (0-shot, rank classification) 61.5 – Paper Code 2022 7 of 15 ran · 8 unverified report
56 BERT-large 340M 61.4 – Paper – 2021 no code linked report
57 BERT-base 110M + MAS 60.3 – Paper Code 2019 linked, not harvested report
58 longdoc S (OntoNotes + PreCo + LitBank) 60.1 – Paper Code 2021 linked, not harvested report
59 longdoc S (ON + PreCo + LitBank + 30k pseudo-singletons) 59.4 – Paper Code 2021 linked, not harvested report
60 UDSSM-II 59.2 – Paper – 2019 no code linked report
61 LaMini-T5 738M 59 – Paper Code 2023 linked, not harvested report
62 Flipped-3B 58.37 – Paper Code 2022 linked, not harvested report
63 KEE+NKAM winner of the WSC2016 58.3 – Paper – 2016 no code linked report
64 Char-level CNN+LSTM (partial scoring) 57.9 – Paper Code 2018 0 of 7 ran · 7 unverified report
65 UDSSM-I (ensemble) 57.1 – Paper – 2019 no code linked report
66 Knowledge Hunter 57.1 – Paper – 2018 no code linked report
67 WKH 57.1 – Paper Code 2019 linked, not harvested report
68 BERT-base 110M 56.5 – Paper – 2021 no code linked report
69 GPT-2 Small 117M (full scoring) 55.7 – Paper Code 2018 linked, not harvested report
70 ALBERT-base 11M 55.4 – Paper – 2021 no code linked report
71 Pythia 12B (0-shot) 54.8 – Paper Code 2023 linked, not harvested report
72 UDSSM-I 54.5 – Paper – 2019 no code linked report
73 Subword-level Transformer LM 54.1 – Paper Code 2017 600 of 946 ran · 346 unverified report
74 USSM + Supervised DeepNet + KB 52.8 – Paper Code 2019 linked, not harvested report
75 KEE+NKAM on WinoGrande 52.8 – Paper Code 2019 linked, not harvested report
76 USSM + KB 52 – Paper Code 2019 linked, not harvested report
77 Random chance baseline 50 – Paper – 2021 no code linked report
78 Hybrid H3 125M (3-shot, logit scoring) 43.3 – Paper Code 2022 7 of 15 ran · 8 unverified report
79 Pythia 2.8B (0-shot) 38.5 – Paper Code 2023 linked, not harvested report
80 Neo-6B (few-shot) 36.5 – Paper Code 2022 2 of 2 ran · 0 unverified report
81 Pythia 6.9B (0-shot) 36.5 – Paper Code 2023 linked, not harvested report
82 Pythia 12B (5-shot) 36.5 – Paper Code 2023 linked, not harvested report

All 82 rows shown. 82 link to a paper page on this site; 0 are marked as using additional training data in the archive. No GitHub stars are tracked; "Code" is the first repository the archive lists for the row. The archive carries no row tags, review links or community-submitted rows for this table; none are shown. archive 2025-07-28

Syntology Ran reads "N of M ran · U unverified": of the M code samples Syntology harvested from repositories linked to that row's paper (joined by arXiv id), N executed on a synthesized input and the other U = M−N are unverified (harvested, no recorded run). It counts code from repositories linked to that row's paper, not this result: the row's number was not reproduced and nothing here is a correctness claim. The other cell texts mean no graph line for the row: "linked, not harvested" (the archive links code, Syntology has not harvested it), "no code linked" (no code link in the archive), "not matched" (the row's paper URL matched no paper on this site). 32 rows have a graph line, from 19 distinct papers; 21 rows (13 papers) have at least one sample that ran. Counting each paper once: Syntology ran 883 of 1,839 samples; 956 unverified. Separately, 621 of those 1,839 are pointer-only (licence): the site points at that code rather than redistributing it, a licence property recorded for ran and unverified samples alike; each cell's tooltip carries the row's own pointer-only count. Read from the graph 2026-09-24. Per-sample status is on the paper page.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections