Datasets › WSC

WSC (Winograd Schema Challenge)

Introduced in The Winograd Schema Challenge1 Jan 2012 archive 2025-07-28

The Winograd Schema Challenge was introduced both as an alternative to the Turing Test and as a test of a system’s ability to do commonsense reasoning. A Winograd schema is a pair of sentences differing in one or two words with a highly ambiguous pronoun, resolved differently in the two sentences, that appears to require commonsense knowledge to be resolved correctly. The examples were designed to be easily solvable by humans but difficult for machines, in principle requiring a deep understanding of the content of the text and the situation it describes.

The original Winograd Schema Challenge dataset consisted of 100 Winograd schemas constructed manually by AI experts. As of 2020 there are 285 examples available; however, the last 12 examples were only added recently. To ensure consistency with earlier models, several authors often prefer to report the performance on the first 273 examples only. These datasets are usually referred to as WSC285 and WSC273, respectively.

Source: https://arxiv.org/pdf/2004.13831.pdf Image Source: https://arxiv.org/pdf/1907.11983.pdf

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Coreference Resolution Winograd Schema Challenge PaLM 540B (fine-tuned) Accuracy 100 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 82 Compare
Classification WSC OPT-1.3B Test Accuracy 64.16% Achieving Dimension-Free Communication in Federated... ZidongLiu/DeComFL 2 Compare

Papers archive 2025-07-28

30 shown of 38 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 361. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Achieving Dimension-Free Communication in Federated Learning via Zeroth-Order Optimization 1 2 24 May 2024 ran 1 of 1 samples (0 unverified)
The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning 2 1 23 May 2023 not harvested
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions 1 5 27 Apr 2023 not harvested
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling 4 4 3 Apr 2023 not harvested
Exploring the Benefits of Training Expert Language Models over Instruction Tuning 2 1 7 Feb 2023 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Hungry Hungry Hippos: Towards Language Modeling with State Space Models 3 3 28 Dec 2022 ran 7 of 15 samples (8 unverified)
Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE 0 2 4 Dec 2022 not harvested
Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models 0 1 28 Oct 2022 not harvested
Scaling Instruction-Finetuned Language Models 9 1 20 Oct 2022 ran 8 of 17 samples (9 unverified; 2 pointer-only for licence)
Guess the Instruction! Flipped Learning Makes Language Models Stronger Zero-Shot Learners 1 1 6 Oct 2022 not harvested
Ask Me Anything: A simple strategy for prompting language models 3 3 5 Oct 2022 ran 2 of 2 samples (0 unverified)
AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model 1 1 2 Aug 2022 ran 1 of 1 samples (0 unverified)
N-Grammer: Augmenting Transformers with latent n-grams 2 1 13 Jul 2022 ran 0 of 6 samples (6 unverified)
UL2: Unifying Language Learning Paradigms 2 2 10 May 2022 ran 0 of 16 samples (16 unverified)
PaLM: Scaling Language Modeling with Pathways 7 4 5 Apr 2022 ran 30 of 37 samples (7 unverified)
ST-MoE: Designing Stable and Transferable Sparse Expert Models 3 2 17 Feb 2022 ran 5 of 5 samples (0 unverified; 5 pointer-only for licence)
On Generalization in Coreference Resolution 2 2 20 Sep 2021 not harvested
Finetuned Language Models Are Zero-Shot Learners 8 2 3 Sep 2021 ran 0 of 1 samples (1 unverified)
Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd Schema 0 7 16 Apr 2021 not harvested
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
Language Models are Few-Shot Learners 67 1 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
Generative Data Augmentation for Commonsense Reasoning 1 1 24 Apr 2020 ran 2 of 4 samples (2 unverified; 4 pointer-only for licence)
TTTTTackling WinoGrande Schemas 0 1 18 Mar 2020 not harvested
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 57 1 23 Oct 2019 ran 2 of 31 samples (29 unverified)
A Hybrid Neural Network Model for Commonsense Reasoning 3 1 27 Jul 2019 ran 0 of 1 samples (1 unverified)
WinoGrande: An Adversarial Winograd Schema Challenge at Scale 10 4 24 Jul 2019 not harvested
Attention Is (not) All You Need for Commonsense Reasoning 2 3 31 May 2019 not harvested
A Surprisingly Robust Trick for Winograd Schema Challenge 2 4 15 May 2019 not harvested
SocialIQA: Commonsense Reasoning about Social Interactions 1 2 22 Apr 2019 ran 0 of 10 samples (10 unverified)

The full list of 38 is in the JSON twin.

Dataset loaders archive 2025-07-28

6 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Winograd Schema Challenge
  • WSC
  • winograd_wsc

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections