Datasets › WinoGrande

WinoGrande

Introduced by Keisuke Sakaguchi et al. in WinoGrande: An Adversarial Winograd Schema Challenge at Scale archive 2025-07-28

WinoGrande is a large-scale dataset of 44k problems, inspired by the original WSC design, but adjusted to improve both the scale and the hardness of the dataset. The key steps of the dataset construction consist of (1) a carefully designed crowdsourcing procedure, followed by (2) systematic bias reduction using a novel AfLite algorithm that generalizes human-detectable word associations to machine-detectable embedding associations.

Source: WinoGrande: An Adversarial Winograd Schema Challenge at Scale Image Source: https://winogrande.allenai.org/

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Common Sense Reasoning WinoGrande ST-MoE-32B 269B (fine-tuned) Accuracy 96.1 ST-MoE: Designing Stable and Transferable Sparse Expert Models tensorflow/mesh +2 77 Compare
parameter-efficient fine-tuning WinoGrande LLaMA2-7b Accuracy (% ) 70.80 GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient... On-Point-RND/GIFT_SW 3 Compare
Text Generation WinoGrande no rows — — 0 Compare
Text Generation Winogrande (5-shot) no rows — — 0 Compare
Text Generation Winogrande TR no rows — — 0 Compare
Text Generation Winogrande TR v0.2 no rows — — 0 Compare
Winogrande WinoGrande no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 34 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 703. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs 1 1 27 Aug 2024 ran 3 of 5 samples (2 unverified; 3 pointer-only for licence)
Mixture-of-Subspaces in Low-Rank Adaptation 1 1 16 Jun 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts 2 3 22 Apr 2024 ran 6 of 11 samples (5 unverified)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM 1 1 12 Mar 2024 not harvested
The Claude 3 Model Family: Opus, Sonnet, Haiku 0 3 4 Mar 2024 not harvested
DoRA: Weight-Decomposed Low-Rank Adaptation 5 1 14 Feb 2024 ran 8 of 15 samples (7 unverified; 14 pointer-only for licence)
Mixtral of Experts 6 2 8 Jan 2024 ran 5 of 5 samples (0 unverified)
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks 2 1 5 Jan 2024 not harvested
Mistral 7B 6 1 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
Textbooks Are All You Need II: phi-1.5 technical report 1 1 11 Sep 2023 not harvested
The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning 2 1 23 May 2023 not harvested
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions 1 6 27 Apr 2023 not harvested
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling 4 4 3 Apr 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
GPT-4 Technical Report 11 2 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
Exploring the Benefits of Training Expert Language Models over Instruction Tuning 2 1 7 Feb 2023 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models 0 1 28 Oct 2022 not harvested
Task Compass: Scaling Multi-task Pre-training with Task Prefix 1 3 12 Oct 2022 not harvested
Guess the Instruction! Flipped Learning Makes Language Models Stronger Zero-Shot Learners 1 1 6 Oct 2022 not harvested
PaLM: Scaling Language Modeling with Pathways 7 3 5 Apr 2022 ran 30 of 37 samples (7 unverified)
Training Compute-Optimal Large Language Models 2 1 29 Mar 2022 ran 8 of 11 samples (3 unverified; 4 pointer-only for licence)
Efficient Language Modeling with Sparse all-MLP 0 5 14 Mar 2022 not harvested
ST-MoE: Designing Stable and Transferable Sparse Expert Models 3 2 17 Feb 2022 ran 5 of 5 samples (0 unverified; 5 pointer-only for licence)
Scaling Language Models: Methods, Analysis & Insights from Training Gopher 3 1 8 Dec 2021 not harvested
Finetuned Language Models Are Zero-Shot Learners 8 2 3 Sep 2021 ran 0 of 1 samples (1 unverified)
LoRA: Low-Rank Adaptation of Large Language Models 74 1 17 Jun 2021 ran 34 of 84 samples (50 unverified; 28 pointer-only for licence)
Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd Schema 0 7 16 Apr 2021 not harvested
UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark 1 1 24 Mar 2021 ran 0 of 5 samples (5 unverified)

The full list of 34 is in the JSON twin.

Dataset loaders archive 2025-07-28

6 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC-BY

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • WinoGrande
  • Winogrande
  • Winogrande (5-shot)
  • Winogrande TR v0.2
  • Winogrande TR

5 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections