Datasets › SWAG

SWAG (Situations With Adversarial Generations)

Introduced by Rowan Zellers et al. in SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference16 Aug 2018 archive 2025-07-28

Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine"). SWAG (Situations With Adversarial Generations) is a large-scale dataset for this task of grounded commonsense inference, unifying natural language inference and physically grounded reasoning.

The dataset consists of 113k multiple choice questions about grounded situations. Each question is a video caption from LSMDC or ActivityNet Captions, with four answer choices about what might happen next in the scene. The correct answer is the (real) video caption for the next event in the video; the three incorrect answers are adversarially generated and human verified, so as to fool machines but not humans. The authors aim for SWAG to be a benchmark for evaluating grounded commonsense NLI and for learning representations.

Source: SWAG Image Source: Zellers et al

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Common Sense Reasoning SWAG DeBERTalarge Test 90.8 DeBERTa: Decoding-enhanced BERT with Disentangled Attention huggingface/transformers +13 5 Compare
Question Answering SWAG DeBERTaV3large Accuracy 93.4 DeBERTaV3: Improving DeBERTa using ELECTRA-Style... microsoft/DeBERTa +2 1 Compare

Papers archive 2025-07-28

5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 163. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing 3 1 18 Nov 2021 ran 0 of 7 samples (7 unverified)
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 1 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 534 1 11 Oct 2018 ran 204 of 659 samples (455 unverified; 149 pointer-only for licence)
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference 0 2 16 Aug 2018 not harvested

Dataset loaders archive 2025-07-28

4 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • SWAG

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections