Papers › SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference

SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference

16 Aug 2018EMNLP 2018 10arXiv:1808.05326archive 2025-07-28

Rowan Zellers, Yonatan Bisk, Roy Schwartz, Yejin Choi

Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine"). In this paper, we introduce the task of grounded commonsense inference, unifying natural language inference and commonsense reasoning. We present SWAG, a new dataset with 113k multiple choice questions about a rich spectrum of grounded situations. To address the recurring challenges of the annotation artifacts and human biases found in many existing datasets, we propose Adversarial Filtering (AF), a novel procedure that constructs a de-biased dataset by iteratively training an ensemble of stylistic classifiers, and using them to filter the data. To account for the aggressive adversarial filtering, we use state-of-the-art language models to massively oversample a diverse set of potential counterfactuals. Empirical results demonstrate that while humans can solve the resulting inference problems with high accuracy (88%), various competitive models struggle on our task. We provide comprehensive analysis that indicates significant opportunities for future research.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Common Sense ReasoningMultiple-choiceNatural Language InferenceQuestion Answering

Datasets

Introduced by this paper, per the archive.

SWAG

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Common Sense Reasoning SWAG ESIM + ELMo Dev 59.1 #4 of 5 Archive leaderboard report
Common Sense Reasoning SWAG ESIM + ELMo Test 59.2 #4 of 5 Archive leaderboard report
Common Sense Reasoning SWAG ESIM + GloVe Dev 51.9 #5 of 5 Archive leaderboard report
Common Sense Reasoning SWAG ESIM + GloVe Test 52.7 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections