Papers › How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the Winograd Schema...

How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the Winograd Schema Challenge and SWAG

5 Nov 2018IJCNLP 2019 11arXiv:1811.01778archive 2025-07-28

Paul Trichelair, Ali Emami, Adam Trischler, Kaheer Suleman, Jackie Chi Kit Cheung

Recent studies have significantly improved the state-of-the-art on common-sense reasoning (CSR) benchmarks like the Winograd Schema Challenge (WSC) and SWAG. The question we ask in this paper is whether improved performance on these benchmarks represents genuine progress towards common-sense-enabled systems. We make case studies of both benchmarks and design protocols that clarify and qualify the results of previous work by analyzing threats to the validity of previous experimental designs. Our protocols account for several properties prevalent in common-sense benchmarks including size limitations, structural regularities, and variable instance difficulty.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Common Sense ReasoningCoreference Resolution

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Coreference Resolution Winograd Schema Challenge GPT-2 Medium 774M (partial scoring) Accuracy 69.2 #36 of 82 Archive leaderboard report
Coreference Resolution Winograd Schema Challenge GPT-2 Medium 774M (full scoring) Accuracy 64.5 #43 of 82 Archive leaderboard report
Coreference Resolution Winograd Schema Challenge GPT-2 Small 117M (partial scoring) Accuracy 61.5 #54 of 82 Archive leaderboard report
Coreference Resolution Winograd Schema Challenge GPT-2 Small 117M (full scoring) Accuracy 55.7 #69 of 82 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections