Papers › A Surprisingly Robust Trick for Winograd Schema Challenge

A Surprisingly Robust Trick for Winograd Schema Challenge

15 May 2019arXiv:1905.06290archive 2025-07-28

Vid Kocijan, Ana-Maria Cretu, Oana-Maria Camburu, Yordan Yordanov, Thomas Lukasiewicz

The Winograd Schema Challenge (WSC) dataset WSC273 and its inference counterpart WNLI are popular benchmarks for natural language understanding and commonsense reasoning. In this paper, we show that the performance of three language models on WSC273 strongly improves when fine-tuned on a similar pronoun disambiguation problem dataset (denoted WSCR). We additionally generate a large unsupervised WSC-like dataset. By fine-tuning the BERT language model both on the introduced and on the WSCR dataset, we achieve overall accuracies of 72.5% and 74.7% on WSC273 and WNLI, improving the previous state-of-the-art solutions by 8.8% and 9.6%, respectively. Furthermore, our fine-tuned models are also consistently more robust on the "complex" subsets of WSC273, introduced by Trichelair et al. (2018).

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

vid-koci/bert-commonsense officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Common Sense ReasoningCoreference ResolutionLanguage ModelingLanguage ModellingNatural Language InferenceNatural Language UnderstandingWNLI

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Coreference Resolution Winograd Schema Challenge BERTwiki 340M (fine-tuned on WSCR) Accuracy 72.5 #30 of 82 Archive leaderboard report
Coreference Resolution Winograd Schema Challenge BERT-large 340M (fine-tuned on WSCR) Accuracy 71.4 #32 of 82 Archive leaderboard report
Coreference Resolution Winograd Schema Challenge BERTwiki 340M (fine-tuned on half of WSCR) Accuracy 70.3 #34 of 82 Archive leaderboard report
Coreference Resolution Winograd Schema Challenge BERT-base 110M (fine-tuned on WSCR) Accuracy 62.3 #51 of 82 Archive leaderboard report
Natural Language Inference WNLI BERTwiki 340M (fine-tuned on WSCR) Accuracy 74.7 #13 of 23 Archive leaderboard report
Natural Language Inference WNLI BERT-large 340M (fine-tuned on WSCR) Accuracy 71.9 #15 of 23 Archive leaderboard report
Natural Language Inference WNLI BERT-base 110M (fine-tuned on WSCR) Accuracy 70.5 #16 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections