Papers › Attention Is (not) All You Need for Commonsense Reasoning

Attention Is (not) All You Need for Commonsense Reasoning

31 May 2019ACL 2019 7arXiv:1905.13497archive 2025-07-28

Tassilo Klein, Moin Nabi

The recently introduced BERT model exhibits strong performance on several language understanding benchmarks. In this paper, we describe a simple re-implementation of BERT for commonsense reasoning. We show that the attentions produced by BERT can be directly utilized for tasks such as the Pronoun Disambiguation Problem and Winograd Schema Challenge. Our proposed attention-guided commonsense reasoning method is conceptually simple yet empirically powerful. Experimental analysis on multiple datasets demonstrates that our proposed system performs remarkably well on all cases while outperforming the previously reported state of the art by a margin. While results suggest that BERT seems to implicitly learn to establish complex relationships between entities, solving commonsense reasoning tasks might require more than unsupervised models learned from huge text corpora.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

SAP-samples/acl2020-commonsense mentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AllCoreference ResolutionNatural Language Understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Coreference Resolution Winograd Schema Challenge BERT-base 110M + MAS Accuracy 60.3 #57 of 82 Archive leaderboard report
Coreference Resolution Winograd Schema Challenge USSM + Supervised DeepNet + KB Accuracy 52.8 #74 of 82 Archive leaderboard report
Coreference Resolution Winograd Schema Challenge USSM + KB Accuracy 52 #76 of 82 Archive leaderboard report
Natural Language Understanding PDP60 BERT-base 110M + MAS Accuracy 68.3 #7 of 13 Archive leaderboard report
Natural Language Understanding PDP60 USSM + Supervised Deepnet + 3 Knowledge Bases Accuracy 66.7 #8 of 13 Archive leaderboard report
Natural Language Understanding PDP60 USSM + Supervised Deepnet Accuracy 53.3 #13 of 13 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections