Papers › Unreasonable Effectiveness of Rule-Based Heuristics in Solving Russian SuperGLUE Tasks

Unreasonable Effectiveness of Rule-Based Heuristics in Solving Russian SuperGLUE Tasks

3 May 2021arXiv:2105.01192archive 2025-07-28

Tatyana Iazykova, Denis Kapelyushnik, Olga Bystrova, Andrey Kutuzov

Leader-boards like SuperGLUE are seen as important incentives for active development of NLP, since they provide standard benchmarks for fair comparison of modern language models. They have driven the world's best engineering teams as well as their resources to collaborate and solve a set of tasks for general language understanding. Their performance scores are often claimed to be close to or even higher than the human performance. These results encouraged more thorough analysis of whether the benchmark datasets featured any statistical cues that machine learning based language models can exploit. For English datasets, it was shown that they often contain annotation artifacts. This allows solving certain tasks with very simple rules and achieving competitive rankings. In this paper, a similar analysis was done for the Russian SuperGLUE (RSG), a recently published benchmark set and leader-board for Russian natural language understanding. We show that its test datasets are vulnerable to shallow heuristics. Often approaches based on simple rules outperform or come close to the results of the notorious pre-trained language models like GPT-3 or BERT. It is likely (as the simplest explanation) that a significant part of the SOTA models performance in the RSG leader-board is due to exploiting these shallow heuristics and that has nothing in common with real language understanding. We provide a set of recommendations on how to improve these datasets, making the RSG leader-board even more representative of the real progress in Russian NLU.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Common Sense ReasoningNatural Language InferenceNatural Language UnderstandingQuestion AnsweringReading ComprehensionWord Sense Disambiguation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Common Sense Reasoning PARus majority_class Accuracy 0.498 #17 of 22 Archive leaderboard report
Common Sense Reasoning PARus Random weighted Accuracy 0.48 #20 of 22 Archive leaderboard report
Common Sense Reasoning PARus heuristic majority Accuracy 0.478 #21 of 22 Archive leaderboard report
Common Sense Reasoning RWSD Random weighted Accuracy 0.597 #3 of 22 Archive leaderboard report
Common Sense Reasoning RWSD heuristic majority Accuracy 0.669 #17 of 22 Archive leaderboard report
Common Sense Reasoning RWSD majority_class Accuracy 0.669 #20 of 22 Archive leaderboard report
Common Sense Reasoning RuCoS heuristic majority Average F1 0.26 #15 of 22 Archive leaderboard report
Common Sense Reasoning RuCoS heuristic majority EM 0.257 #15 of 22 Archive leaderboard report
Common Sense Reasoning RuCoS Random weighted Average F1 0.25 #17 of 22 Archive leaderboard report
Common Sense Reasoning RuCoS Random weighted EM 0.247 #17 of 22 Archive leaderboard report
Common Sense Reasoning RuCoS majority_class Average F1 0.25 #18 of 22 Archive leaderboard report
Common Sense Reasoning RuCoS majority_class EM 0.247 #18 of 22 Archive leaderboard report
Natural Language Inference LiDiRus heuristic majority MCC 0.147 #13 of 22 Archive leaderboard report
Natural Language Inference LiDiRus Random weighted MCC 0 #20 of 22 Archive leaderboard report
Natural Language Inference LiDiRus majority_class MCC 0 #21 of 22 Archive leaderboard report
Natural Language Inference RCB heuristic majority Accuracy 0.438 #6 of 22 Archive leaderboard report
Natural Language Inference RCB heuristic majority Average F1 0.4 #6 of 22 Archive leaderboard report
Natural Language Inference RCB Random weighted Accuracy 0.374 #17 of 22 Archive leaderboard report
Natural Language Inference RCB Random weighted Average F1 0.319 #17 of 22 Archive leaderboard report
Natural Language Inference RCB majority_class Accuracy 0.484 #22 of 22 Archive leaderboard report
Natural Language Inference RCB majority_class Average F1 0.217 #22 of 22 Archive leaderboard report
Natural Language Inference TERRa heuristic majority Accuracy 0.549 #17 of 22 Archive leaderboard report
Natural Language Inference TERRa majority_class Accuracy 0.513 #18 of 22 Archive leaderboard report
Natural Language Inference TERRa Random weighted Accuracy 0.483 #21 of 22 Archive leaderboard report
Question Answering DaNetQA heuristic majority Accuracy 0.642 #11 of 22 Archive leaderboard report
Question Answering DaNetQA Random weighted Accuracy 0.52 #21 of 22 Archive leaderboard report
Question Answering DaNetQA majority_class Accuracy 0.503 #22 of 22 Archive leaderboard report
Reading Comprehension MuSeRC heuristic majority Average F1 0.671 #15 of 22 Archive leaderboard report
Reading Comprehension MuSeRC heuristic majority EM 0.237 #15 of 22 Archive leaderboard report
Reading Comprehension MuSeRC Random weighted Average F1 0.45 #21 of 22 Archive leaderboard report
Reading Comprehension MuSeRC Random weighted EM 0.071 #21 of 22 Archive leaderboard report
Reading Comprehension MuSeRC majority_class Average F1 0.0 #22 of 22 Archive leaderboard report
Reading Comprehension MuSeRC majority_class EM 0.0 #22 of 22 Archive leaderboard report
Word Sense Disambiguation RUSSE heuristic majority Accuracy 0.595 #15 of 22 Archive leaderboard report
Word Sense Disambiguation RUSSE majority_class Accuracy 0.587 #18 of 22 Archive leaderboard report
Word Sense Disambiguation RUSSE Random weighted Accuracy 0.528 #22 of 22 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTBPECosine AnnealingDense ConnectionsDropoutGPT-3Layer NormalizationLinear LayerLinear Warmup With Cosine AnnealingLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections