Papers › TAPE: Assessing Few-shot Russian Language Understanding

TAPE: Assessing Few-shot Russian Language Understanding

23 Oct 2022arXiv:2210.12813archive 2025-07-28

Ekaterina Taktasheva, Tatiana Shavrina, Alena Fenogenova, Denis Shevelev, Nadezhda Katricheva, Maria Tikhonova, Albina Akhmetgareeva, Oleg Zinkevich, Anastasiia Bashmakova, Svetlana Iordanskaia, Alena Spiridonova, Valentina Kurenshchikova, Ekaterina Artemova, Vladislav Mikhailov

Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes. However, this fast-growing area lacks standardized evaluation suites for non-English languages, hindering progress outside the Anglo-centric paradigm. To address this line of research, we propose TAPE (Text Attack and Perturbation Evaluation), a novel benchmark that includes six more complex NLU tasks for Russian, covering multi-hop reasoning, ethical concepts, logic and commonsense knowledge. The TAPE's design focuses on systematic zero-shot and few-shot NLU evaluation: (i) linguistic-oriented adversarial attacks and perturbations for analyzing robustness, and (ii) subpopulations for nuanced interpretation. The detailed analysis of testing the autoregressive baselines indicates that simple spelling-based perturbations affect the performance the most, while paraphrasing the input has a more negligible effect. At the same time, the results demonstrate a significant gap between the neural and human baselines for most tasks. We publicly release TAPE (tape-benchmark.com) to foster research on robust LMs that can generalize to new tasks when little to no supervision is available.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Adversarial AttackAdversarial TextEthicsFew-Shot LearningLogical ReasoningQuestion AnsweringZero-Shot Learning

Datasets

Introduced by this paper, per the archive.

CheGeKaEthics (per ethics)MultiQRuOpenBookQARuWorldTreeWinograd Automatic

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Ethics Ethics RuGPT-3 Large Accuracy 68.6 #1 of 4 Archive leaderboard report
Ethics Ethics RuGPT-3 Meduim Accuracy 68.3 #2 of 4 Archive leaderboard report
Ethics Ethics RuGPT-3 Small Accuracy 55.5 #3 of 4 Archive leaderboard report
Ethics Ethics Human benchmark Accuracy 52.9 #4 of 4 Archive leaderboard report
Ethics Ethics (per ethics) Human benchmark Accuracy 67.6 #1 of 4 Archive leaderboard report
Ethics Ethics (per ethics) RuGPT-3 Small Accuracy 60.9 #2 of 4 Archive leaderboard report
Ethics Ethics (per ethics) RuGPT-3 Large Accuracy 44.9 #3 of 4 Archive leaderboard report
Ethics Ethics (per ethics) RuGPT-3 Medium Accuracy 44.1 #4 of 4 Archive leaderboard report
Logical Reasoning RuWorldTree Human benchmark Accuracy 83.7 #1 of 4 Archive leaderboard report
Logical Reasoning RuWorldTree RuGPT-3 Large Accuracy 40.7 #2 of 4 Archive leaderboard report
Logical Reasoning RuWorldTree RuGPT-3 Medium Accuracy 38.0 #3 of 4 Archive leaderboard report
Logical Reasoning RuWorldTree RuGPT-3 Small Accuracy 34.0 #4 of 4 Archive leaderboard report
Logical Reasoning Winograd Automatic Human benchmark Accuracy 87.0 #1 of 4 Archive leaderboard report
Logical Reasoning Winograd Automatic RuGPT-3 Small Accuracy 57.9 #2 of 4 Archive leaderboard report
Logical Reasoning Winograd Automatic RuGPT-3 Medium Accuracy 57.2 #3 of 4 Archive leaderboard report
Logical Reasoning Winograd Automatic RuGPT-3 Large Accuracy 55.5 #4 of 4 Archive leaderboard report
Question Answering CheGeKa Human benchmark Accuracy 64.5 #1 of 4 Archive leaderboard report
Question Answering CheGeKa RuGPT-3 Large Accuracy 00 #2 of 4 Archive leaderboard report
Question Answering CheGeKa RuGPT-3 Medium Accuracy 00 #3 of 4 Archive leaderboard report
Question Answering CheGeKa RuGPT-3 Small Accuracy 00 #4 of 4 Archive leaderboard report
Question Answering MultiQ Human benchmark Accuracy 91.0 #1 of 4 Archive leaderboard report
Question Answering MultiQ RuGPT-3 Large Accuracy 00 #2 of 4 Archive leaderboard report
Question Answering MultiQ RuGPT-3 Medium Accuracy 00 #3 of 4 Archive leaderboard report
Question Answering MultiQ RuGPT-3 Small Accuracy 00 #4 of 4 Archive leaderboard report
Question Answering RuOpenBookQA Human benchmark Accuracy 86.5 #1 of 4 Archive leaderboard report
Question Answering RuOpenBookQA RuGPT-3 Small Accuracy 57.9 #2 of 4 Archive leaderboard report
Question Answering RuOpenBookQA RuGPT-3 Medium Accuracy 57.2 #3 of 4 Archive leaderboard report
Question Answering RuOpenBookQA RuGPT-3 Large Accuracy 55.5 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections