Home › Datasets › language › Russian
Russian datasets
archive 2025-07-28
143 datasets carry the language tag "Russian", ordered by the archive's paper count. Page 2 of 3: 48 shown of 143. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Russian datasets 49–96 of 143
RuBQ (Russian Knowledge Base Questions)
The first Russian knowledge base question answering (KBQA) dataset.
7 papers · 0 benchmarks
TERRa (Textual Entailment Recognition for Russian)
Textual Entailment Recognition has been proposed recently as a generic task that captures major semantic inference needs across many NLP applications, such as Question Answering, Information Retrieval, Information Extraction, and Text…
7 papers · 1 benchmark
CLAMS (Cross-linguistic Analysis of Models on Syntax)
Targeted syntactic evaluation datasets in 5 languages: English, French, German, Russian, and Hebrew.
6 papers · 0 benchmarks
Concepticon (Concepticon. A Resource for the Linking of Concept Lists)
This resource, our Concepticon, links concept labels from different conceptlists to concept sets.
6 papers · 0 benchmarks
GTSinger (GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks)
The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages…
6 papers · 0 benchmarks
MuSeRC (Russian Multi-Sentence Reading Comprehension)
We present a reading comprehension challenge in which questions can only be answered by taking into account information from multiple sentences.
6 papers · 1 benchmark
RCB (Russian Commitment Bank)
The Russian Commitment Bank is a corpus of naturally occurring discourses whose final sentence contains a clause-embedding predicate under an entailment cancelling operator (question, modal, negation, antecedent of conditional).
6 papers · 1 benchmark
WikiNEuRal is a high-quality automatically-generated dataset for Multilingual Named Entity Recognition.
6 papers · 0 benchmarks
XL-BEL is a benchmark for cross-lingual biomedical entity linking (XL-BEL).
6 papers · 0 benchmarks
XQA is a data which consists of a total amount of 90k question-answer pairs in nine languages for cross-lingual open-domain question answering.
6 papers · 0 benchmarks
A clickthrough prediction dataset, for more information please see the Kaggle page
5 papers · 1 benchmark
LiDiRus (Linguistic Diagnostic for Russian)
LiDiRus is a diagnostic dataset that covers a large volume of linguistic phenomena, while allowing you to evaluate information systems on a simple test of textual entailment recognition.
5 papers · 1 benchmark
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
The Russian Corpus of Linguistic Acceptability (RuCoLA) is built from the ground up under the well-established binary LA approach.
5 papers · 1 benchmark
RuCoS (Russian Reading Comprehension with Commonsense Reasoning)
Russian reading comprehension with Commonsense reasoning (RuCoS) is a large-scale reading comprehension dataset that requires commonsense reasoning.
5 papers · 1 benchmark
Taiga is a corpus, where text sources and their meta-information are collected according to popular ML tasks.
5 papers · 0 benchmarks
DISRPT2019 (DISRPT2019 shared task on Discourse Unit Segmentation and Connective Detection)
The DISRPT 2019 workshop introduces the first iteration of a cross-formalism shared task on discourse unit segmentation.
4 papers · 0 benchmarks
Gazeta is a dataset for automatic summarization of Russian news.
4 papers · 1 benchmark
GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets.
4 papers · 0 benchmarks
HKR (Handwritten Kazakh and Russian (HKR) Database for Text Recognition)
The database is written in Cyrillic and shares the same 33 characters.
4 papers · 1 benchmark
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
DISRPT2021 (DISRPT2021 shared task on Discourse Unit Segmentation, Connective Detection and Discourse Relation Classification)
The DISRPT 2021 shared task, co-located with CODI 2021 at EMNLP, introduces the second iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the first iteration of a cross-formalism…
3 papers · 0 benchmarks
Digital Peter is a dataset of Peter the Great's manuscripts annotated for segmentation and text recognition.
3 papers · 1 benchmark
HeadlineCause is a dataset for detecting implicit causal relations between pairs of news headlines.
3 papers · 0 benchmarks
MRS (Multilingual Reply Suggestion)
MRS, a multilingual reply suggestion dataset with ten languages.
3 papers · 0 benchmarks
Frame-to-frame video alignment/synchronization
3 papers · 1 benchmark
MultiQ is a multi-hop QA dataset for Russian, suitable for general open-domain question answering, information retrieval, and reading comprehension tasks.
3 papers · 1 benchmark
PatTR (Patent Translation Resource)
PatTR is a sentence-parallel corpus extracted from the MAREC patent collection.
3 papers · 0 benchmarks
QALD-9-Plus Dataset Description QALD-9-Plus is the dataset for Knowledge Graph Question Answering (KGQA) based on well-known QALD-9.
3 papers · 1 benchmark
RESD (Russian Emotional Speech Dialogs with annotated text)
Russian dataset of emotional speech dialogues.
3 papers · 1 benchmark
RuShiftEval is a manually annotated lexical semantic change dataset for Russian.
3 papers · 0 benchmarks
TyDiP (A Dataset for Politeness Classification in Nine Typologically Diverse Languages)
A Dataset for Politeness Classification in Nine Typologically Diverse Languages (TyDiP) is a dataset containing three-way politeness annotations for 500 examples in each language, totaling 4.5K examples.
3 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
CLSE (Corpus of Linguistically Significant Entities)
2 papers · 0 benchmarks
CheGeKa is a Jeopardy!-like Russian QA dataset collected from the official Russian quiz database ChGK.
2 papers · 1 benchmark
Dusha (Dusha Crowd, Dusha Podcast)
Dusha is a dataset for speech emotion recognition (SER) tasks.
2 papers · 2 benchmarks
Ethics (per ethics) dataset is created to test the knowledge of the basic concepts of morality.
2 papers · 1 benchmark
Golos is a Russian speech dataset suitable for speech research.
2 papers · 0 benchmarks
Kaleidoscope (Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation)
The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage.
2 papers · 0 benchmarks
MGSM8KInstruct, the multilingual math reasoning instruction dataset, encompassing ten distinct languages, thus addressing the issue of training data scarcity in multilingual math reasoning.
2 papers · 0 benchmarks
Morph Call is a suite of 46 probing tasks for four Indo-European languages that fall under different morphology: Russian, French, English, and German.
2 papers · 0 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
RUSLAN is a Russian spoken language corpus for text-to-speech task.
2 papers · 0 benchmarks
RuMedBench is a benchmark dataset for Russian medical language understanding.
2 papers · 0 benchmarks
RuOpenBookQA is a QA dataset with multiple-choice elementary-level science questions which probe the understanding of core science facts.
2 papers · 1 benchmark
https://github.com/dialogue-evaluation/RuSentNE-evaluation
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.