Home › Datasets › language › Russian
Russian datasets
archive 2025-07-28
143 datasets carry the language tag "Russian", ordered by the archive's paper count. Page 3 of 3: 47 shown of 143. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Russian datasets 97–143 of 143
Includes Russian tweets and news comments from multiple sources, covering multiple stories, as well as text classification approaches to stance detection as benchmarks over this data in this language.
2 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
RuWorldTree is a QA dataset with multiple-choice elementary-level science questions, which evaluate the understanding of core science facts.
2 papers · 1 benchmark
TNCR Dataset (Table Net Detection and Classification Dataset)
We present TNCR, a new table dataset with varying image quality collected from free open source websites.
2 papers · 0 benchmarks
WikiCaps is a large-scale multilingual but non-parallel data set for multimodal machine translation and retrieval.
2 papers · 0 benchmarks
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years.
1 paper · 0 benchmarks
Bukva (Bukva: Russian Sign Language Alphabet)
We introduce a video dataset Bukva for Russian Dactyl Recognition task.
1 paper · 1 benchmark
A dataset of abdominal CT studies in NifTi format from the open-source medical data repository Medical Decathlon was utilized.
1 paper · 1 benchmark
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
1 paper · 0 benchmarks
Dataset contains CS/Math articles abstracts (in Russian) obtained from two online sources.
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
The Lenta Short Sentences dataset is a text dataset for language modelling for the Russian language.
1 paper · 0 benchmarks
The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis.
1 paper · 1 benchmark
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
MAKED (MultiModal MultiLingual Summarization and Keyword Extraction Dataset)
Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification.
1 paper · 0 benchmarks
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
We present a comprehensive dataset comprising a vast collection of raw mineral samples for the purpose of mineral recognition.
1 paper · 0 benchmarks
NEREL-BIO is an annotation scheme and corpus of PubMed abstracts in Russian and English.
1 paper · 0 benchmarks
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
1 paper · 0 benchmarks
This dataset contains orthographic samples of words in 19 languages (ar, br, de, en, eno, ent, eo, es, fi, fr, fro, it, ko, nl, pt, ru, sh, tr, zh).
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
RFSD (Russian Financial Statements Database)
The Russian Financial Statements Database (RFSD) The Russian Financial Statements Database (RFSD) is an open, harmonized collection of annual unconsolidated financial statements of the universe of Russian firms.
1 paper · 0 benchmarks
RRG (Russian RST dataset from GUM v9.1 corpus)
Parallel version of annotations in GUM RST v9.1.
1 paper · 0 benchmarks
Multilingual explainable fact-checking dataset on Russia-Ukraine Conflict 2022
1 paper · 0 benchmarks
https://github.com/dialogue-evaluation/RuOpinionNE-2024
1 paper · 0 benchmarks
CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024…
1 paper · 2 benchmarks
CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024…
1 paper · 2 benchmarks
CL-RuTerm3 dataset is a novel resource featuring nested term annotations across six domains (the main one is computational linguistics, also mathematics, medicine, economics, literature studies, and agrochemistry), and the RuTermEval-2024…
1 paper · 2 benchmarks
The work provides a comprehensive overview of the corpus for the Russian language for the commonsense inference task.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Synthetic Question Answering dataset in Serbian, acquired by automatic translation of SQuAD.
1 paper · 0 benchmarks
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
533 parallel examples sampled from TACRED, translated into Russian and Korean (and 3 additional examples in Russian), accompanied with tranlsation of a list of trigger words collected for the different relations.
1 paper · 0 benchmarks
UNER v1 adds an NER annotation layer to 18 datasets (primarily treebanks from UD) and covers 12 geneologically and ty- pologically diverse languages: Cebuano, Danish, German, English, Croatian, Portuguese, Russian, Slovak, Serbian,…
1 paper · 31 benchmarks
WEATHub is a dataset containing 24 languages.
1 paper · 0 benchmarks
News translation is a recurring WMT task.
1 paper · 0 benchmarks
The Winograd schema challenge composes tasks with syntactic ambiguity, which can be resolved with logic and reasoning.
1 paper · 1 benchmark
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
The first and the one open dataset for Russian finger- spelling, contained 1,593 annotated phrases and over 37 thousand HD+ videos.
1 paper · 1 benchmark
This dataset contains human-annotated sense identifiers for 2562 contexts of 20 words used in the RUSSE'2018 shared task on Word Sense Induction and Disambiguation for the Russian language.
0 papers · 0 benchmarks
LRWC (Lexical Relations from the Wisdom of the Crowd)
This dataset contains the opinions of Russian native speakers about the relationship between a generic term (hypernym) and a specific instance of it (hyponym).
0 papers · 0 benchmarks
RuADReCT (The Russian Adverse Drug Reaction Corpus of Tweets)
Created as part of the Social Media Mining for Health Applications (#SMM4H '20) shared tasks, this dataset consists of 9515 tweets describing health issues.
0 papers · 0 benchmarks
This dataset contains annotations of semantic frames and intra-frame syntax for 1500 Russian sentences.
0 papers · 0 benchmarks
Video Dataset (Storytelling Video Dataset (Russian, Emotion, Gesture, Speech))
The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.