{"url":"/dataset/russe","name":"RUSSE","full_name":"Russian Words in Context (based on RUSSE)","description_markdown":"WiC: The Word-in-Context Dataset A reliable benchmark for the evaluation of context-sensitive word embeddings.\r\n\r\nDepending on its context, an ambiguous word can refer to multiple, potentially unrelated, meanings. Mainstream static word embeddings, such as Word2vec and GloVe, are unable to reflect this dynamic semantic nature. Contextualised word embeddings are an attempt at addressing this limitation by computing dynamic representations for words which can adapt based on context.\r\n\r\nRussian SuperGLUE task borrows original data from the Russe project, Word Sense Induction and Disambiguation shared task (2018)\r\n\r\n### Task Type\r\nReading Comprehension. Binary Classification: true/false\r\n\r\n### Example\r\n```\r\n{\r\n  \"idx\" : 8,\r\n  \"word\" : \"дорожка\",\r\n  \"sentence1\" : \"Бурые ковровые дорожки заглушали шаги\",\r\n  \"sentence2\" : \"Приятели решили выпить на дорожку в местном баре\",\r\n  \"start1\" : 15,\r\n  \"end1\" : 23,\r\n  \"start2\" : 26,\r\n  \"end2\" : 34,\r\n  \"label\" : false,\r\n  \"gold_sense1\" : 1,\r\n  \"gold_sense2\" : 2\r\n}\r\n```\r\n                      \r\n### How did we collect data? \r\nAll text examples were collected from Russe original dataset, already collected by Russian Semantic Evaluation at ACL SIGSLAV. Human assessment was carried out on Yandex.Toloka.\r\n\r\nIn version 2, we have manually collected in the same format testset.","description_withheld":null,"homepage":"https://github.com/RussianNLP/RussianSuperGLUE","introduced_date":"2018-03-15","introduced_date_note":null,"introduced_by":{"paper":"/paper/russe2018-a-shared-task-on-word-sense","title":"RUSSE'2018: A Shared Task on Word Sense Induction for the Russian Language","first_author":"Alexander Panchenko","url":null},"license":{"name":"MIT License","url":"https://github.com/RussianNLP/RussianSuperGLUE/blob/master/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Reading Comprehension","url":"/task/reading-comprehension","datasets_with_task":"/datasets/task/reading-comprehension"},{"name":"Word Sense Disambiguation","url":"/task/word-sense-disambiguation","datasets_with_task":"/datasets/task/word-sense-disambiguation"}],"languages":[{"name":"Russian","url":"/datasets/language/russian"}],"variants":["RUSSE"],"data_loaders":[],"num_papers_in_archive":8,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/word-sense-disambiguation-on-russe","task":"Word Sense Disambiguation","dataset_variant":"RUSSE","rows":22,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"Human Benchmark","paper":"/paper/russiansuperglue-a-russian-language","metrics":{"Accuracy":"0.805"},"code_links":[{"title":"RussianNLP/RussianSuperGLUE","url":"https://github.com/RussianNLP/RussianSuperGLUE"},{"title":"RussianNLP/MOROCCO","url":"https://github.com/RussianNLP/MOROCCO"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/unreasonable-effectiveness-of-rule-based","title":"Unreasonable Effectiveness of Rule-Based Heuristics in Solving Russian SuperGLUE Tasks","date":"2021-05-03","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/russiansuperglue-a-russian-language","title":"RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark","date":"2020-10-29","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}