{"url":"/dataset/parus","name":"PARus","full_name":"Choice of Plausible Alternatives for Russian language","description_markdown":"Choice of Plausible Alternatives for Russian language (PARus) evaluation provides researchers with a tool for assessing progress in open-domain commonsense causal reasoning. Each question in PARus is composed of a premise and two alternatives, where the task is to select the alternative that more plausibly has a causal relation with the premise. The correct alternative is randomized so that the expected performance of randomly guessing is 50%.\r\n\r\n### Task Type\r\nEvaluation of commonsense causal reasoning\r\n\r\nSentence Pair Classification: suitable - not suitable\r\n\r\n### Example\r\n```\r\n{\r\n  \"premise\": \"Гости вечеринки прятались за диваном.\",\r\n  \"choice1\": \"Это была вечеринка-сюрприз.\",\r\n  \"choice2\":\"Это был день рождения.\",\r\n  \"question\": \"cause\",\r\n  \"label\": 0,\r\n  \"idx\": 4\r\n}\r\n```\r\n### How did we collect data? \r\nAll text examples were collected from open news sources and literary magazines, then manually reviewed and supplemented by a human assessment on Yandex.Toloka","description_withheld":null,"homepage":"https://github.com/RussianNLP/RussianSuperGLUE","introduced_date":"2020-10-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/russiansuperglue-a-russian-language","title":"RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark","first_author":"Tatiana Shavrina","url":null},"license":{"name":"MIT License","url":"https://github.com/RussianNLP/RussianSuperGLUE/blob/master/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Common Sense Reasoning","url":"/task/common-sense-reasoning","datasets_with_task":"/datasets/task/common-sense-reasoning"}],"languages":[{"name":"Russian","url":"/datasets/language/russian"}],"variants":["PARus"],"data_loaders":[],"num_papers_in_archive":7,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/common-sense-reasoning-on-parus","task":"Common Sense Reasoning","dataset_variant":"PARus","rows":22,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"Human Benchmark","paper":"/paper/russiansuperglue-a-russian-language","metrics":{"Accuracy":"0.982"},"code_links":[{"title":"RussianNLP/RussianSuperGLUE","url":"https://github.com/RussianNLP/RussianSuperGLUE"},{"title":"RussianNLP/MOROCCO","url":"https://github.com/RussianNLP/MOROCCO"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/unreasonable-effectiveness-of-rule-based","title":"Unreasonable Effectiveness of Rule-Based Heuristics in Solving Russian SuperGLUE Tasks","date":"2021-05-03","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/russiansuperglue-a-russian-language","title":"RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark","date":"2020-10-29","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/mt5-a-massively-multilingual-pre-trained-text","title":"mT5: A massively multilingual pre-trained text-to-text transformer","date":"2020-10-22","rows_on_this_dataset":1,"code_links":8,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":0,"samples_unverified":13,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":2,"samples_harvested":14,"samples_ran":1,"samples_unverified":13,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}