{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/russe2018-a-shared-task-on-word-sense","title":"RUSSE'2018: A Shared Task on Word Sense Induction for the Russian Language","arxiv_id":"1803.05795","date":"2018-03-15","proceeding":null,"authors":["Alexander Panchenko","Anastasiya Lopukhina","Dmitry Ustalov","Konstantin Lopukhin","Nikolay Arefyev","Alexey Leontyev","Natalia Loukachevitch"],"abstract":"The paper describes the results of the first shared task on word sense\ninduction (WSI) for the Russian language. While similar shared tasks were\nconducted in the past for some Romance and Germanic languages, we explore the\nperformance of sense induction and disambiguation methods for a Slavic language\nthat shares many features with other Slavic languages, such as rich morphology\nand virtually free word order. The participants were asked to group contexts of\na given word in accordance with its senses that were not provided beforehand.\nFor instance, given a word \"bank\" and a set of contexts for this word, e.g.\n\"bank is a financial institution that accepts deposits\" and \"river bank is a\nslope beside a body of water\", a participant was asked to cluster such contexts\nin the unknown in advance number of clusters corresponding to, in this case,\nthe \"company\" and the \"area\" senses of the word \"bank\". For the purpose of this\nevaluation campaign, we developed three new evaluation datasets based on sense\ninventories that have different sense granularity. The contexts in these\ndatasets were sampled from texts of Wikipedia, the academic corpus of Russian,\nand an explanatory dictionary of Russian. Overall, 18 teams participated in the\ncompetition submitting 383 models. Multiple teams managed to substantially\noutperform competitive state-of-the-art baselines from the previous years based\non sense embeddings.","url_abs":"http://arxiv.org/abs/1803.05795v3","url_pdf":"http://arxiv.org/pdf/1803.05795v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"word-sense-induction","task_name":"Word Sense Induction"}],"methods":[],"datasets_introduced":[{"slug":"russe","name":"RUSSE","full_name":"Russian Words in Context (based on RUSSE)"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}