{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sart-similarity-analogies-and-relatedness-for","title":"SART - Similarity, Analogies, and Relatedness for Tatar Language: New Benchmark Datasets for Word Embeddings Evaluation","arxiv_id":"1904.00365","date":"2019-03-31","proceeding":null,"authors":["Albina Khusainova","Adil Khan","Adín Ramírez Rivera"],"abstract":"There is a huge imbalance between languages currently spoken and\ncorresponding resources to study them. Most of the attention naturally goes to\nthe \"big\" languages: those which have the largest presence in terms of media\nand number of speakers. Other less represented languages sometimes do not even\nhave a good quality corpus to study them. In this paper, we tackle this\nimbalance by presenting a new set of evaluation resources for Tatar, a language\nof the Turkic language family which is mainly spoken in Tatarstan Republic,\nRussia.\n  We present three datasets: Similarity and Relatedness datasets that consist\nof human scored word pairs and can be used to evaluate semantic models; and\nAnalogies dataset that comprises analogy questions and allows to explore\nsemantic, syntactic, and morphological aspects of language modeling. All three\ndatasets build upon existing datasets for the English language and follow the\nsame structure. However, they are not mere translations. They take into account\nspecifics of the Tatar language and expand beyond the original datasets. We\nevaluate state-of-the-art word embedding models for two languages using our\nproposed datasets for Tatar and the original datasets for English and report\nour findings on performance comparison.","url_abs":"http://arxiv.org/abs/1904.00365v1","url_pdf":"http://arxiv.org/pdf/1904.00365v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sart-similarity-analogies-and-relatedness-for","repo_url":"https://github.com/tat-nlp/SART","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"embeddings-evaluation","task_name":"Embeddings Evaluation"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[{"slug":"sart","name":"SART","full_name":null}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}