{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/size-vs-structure-in-training-corpora-for","title":"Size vs. Structure in Training Corpora for Word Embedding Models: Araneum Russicum Maximum and Russian National Corpus","arxiv_id":"1801.06407","date":"2018-01-19","proceeding":null,"authors":["Andrey Kutuzov","Maria Kunilovskaya"],"abstract":"In this paper, we present a distributional word embedding model trained on\none of the largest available Russian corpora: Araneum Russicum Maximum (over 10\nbillion words crawled from the web). We compare this model to the model trained\non the Russian National Corpus (RNC). The two corpora are much different in\ntheir size and compilation procedures. We test these differences by evaluating\nthe trained models against the Russian part of the Multilingual SimLex999\nsemantic similarity dataset. We detect and describe numerous issues in this\ndataset and publish a new corrected version. Aside from the already known fact\nthat the RNC is generally a better training corpus than web corpora, we\nenumerate and explain fine differences in how the models process semantic\nsimilarity task, what parts of the evaluation set are difficult for particular\nmodels and why. Additionally, the learning curves for both models are\ndescribed, showing that the RNC is generally more robust as training material\nfor this task.","url_abs":"http://arxiv.org/abs/1801.06407v1","url_pdf":"http://arxiv.org/pdf/1801.06407v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"size-vs-structure-in-training-corpora-for","repo_url":"https://github.com/natasha/navec","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}