{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evaluation-of-croatian-word-embeddings","title":"Evaluation of Croatian Word Embeddings","arxiv_id":"1711.01804","date":"2017-11-06","proceeding":"LREC 2018 5","authors":["Lukas Svoboda","Slobodan Beliga"],"abstract":"Croatian is poorly resourced and highly inflected language from Slavic\nlanguage family. Nowadays, research is focusing mostly on English. We created a\nnew word analogy corpus based on the original English Word2vec word analogy\ncorpus and added some of the specific linguistic aspects from Croatian\nlanguage. Next, we created Croatian WordSim353 and RG65 corpora for a basic\nevaluation of word similarities. We compared created corpora on two popular\nword representation models, based on Word2Vec tool and fastText tool. Models\nhas been trained on 1.37B tokens training data corpus and tested on a new\nrobust Croatian word analogy corpus. Results show that models are able to\ncreate meaningful word representation. This research has shown that free word\norder and the higher morphological complexity of Croatian language influences\nthe quality of resulting word embeddings.","url_abs":"http://arxiv.org/abs/1711.01804v2","url_pdf":"http://arxiv.org/pdf/1711.01804v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evaluation-of-croatian-word-embeddings","repo_url":"https://github.com/Svobikl/cr-analogy","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[{"method_slug":"fasttext","method_name":"fastText"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}