Papers › Evaluation of Croatian Word Embeddings

Evaluation of Croatian Word Embeddings

6 Nov 2017LREC 2018 5arXiv:1711.01804archive 2025-07-28

Lukas Svoboda, Slobodan Beliga

Croatian is poorly resourced and highly inflected language from Slavic language family. Nowadays, research is focusing mostly on English. We created a new word analogy corpus based on the original English Word2vec word analogy corpus and added some of the specific linguistic aspects from Croatian language. Next, we created Croatian WordSim353 and RG65 corpora for a basic evaluation of word similarities. We compared created corpora on two popular word representation models, based on Word2Vec tool and fastText tool. Models has been trained on 1.37B tokens training data corpus and tested on a new robust Croatian word analogy corpus. Results show that models are able to create meaningful word representation. This research has shown that free word order and the higher morphological complexity of Croatian language influences the quality of resulting word embeddings.

PaperPDFConference PDFCode

Code

Svobikl/cr-analogy officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Word Embeddings

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

fastText

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections