Papers › Word Embeddings for the Armenian Language: Intrinsic and Extrinsic Evaluation

Word Embeddings for the Armenian Language: Intrinsic and Extrinsic Evaluation

7 Jun 2019arXiv:1906.03134archive 2025-07-28

Karen Avetisyan, Tsolak Ghukasyan

In this work, we intrinsically and extrinsically evaluate and compare existing word embedding models for the Armenian language. Alongside, new embeddings are presented, trained using GloVe, fastText, CBOW, SkipGram algorithms. We adapt and use the word analogy task in intrinsic evaluation of embeddings. For extrinsic evaluation, two tasks are employed: morphological tagging and text classification. Tagging is performed on a deep neural network, using ArmTDP v2.3 dataset. For text classification, we propose a corpus of news articles categorized into 7 classes. The datasets are made public to serve as benchmarks for future models.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

ispras-texterra/word-embeddings-eval-hy officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ArticlesClassificationGeneral ClassificationMorphological TaggingText ClassificationWord Embeddingstext-classification

Datasets

Introduced by this paper, per the archive.

iLur News Texts

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

GloVefastText

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections