Papers › HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings

HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings

16 Nov 2024arXiv:2411.10724archive 2025-07-28

Anton Alekseev, Gulnara Kabaeva

One of the key tasks in modern applied computational linguistics is constructing word vector representations (word embeddings), which are widely used to address natural language processing tasks such as sentiment analysis, information extraction, and more. To choose an appropriate method for generating these word embeddings, quality assessment techniques are often necessary. A standard approach involves calculating distances between vectors for words with expert-assessed 'similarity'. This work introduces the first 'silver standard' dataset for such tasks in the Kyrgyz language, alongside training corresponding models and validating the dataset's suitability through quality evaluation metrics.

PaperPDFCode

Code

alexeyev/awesome-kyrgyz-nlp officialmentioned on GitHub report
alexeyev/kyrgyz-embedding-evaluation officialmentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Sentiment AnalysisWord Embeddings

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections