Papers › SimRelUz: Similarity and Relatedness scores as a Semantic Evaluation dataset for Uzbek language

SimRelUz: Similarity and Relatedness scores as a Semantic Evaluation dataset for Uzbek language

12 May 2022SIGUL (LREC) 2022 6arXiv:2205.06072archive 2025-07-28

Ulugbek Salaev, Elmurod Kuriyozov, Carlos Gómez-Rodríguez

Semantic relatedness between words is one of the core concepts in natural language processing, thus making semantic evaluation an important task. In this paper, we present a semantic model evaluation dataset: SimRelUz - a collection of similarity and relatedness scores of word pairs for the low-resource Uzbek language. The dataset consists of more than a thousand pairs of words carefully selected based on their morphological features, occurrence frequency, semantic relation, as well as annotated by eleven native Uzbek speakers from different age groups and gender. We also paid attention to the problem of dealing with rare words and out-of-vocabulary words to thoroughly evaluate the robustness of semantic models.

PaperPDFConference PDFCode

Code

ulugbeksalaev/simreluz officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections