{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/word-similarity-datasets-for-thai","title":"Word Similarity Datasets for Thai: Construction and Evaluation","arxiv_id":"1904.04307","date":"2019-04-08","proceeding":null,"authors":["Ponrudee Netisopakul","Gerhard Wohlgenannt","Aleksei Pulich"],"abstract":"Distributional semantics in the form of word embeddings are an essential\ningredient to many modern natural language processing systems. The\nquantification of semantic similarity between words can be used to evaluate the\nability of a system to perform semantic interpretation. To this end, a number\nof word similarity datasets have been created for the English language over the\nlast decades. For Thai language few such resources are available. In this work,\nwe create three Thai word similarity datasets by translating and re-rating the\npopular WordSim-353, SimLex-999 and SemEval-2017-Task-2 datasets. The three\ndatasets contain 1852 word pairs in total and have different characteristics in\nterms of difficulty, domain coverage, and notion of similarity (relatedness\nvs.~similarity). These features help to gain a broader picture of the\nproperties of an evaluated word embedding model. We include baseline\nevaluations with existing Thai embedding models, and identify the high ratio of\nout-of-vocabulary words as one of the biggest challenges. All datasets,\nevaluation results, and a tool for easy evaluation of new Thai embedding models\nare available to the NLP community online.","url_abs":"http://arxiv.org/abs/1904.04307v1","url_pdf":"http://arxiv.org/pdf/1904.04307v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"word-similarity-datasets-for-thai","repo_url":"https://github.com/gwohlgen/thai_word_similarity","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"task-2","task_name":"Task 2"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"word-similarity","task_name":"Word Similarity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}