{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cosimlex-a-resource-for-evaluating-graded","title":"CoSimLex: A Resource for Evaluating Graded Word Similarity in Context","arxiv_id":"1912.05320","date":"2019-12-11","proceeding":"LREC 2020 5","authors":["Carlos Santos Armendariz","Matthew Purver","Matej Ulčar","Senja Pollak","Nikola Ljubešić","Marko Robnik-Šikonja","Mark Granroth-Wilding","Kristiina Vaik"],"abstract":"State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic evaluation of embeddings are based on judgements of similarity, but ignore context; standard tasks for word sense disambiguation take account of context but do not provide continuous measures of meaning similarity. This paper describes an effort to build a new dataset, CoSimLex, intended to fill this gap. Building on the standard pairwise similarity task of SimLex-999, it provides context-dependent similarity measures; covers not only discrete differences in word sense but more subtle, graded changes in meaning; and covers not only a well-resourced language (English) but a number of less-resourced languages. We define the task and evaluation metrics, outline the dataset collection methodology, and describe the status of the dataset so far.","url_abs":"https://arxiv.org/abs/1912.05320v3","url_pdf":"https://arxiv.org/pdf/1912.05320v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cosimlex-a-resource-for-evaluating-graded","repo_url":"https://github.com/lilytang2017/semeval2020","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"word-sense-disambiguation","task_name":"Word Sense Disambiguation"},{"task_slug":"word-similarity","task_name":"Word Similarity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1912.05320","atlas_url":"https://app.syntology.ai/?focus=1912.05320","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}