{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/card-660-cambridge-rare-word-dataset-a","title":"Card-660: Cambridge Rare Word Dataset - a Reliable Benchmark for Infrequent Word Representation Models","arxiv_id":"1808.09308","date":"2018-08-28","proceeding":"EMNLP 2018 10","authors":["Mohammad Taher Pilehvar","Dimitri Kartsaklis","Victor Prokhorov","Nigel Collier"],"abstract":"Rare word representation has recently enjoyed a surge of interest, owing to\nthe crucial role that effective handling of infrequent words can play in\naccurate semantic understanding. However, there is a paucity of reliable\nbenchmarks for evaluation and comparison of these techniques. We show in this\npaper that the only existing benchmark (the Stanford Rare Word dataset) suffers\nfrom low-confidence annotations and limited vocabulary; hence, it does not\nconstitute a solid comparison framework. In order to fill this evaluation gap,\nwe propose CAmbridge Rare word Dataset (Card-660), an expert-annotated word\nsimilarity dataset which provides a highly reliable, yet challenging, benchmark\nfor rare word representation techniques. Through a set of experiments we show\nthat even the best mainstream word embeddings, with millions of words in their\nvocabularies, are unable to achieve performances higher than 0.43 (Pearson\ncorrelation) on the dataset, compared to a human-level upperbound of 0.90. We\nrelease the dataset and the annotation materials at\nhttps://pilehvar.github.io/card-660/.","url_abs":"http://arxiv.org/abs/1808.09308v1","url_pdf":"http://arxiv.org/pdf/1808.09308v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"word-similarity","task_name":"Word Similarity"}],"methods":[],"datasets_introduced":[{"slug":"card-660","name":"CARD-660","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1808.09308","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}